← All work

CNN robustness & calibration / AI & apps

Independent controlled experiment

CNN confidence under corruption.

Six ResNet-18 runs test how image corruption changes accuracy and confidence. Mixup improved the measured averages, while calibration fitted on clean images helped the baseline and hurt mixup under corruption.

My role
Experiment design / training / evaluation
When
October 2026
Built with
Python · PyTorch · ResNet-18 · NumPy · CIFAR-10-C · Matplotlib

Outcome

Mixup raised mean clean accuracy from 94.19% to 94.62% and mean corruption accuracy from 73.64% to 75.88%. Unscaled corruption NLL fell from 1.274 to 0.853. Clean-fitted temperature scaling then reduced baseline corruption NLL to 0.967 but increased mixup's to 0.907: a benefit on clean images did not transfer uniformly to corrupted inputs.

The interactive explorer and figures use measured final-test results. Means and sample standard deviations describe three paired training seeds; they are not confidence intervals. Corrupted variants share the same source images, and this study does not establish general real-world or adversarial robustness.

Measured accuracy, negative log-likelihood, and calibration error across five corruption severities for baseline and mixup, before and after temperature scaling
CIFAR-10-C · means across 15 corruption types at each severity; bars show sample standard deviation across three training seeds · Open full size ↗

The challenge

A classifier can become confidently wrong when its inputs change. I wanted to measure whether mixup improves a CNN's predictions under common image corruptions, and whether a temperature fitted on clean held-out data keeps helping when those images shift.

Approach

Keep each data role separate

I froze the comparison before full training and used 40,000 CIFAR-10 images for training, 5,000 for checkpoint selection, and 5,000 for calibration. The official 10,000-image test set and its corrupted versions were reserved for final evaluation.

Compare six controlled training runs

I trained baseline and mixup ResNet-18 models from random initialization for 100 epochs each, pairing seeds 42, 43, and 44. Architecture, optimizer, augmentation, partitions, and checkpoint selection were held fixed; mixup used alpha 0.2.

Test confidence under the same shifts

Each selected checkpoint was evaluated on clean test images and all 15 standard CIFAR-10-C corruption types at five severities. I measured accuracy, negative log-likelihood, Brier score, and 15-bin calibration error before and after a temperature fitted only on the clean calibration split.

Retain the evidence behind the averages

Saved predictions, checkpoint identities, source snapshots, and condition-level metrics support the analysis. Corruption metrics are averaged within each run, then compared within each paired seed. The figures show seed variability and preserve the negative finding about calibration transfer.

Sources & project context

Explore how it works

Try it yourself ↓
Recorded experiment ResNet-18 · CIFAR-10-C

Confidence under corruption.

Follow the same image through the test. See what each model predicts, then compare results across three training seeds.

Temperature scaling

Scaling changes confidence, not the predicted class. Temperatures were fitted on separate clean calibration data.

One image, observed closely

Saved predictions · seed 42 only

Original CIFAR-10 test image 0, labeled cat
Original ID 0
cat test image 0: Gaussian noise, severity 3
Severity 3 True label: cat
BaselineIncorrect
frog99.8%

Confidence · T = 1.00

MixupIncorrect
frog53.4%

Confidence · T = 1.00

Explore severityOriginal → strongest

Fixed test IDs 0, 1 and 4 are illustrative examples, not a representative sample. Confidence is the probability assigned to the predicted class.

The full test, across three seeds

Gaussian noise · severity 3 · 10,000 images per seed

42 / 43 / 44
Mean and sample standard deviation across three training seeds, Gaussian noise · severity 3, before temperature scaling.
MetricBaselineMixup
Accuracy ↑42.34%± 4.20%43.98%± 4.68%
NLL ↓3.205± 0.2532.109± 0.347
ECE ↓43.83%± 3.23%33.01%± 6.48%

Mean ± sample SD across seeds. SD is not a confidence interval. NLL measures log loss; ECE measures calibration error using 15 bins.

Track the change

Accuracy from clean to severity 5, Gaussian noise. Whiskers show sample SD across seeds.Baseline: clean 94.19%, severity 1 80.82%, severity 2 62.17%, severity 3 42.34%, severity 4 34.46%, severity 5 28.89%. Mixup: clean 94.62%, severity 1 81.47%, severity 2 63.07%, severity 3 43.98%, severity 4 36.87%, severity 5 30.91%.0%50%100%Clean12345
BaselineMixupWhiskers: seed SD · x-axis: severity

Measured, saved results. This explorer runs no model. Six local training runs; three paired seeds. Corruptions of the same image are not independent samples.

Data, attribution and limits

ResNet-18 trained from scratch with a CIFAR stem. Baseline and Mixup use the same fixed train, validation and calibration partitions. A positive temperature preserves predicted labels; clean calibration can transfer poorly to corrupted inputs. Three seeds support a descriptive comparison, not a significance claim.

Original pixels: CIFAR-10 · Krizhevsky, Nair & Hinton. Corrupted images: CIFAR-10-C · Hendrycks & Dietterich, distributed under the record’s CC BY 4.0. The lossless PNG exports preserve all 32 × 32 RGB pixels.

Gaussian noise · severity 3. cat example, original ID 0. Before temperature scaling.

More from the project

Measured changes in NLL, Brier score, and calibration error showing clean-fitted temperature scaling helping the baseline and hurting mixup under corruption
Calibration transfer · scaled minus raw metrics from the completed study; negative values reduce the metric · Open full size ↗
Measured clean-test and mean-corruption accuracy, NLL, Brier score, and calibration error with individual seed results
Clean and corrupted test images · original analysis output, with seed means and individual seed points · Open full size ↗
Next projectNexus listings engine ↗