EDUARDO CEPEDA · INDEPENDENT CONTROLLED EXPERIMENT · 2026

CNN confidence under corruption.

Mixup improved the measured accuracy and mean corruption NLL. Temperature scaling fitted on clean images helped the baseline under corruption, but worsened mixup's corruption NLL, Brier score, and calibration error.

Completed study: full-2026-10-08. All results below come from six completed training runs and final saved predictions.

What was compared

Six CIFAR-adapted ResNet-18 models trained from random initialization for 100 epochs each: baseline cross-entropy and mixup with alpha 0.2, paired across seeds 42, 43, and 44. Architecture, optimization, augmentation, and data partitions were held fixed. Checkpoints were selected by minimum validation NLL.

A fixed stratified split reserved 40,000 CIFAR-10 images for training, 5,000 for validation, and 5,000 for fitting temperature. Each checkpoint was evaluated on all 10,000 official test images and the 15 standard CIFAR-10-C corruption types at five severities. No temperature was fitted on test or corrupted data.

The primary paired comparison

Mixup minus baseline, after averaging the 75 corruption conditions within each run. Negative NLL differences favor mixup; positive accuracy differences favor mixup. Values are means ± sample standard deviations across three paired seed differences.

Mean corruption results: paired differences
ScalingΔ NLLΔ accuracy (percentage points)
Raw-0.4207 ± 0.04252.2348 ± 0.7332
With T-0.0602 ± 0.02892.2348 ± 0.7332
Each of the three paired effects
SeedΔ NLL, rawΔ NLL, with TΔ accuracy (pp)
42-0.449991-0.0830652.9323
43-0.440005-0.0699072.3017
44-0.371968-0.0277321.4705

Mean clean accuracy rose from 94.19% to 94.62%; mean corruption accuracy rose from 73.64% to 75.88%. Raw corruption NLL fell from 1.274 to 0.853. The direction of the primary NLL contrast favored mixup in all three seeds, both before and after scaling.

Calibration did not transfer uniformly

Temperature scaling reduced clean-test NLL for both methods. Under corruption, baseline NLL improved from 1.274 to 0.967, while mixup NLL increased from 0.853 to 0.907. This opposite direction occurred in all three seeds. For mixup, mean corruption Brier score increased from 0.3593 to 0.3701 and ECE15 from 0.1019 to 0.1219.

Positive scalar temperature leaves argmax predictions and accuracy unchanged. The fitted temperatures were above 1 for baseline and below 1 for mixup. They soften baseline probabilities and sharpen mixup probabilities; the measured effects depend on the test condition and metric.

Clean and corrupted test results: mean ± seed sample SD
Test conditionMethodScalingAccuracy (%)NLLBrierECE15
Clean testBaselineRaw94.1933 ± 0.18580.2235 ± 0.00290.0930 ± 0.00260.0333 ± 0.0014
Clean testMixupRaw94.6167 ± 0.08960.1941 ± 0.00820.0825 ± 0.00170.0262 ± 0.0070
Clean testBaselineWith T94.1933 ± 0.18580.1899 ± 0.00390.0878 ± 0.00220.0092 ± 0.0008
Clean testMixupWith T94.6167 ± 0.08960.1868 ± 0.00280.0832 ± 0.00160.0134 ± 0.0028
75-condition meanBaselineRaw73.6419 ± 0.45221.2739 ± 0.02230.4289 ± 0.00500.1766 ± 0.0011
75-condition meanMixupRaw75.8767 ± 0.79050.8532 ± 0.06410.3593 ± 0.01760.1019 ± 0.0165
75-condition meanBaselineWith T73.6419 ± 0.45220.9674 ± 0.02010.3945 ± 0.00520.1218 ± 0.0016
75-condition meanMixupWith T75.8767 ± 0.79050.9071 ± 0.04350.3701 ± 0.01270.1219 ± 0.0068

NLL is mean negative natural log probability of the true class. Brier is the mean sum of squared probability errors over all ten classes. ECE15 uses 15 equal-width confidence bins. Lower values are better for these three metrics; each measures a different aspect of the predictions.

Original analysis figures

Accuracy and confidence across corruption severities
Accuracy and confidence across corruption severities. Each point averages the 15 corruption types at one severity within each run, then averages the three seeds. Bars are seed sample standard deviations, not confidence intervals. Open full size.
When clean-fitted calibration meets corrupted inputs
When clean-fitted calibration meets corrupted inputs. Scaled minus raw NLL, Brier score, and ECE15. Negative values reduce the metric. The temperature is fitted on clean calibration images and kept fixed for every test condition. Open full size.
Clean test and mean corruption results
Clean test and mean corruption results. Small points show individual seeds. The 75-condition mean is computed within each run; corrupted variants are views of the same 10,000 source images. Open full size.

What the results support

Run evidence

Selected checkpoints after 100 epochs per run
MethodSeedSelected epochTemperature
Baseline42971.468553
Mixup42990.817504
Baseline431001.507146
Mixup43990.862529
Baseline44971.488337
Mixup44940.945717

This page is derived from the completed study's comparison.json, comparison.csv, and original generated figures. The implementation preserves per-condition predictions, IDs, checkpoint and source hashes, calibration records, and evaluation identities. The public repository provides the code and frozen protocol. The full technical report preserves the measured tables and provenance; the results discussion explains the findings and limitations.