Investigating whether cross-architecture model disagreement is a useful, training-free signal for detecting mislabeled images, and specifically when it beats a training-based detector like Cleanlab.
The core question: when several frozen, pretrained vision-language models with different inductive biases and training data all disagree on an image's label, does that disagreement predict annotation errors? And is it competitive with methods that require training a classifier on the noisy labels?
Short answer, after three datasets: no, and the reason is the interesting part.
- Training-free cross-architecture disagreement is dominated by a frozen-feature + Cleanlab baseline on every dataset tested (CIFAR-10N, and gate-tested on CUB-200 and EuroSAT). Adding disagreement on top of Cleanlab dilutes precision rather than complementing it.
- The mechanism: a linear probe on frozen pretrained features reconstructs the true classes almost perfectly, even under 40% label noise and even in domains where zero-shot prompting fails badly. So Cleanlab's "probe confidently disagrees with the given label" flag is near-perfect, and disagreement inherits zero-shot's weakness with no compensating advantage.
- The surprising part (the dissociation): on EuroSAT, zero-shot CLIP is barely functional (~32–49% top-1) yet a linear probe on the same frozen features hits 94.6%. Weak zero-shot does not imply weak features. This is the crux of why the training-free approach has almost no domain where it wins.
- Reusable contribution: a one-number precondition gate — linear-probe accuracy on frozen features — that tells you before building a pipeline whether a training-free detector has any hope in your domain.
This is an honest negative result with a clear mechanism and a reusable diagnostic. The experimental phase is closed.
Four frozen, zero-shot OpenCLIP models, chosen for architectural and training-data diversity:
| Model | Backbone | Pretraining |
|---|---|---|
| RN50 | ResNet (CNN) | OpenAI WIT |
| ViT-B-16 | ViT | LAION-2B |
| convnext_base_w | Modern CNN | LAION-2B |
| ViT-B-32 | ViT (different patch/data) | OpenAI WIT |
All models are frozen. No fine-tuning is used for the disagreement method — "no training" is the entire point of the comparison. Each model produces a softmax over the dataset's class names via zero-shot prompting.
Contention scores (all: higher = more likely mislabeled), computed only from model outputs, never from labels:
hard_disagreement— fraction of model pairs whose argmax predictions differsoft_disagreement— mean symmetric pairwise KL divergence between softmax distributionsconfidence_weighted— hard disagreement weighted by mean model confidence
Each also has a "strong 3" variant that drops the weakest model (RN50).
Baseline (Cleanlab): a lightweight linear probe is trained on frozen ViT-B-16 embeddings using the noisy labels, with 5-fold cross-validation to produce leak-free out-of-fold predictions. cleanlab.rank.get_label_quality_scores then ranks likely errors. The probe is the only trained component in the repo; no vision model is fine-tuned on these images (that would leak, since CIFAR-10N's images are the CIFAR-10 train split with clean-label ground truth).
Evaluation: precision@K and AUROC of each score against the known mislabel mask. Across noise regimes, AUROC and lift-over-baseline are the honest metrics — raw precision@K inflates with the base rate.
| method | P@500 | P@2000 | P@10000 | AUROC |
|---|---|---|---|---|
| cleanlab | 1.000 | 0.993 | 0.444 | 0.990 |
| hard (all 4) | 0.248 | 0.242 | 0.174 | 0.659 |
| soft (strong 3) | 0.130 | 0.158 | 0.170 | 0.681 |
| conf-weighted (all 4) | 0.220 | 0.195 | 0.134 | 0.634 |
Baseline mislabel rate: 9.01%.
| method | P@500 | P@2000 | P@10000 | AUROC |
|---|---|---|---|---|
| cleanlab | 1.000 | 1.000 | 0.999 | 0.988 |
| hard (all 4) | 0.628 | 0.629 | 0.540 | 0.591 |
| soft (strong 3) | 0.488 | 0.514 | 0.540 | 0.618 |
Baseline mislabel rate: 40.21%. (Raw precision rises only because 40% of images are now errors; lift over baseline falls vs the aggregate regime.)
On the mislabeled subset, the frozen-feature probe predicts the clean label ~85% (aggregate) / ~89% (worse) of the time and the noisy label <8%. The features are strong enough to reconstruct the true class even while trained on noisy labels, so "probe confidently disagrees with the given label" is a near-perfect error flag. Verified leak-free: out-of-fold refit matches stored predictions to 3e-8, one prediction per image, no duplicate embeddings.
Implication: cross-architecture disagreement can only outperform a feature-probe baseline where the features themselves are weak — i.e. domains far from the encoders' pretraining distribution. That motivated the hard-domain phase below.
Hypothesis: in fine-grained / out-of-distribution domains where a linear probe on frozen features is weak, training-free disagreement might retain signal the probe-based baseline loses.
Gate: before building a full pipeline on a new dataset, measure linear-probe test accuracy on frozen ViT-B-16 features. Proceed only if the probe is weak (< ~0.60) and zero-shot models carry above-chance signal with meaningful disagreement. If the probe exceeds ~0.70, the domain is too easy for the features and the pipeline would just reproduce the CIFAR loss.
| dataset | probe acc | zero-shot top-1 (range) | mean pairwise disagree | verdict |
|---|---|---|---|---|
| CUB-200-2011 (fine-grained birds) | 0.805 | 0.43 – 0.71 | 0.50 | too easy — do not build |
| EuroSAT RGB (satellite land cover) | 0.946 | 0.32 – 0.49 | 0.56 | too easy — do not build |
Both failed the gate, so the full noise pipeline was not run on either — by design, following a stopping rule fixed before each test.
EuroSAT is the sharpest data point. Zero-shot CLIP can barely read satellite tiles (best model 48.6%, two models near ~32–37%), yet a linear probe on those same frozen embeddings reaches 94.6%. The class structure is fully present in the features; zero-shot prompting simply can't access it, but a cheap trained probe can.
This kills the hard-domain hypothesis in the most informative way. "Find a domain where features are weak" turns out to be much harder than "find a domain where zero-shot is weak," because across easy (CIFAR), fine-grained (CUB), and out-of-distribution (EuroSAT) settings, frozen features remain linearly separable into the true classes. A probe-based detector exploits that separability directly; training-free disagreement does not, so it is dominated essentially everywhere a strong encoder exists.
Cross-architecture disagreement is a real but weak label-noise signal (best AUROC ~0.68 on CIFAR-10N) that is dominated by a frozen-feature linear-probe + Cleanlab across every regime and domain tested. The dominance is not a quirk of low noise: it holds at 40% noise and in domains where zero-shot prompting collapses. The practical takeaway is the precondition gate — check linear-probe accuracy on frozen features first; if it is high, a training-free disagreement detector has no room to win, and no amount of dataset-hunting changes that.
Negative result, cleanly established, with a mechanism and a reusable diagnostic.
disagreement-noise/
├── data/ # CIFAR-10_human.pt, cub/, eurosat/ (all gitignored)
├── inference/ # four zero-shot model scripts -> outputs/*_probs.npy
├── analysis/
│ ├── contention.py # disagreement scoring
│ ├── eval.py # noise loading (regime-aware), precision@K, AUROC
│ ├── overlap.py # disagreement vs Cleanlab overlap
│ └── compare_regimes.py # aggregate vs worse side-by-side
├── baselines/
│ └── cleanlab_baseline.py # frozen-feature probe + Cleanlab (regime-aware)
├── precondition/
│ ├── cub_probe_test.py # gate: CUB-200
│ └── eurosat_probe_test.py # gate: EuroSAT RGB
├── outputs/ # .npy arrays, figures (gitignored)
└── notebooks/
- Run the four inference scripts →
outputs/*_probs.npy(softmax per model) analysis/contention.py→ six disagreement scores (label-independent, fixed)analysis/eval.py→ precision@K + AUROC vs the mislabel maskbaselines/cleanlab_baseline.py→ probe + Cleanlab (per regime)analysis/overlap.py/compare_regimes.py→ head-to-headprecondition/*_probe_test.py→ gate before any new-dataset pipeline
pathlib throughout; seed 42; batch 256, torch.no_grad(); skip-if-exists on all cached arrays; all saved arrays float32. Disagreement scores never depend on labels and are shared across regimes.
- CIFAR-10N:
CIFAR-10_human.ptfrom UCSC-REAL/cifar-10-100n, placed indata/. Keys:aggre_label(~9%),worse_label(~40%),clean_label(ground truth). - CUB-200-2011: standard release in
data/cub/CUB_200_2011/. Not downloaded programmatically. - EuroSAT RGB: 3-band release in
data/eurosat/2750/. Not downloaded programmatically.
- Phase 1 — CIFAR-10N, aggregate: disagreement signal real but weak (best AUROC 0.68); Cleanlab near-perfect (AUROC 0.990). Confirmed leak-free via out-of-fold refit diagnostics.
- Phase 2 — CIFAR-10N, worse (40%): tested whether heavy noise breaks the training-based baseline. It did not — the probe stays accurate on mislabeled images (predicts clean label ~89%), Cleanlab still dominant. Hypothesis rejected.
- Phase 3 — hard-domain pivot, gated: CUB-200 probe 0.805, EuroSAT probe 0.946 — both too easy for the features, pipeline not built. EuroSAT revealed the weak-zero-shot / strong-probe dissociation that explains the whole negative result.
- Phase closed. Deliverable: negative result + precondition-gate methodology. Next step is a short writeup, not more experiments.