Skip to content

About

Investigating whether cross-architecture model disagreement is a useful, training-free signal for detecting mislabeled images, and specifically when it beats a training-based detector like Cleanlab.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

disagreement-noise

Investigating whether cross-architecture model disagreement is a useful, training-free signal for detecting mislabeled images, and specifically when it beats a training-based detector like Cleanlab.

The core question: when several frozen, pretrained vision-language models with different inductive biases and training data all disagree on an image's label, does that disagreement predict annotation errors? And is it competitive with methods that require training a classifier on the noisy labels?

Short answer, after three datasets: no, and the reason is the interesting part.


TL;DR of findings

  • Training-free cross-architecture disagreement is dominated by a frozen-feature + Cleanlab baseline on every dataset tested (CIFAR-10N, and gate-tested on CUB-200 and EuroSAT). Adding disagreement on top of Cleanlab dilutes precision rather than complementing it.
  • The mechanism: a linear probe on frozen pretrained features reconstructs the true classes almost perfectly, even under 40% label noise and even in domains where zero-shot prompting fails badly. So Cleanlab's "probe confidently disagrees with the given label" flag is near-perfect, and disagreement inherits zero-shot's weakness with no compensating advantage.
  • The surprising part (the dissociation): on EuroSAT, zero-shot CLIP is barely functional (~32–49% top-1) yet a linear probe on the same frozen features hits 94.6%. Weak zero-shot does not imply weak features. This is the crux of why the training-free approach has almost no domain where it wins.
  • Reusable contribution: a one-number precondition gate — linear-probe accuracy on frozen features — that tells you before building a pipeline whether a training-free detector has any hope in your domain.

This is an honest negative result with a clear mechanism and a reusable diagnostic. The experimental phase is closed.


Method

Four frozen, zero-shot OpenCLIP models, chosen for architectural and training-data diversity:

Model Backbone Pretraining
RN50 ResNet (CNN) OpenAI WIT
ViT-B-16 ViT LAION-2B
convnext_base_w Modern CNN LAION-2B
ViT-B-32 ViT (different patch/data) OpenAI WIT

All models are frozen. No fine-tuning is used for the disagreement method — "no training" is the entire point of the comparison. Each model produces a softmax over the dataset's class names via zero-shot prompting.

Contention scores (all: higher = more likely mislabeled), computed only from model outputs, never from labels:

  • hard_disagreement — fraction of model pairs whose argmax predictions differ
  • soft_disagreement — mean symmetric pairwise KL divergence between softmax distributions
  • confidence_weighted — hard disagreement weighted by mean model confidence

Each also has a "strong 3" variant that drops the weakest model (RN50).

Baseline (Cleanlab): a lightweight linear probe is trained on frozen ViT-B-16 embeddings using the noisy labels, with 5-fold cross-validation to produce leak-free out-of-fold predictions. cleanlab.rank.get_label_quality_scores then ranks likely errors. The probe is the only trained component in the repo; no vision model is fine-tuned on these images (that would leak, since CIFAR-10N's images are the CIFAR-10 train split with clean-label ground truth).

Evaluation: precision@K and AUROC of each score against the known mislabel mask. Across noise regimes, AUROC and lift-over-baseline are the honest metrics — raw precision@K inflates with the base rate.


Results: CIFAR-10N

Aggregate regime (~9% noise)

method P@500 P@2000 P@10000 AUROC
cleanlab 1.000 0.993 0.444 0.990
hard (all 4) 0.248 0.242 0.174 0.659
soft (strong 3) 0.130 0.158 0.170 0.681
conf-weighted (all 4) 0.220 0.195 0.134 0.634

Baseline mislabel rate: 9.01%.

Worse regime (~40% noise)

method P@500 P@2000 P@10000 AUROC
cleanlab 1.000 1.000 0.999 0.988
hard (all 4) 0.628 0.629 0.540 0.591
soft (strong 3) 0.488 0.514 0.540 0.618

Baseline mislabel rate: 40.21%. (Raw precision rises only because 40% of images are now errors; lift over baseline falls vs the aggregate regime.)

Why Cleanlab wins (the mechanism)

On the mislabeled subset, the frozen-feature probe predicts the clean label ~85% (aggregate) / ~89% (worse) of the time and the noisy label <8%. The features are strong enough to reconstruct the true class even while trained on noisy labels, so "probe confidently disagrees with the given label" is a near-perfect error flag. Verified leak-free: out-of-fold refit matches stored predictions to 3e-8, one prediction per image, no duplicate embeddings.

Implication: cross-architecture disagreement can only outperform a feature-probe baseline where the features themselves are weak — i.e. domains far from the encoders' pretraining distribution. That motivated the hard-domain phase below.


Hard-domain phase: the precondition gate

Hypothesis: in fine-grained / out-of-distribution domains where a linear probe on frozen features is weak, training-free disagreement might retain signal the probe-based baseline loses.

Gate: before building a full pipeline on a new dataset, measure linear-probe test accuracy on frozen ViT-B-16 features. Proceed only if the probe is weak (< ~0.60) and zero-shot models carry above-chance signal with meaningful disagreement. If the probe exceeds ~0.70, the domain is too easy for the features and the pipeline would just reproduce the CIFAR loss.

Gate results

dataset probe acc zero-shot top-1 (range) mean pairwise disagree verdict
CUB-200-2011 (fine-grained birds) 0.805 0.43 – 0.71 0.50 too easy — do not build
EuroSAT RGB (satellite land cover) 0.946 0.32 – 0.49 0.56 too easy — do not build

Both failed the gate, so the full noise pipeline was not run on either — by design, following a stopping rule fixed before each test.

The dissociation (the key result)

EuroSAT is the sharpest data point. Zero-shot CLIP can barely read satellite tiles (best model 48.6%, two models near ~32–37%), yet a linear probe on those same frozen embeddings reaches 94.6%. The class structure is fully present in the features; zero-shot prompting simply can't access it, but a cheap trained probe can.

This kills the hard-domain hypothesis in the most informative way. "Find a domain where features are weak" turns out to be much harder than "find a domain where zero-shot is weak," because across easy (CIFAR), fine-grained (CUB), and out-of-distribution (EuroSAT) settings, frozen features remain linearly separable into the true classes. A probe-based detector exploits that separability directly; training-free disagreement does not, so it is dominated essentially everywhere a strong encoder exists.


Conclusion

Cross-architecture disagreement is a real but weak label-noise signal (best AUROC ~0.68 on CIFAR-10N) that is dominated by a frozen-feature linear-probe + Cleanlab across every regime and domain tested. The dominance is not a quirk of low noise: it holds at 40% noise and in domains where zero-shot prompting collapses. The practical takeaway is the precondition gate — check linear-probe accuracy on frozen features first; if it is high, a training-free disagreement detector has no room to win, and no amount of dataset-hunting changes that.

Negative result, cleanly established, with a mechanism and a reusable diagnostic.


Repo structure

disagreement-noise/
├── data/                  # CIFAR-10_human.pt, cub/, eurosat/ (all gitignored)
├── inference/             # four zero-shot model scripts -> outputs/*_probs.npy
├── analysis/
│   ├── contention.py      # disagreement scoring
│   ├── eval.py            # noise loading (regime-aware), precision@K, AUROC
│   ├── overlap.py         # disagreement vs Cleanlab overlap
│   └── compare_regimes.py # aggregate vs worse side-by-side
├── baselines/
│   └── cleanlab_baseline.py   # frozen-feature probe + Cleanlab (regime-aware)
├── precondition/
│   ├── cub_probe_test.py       # gate: CUB-200
│   └── eurosat_probe_test.py   # gate: EuroSAT RGB
├── outputs/               # .npy arrays, figures (gitignored)
└── notebooks/

Pipeline order

  1. Run the four inference scripts → outputs/*_probs.npy (softmax per model)
  2. analysis/contention.py → six disagreement scores (label-independent, fixed)
  3. analysis/eval.py → precision@K + AUROC vs the mislabel mask
  4. baselines/cleanlab_baseline.py → probe + Cleanlab (per regime)
  5. analysis/overlap.py / compare_regimes.py → head-to-head
  6. precondition/*_probe_test.py → gate before any new-dataset pipeline

Conventions

pathlib throughout; seed 42; batch 256, torch.no_grad(); skip-if-exists on all cached arrays; all saved arrays float32. Disagreement scores never depend on labels and are shared across regimes.

Data

  • CIFAR-10N: CIFAR-10_human.pt from UCSC-REAL/cifar-10-100n, placed in data/. Keys: aggre_label (~9%), worse_label (~40%), clean_label (ground truth).
  • CUB-200-2011: standard release in data/cub/CUB_200_2011/. Not downloaded programmatically.
  • EuroSAT RGB: 3-band release in data/eurosat/2750/. Not downloaded programmatically.

Status log

  • Phase 1 — CIFAR-10N, aggregate: disagreement signal real but weak (best AUROC 0.68); Cleanlab near-perfect (AUROC 0.990). Confirmed leak-free via out-of-fold refit diagnostics.
  • Phase 2 — CIFAR-10N, worse (40%): tested whether heavy noise breaks the training-based baseline. It did not — the probe stays accurate on mislabeled images (predicts clean label ~89%), Cleanlab still dominant. Hypothesis rejected.
  • Phase 3 — hard-domain pivot, gated: CUB-200 probe 0.805, EuroSAT probe 0.946 — both too easy for the features, pipeline not built. EuroSAT revealed the weak-zero-shot / strong-probe dissociation that explains the whole negative result.
  • Phase closed. Deliverable: negative result + precondition-gate methodology. Next step is a short writeup, not more experiments.

About

Investigating whether cross-architecture model disagreement is a useful, training-free signal for detecting mislabeled images, and specifically when it beats a training-based detector like Cleanlab.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages