Research & Papers

CILN framework creates realistic label noise benchmarks via controlled corruptions

New benchmark exposes AI training failure modes that standard noise tests miss

Deep Dive

A team of researchers (Shadman Islam, Agustinus Kristiadi, Mostafa Milani) introduced CILN (Controlled Instance-dependent Label Noise), a benchmark generation framework that creates realistic label noise by corrupting input data rather than relying on imperfect human annotators. In CILN, a diverse voter pool labels corrupted instances, making both the source and severity of ambiguity explicit and controllable. The team built 90 benchmark settings across CIFAR-10, MNIST, and Adult datasets, spanning multiple corruption families and severity levels.

Experiments show CILN benchmarks exhibit genuine instance-dependent noise with diverse confusion structures. On CIFAR-10, CILN can produce label distributions closer to human uncertainty than existing synthetic IDN benchmarks. Importantly, CILN exposes failure modes in popular noisy-label learning methods like Co-Teaching and DivideMix that aren't observed under comparable levels of rater-fallibility noise. This suggests noise structure—not just noise rate—plays a critical role in benchmark difficulty and algorithm behavior, providing a complementary framework for studying noisy-label learning under diverse sources of instance difficulty.

Key Points
  • CILN framework generates instance-dependent label noise via controlled input corruptions across 90 benchmark settings on CIFAR-10, MNIST, and Adult
  • On CIFAR-10, CILN produces label distributions closer to human uncertainty than existing synthetic IDN benchmarks
  • Exposes failure modes in Co-Teaching and DivideMix that standard rater-fallibility noise benchmarks do not reveal

Why It Matters

Reveals that noise structure, not just noise rate, critically impacts AI robustness and benchmarking

📬 Get the top 10 AI stories daily