Image & Video

JPEG AIC2026 dataset benchmarks learning-based codecs with 9,618 images

New benchmark dataset reveals major flaws in quality assessment for AI-compressed images

Deep Dive

The JPEG committee and a team of researchers from universities including Konstanz, Lisbon, and others have introduced AIC2026, a comprehensive dataset designed for fine-grained evaluation of image coding quality. The dataset was curated from 2,787 candidate images using semantic clustering and inter-metric disagreement among objective IQA methods, resulting in 70 diverse source images. These were compressed using 12 codecs (8 conventional, 4 learning-based) across 17 configurations. Each source-codec pair produced 20 distortion levels calibrated using the ColorVideoVDP (CVVDP) metric to span 0.2 to 4.0 just-noticeable difference (JND) units, creating a total of 9,618 distorted images. This fine-grained sampling allows detailed rate-distortion analysis and objective metric evaluation for subtle quality differences, particularly crucial for evaluating emerging learning-based compression methods.

Extensive objective analysis applied 24 conventional and 12 learning-based IQA methods to the dataset. The results reveal substantial disagreement among current objective metrics when assessing fine-grained quality differences, especially for artifacts introduced by learning-based codecs—which often exhibit unnatural distortions not well captured by traditional metrics. This highlights a critical gap: existing benchmarks and metrics may not reliably predict perceptual quality for AI-driven compression. The AIC2026 dataset is publicly available, providing a standardized testbed for developing better quality metrics and advancing image compression research. For engineers working on streaming, storage, or real-time imaging, this dataset offers a more realistic and challenging benchmark for comparing codec performance at near-threshold distortion levels.

Key Points
  • Dataset includes 70 source images from 2,787 candidates, compressed by 12 codecs (8 conventional + 4 learning-based) across 17 configurations
  • Each source-codec pair has 20 distortion levels calibrated from 0.2 to 4.0 JND units using CVVDP metric, totaling 9,618 images
  • 36 IQA methods (24 conventional, 12 learning-based) show significant disagreement, especially for learning-based codec artifacts

Why It Matters

This dataset enables rigorous benchmarking of image compression quality, critical for streaming and storage optimization in AI-driven workflows.

📬 Get the top 10 AI stories daily