Image & Video

MODEST dataset pushes stereo depth-of-field research forward

50 distinct optical setups, 20MP images, and 20K real-world shots—MODEST is the new benchmark for depth-of-field stereo vision.

Deep Dive

A team led by Nisarg K. Trivedi and Vinayaka A. Belludi at Georgia Tech, in collaboration with Li-Yun Wang, has released MODEST (Multi-Optics Depth-of-Field Stereo Dataset), a groundbreaking resource for computer vision researchers. Published on arXiv and available for non-commercial academic use, MODEST addresses a critical gap in stereo vision datasets by providing 20,000 ultra-high-resolution images (5472×3648 pixels, 20MP) captured across 50 distinct optical configurations using professional stereo DSLRs. Each configuration varies focal length (28–70mm) and aperture (f/2.8–f/22), enabling controlled analysis of geometric and optical effects in shallow depth-of-field (DoF) rendering and defocus deblurring. The dataset includes 10 complex real-world scenes with challenging visual elements such as reflective surfaces, transparent glass, fine-grained details, point lights, and multi-scale depth illusions, making it ideal for stress-testing state-of-the-art (SOTA) models.

MODEST goes beyond raw image data by including full intrinsics and extrinsics calibration images, supporting evolving calibration techniques. The researchers evaluated several SOTA DoF and deblurring methods on MODEST, revealing failure cases and limitations that highlight the dataset’s value in bridging the realism gap between synthetic, low-resolution training data and real-world high-resolution camera optics. The dataset is now available on the project page and Hugging Face for purely academic research, and it promises to accelerate progress in AR/VR, smartphone imaging, industrial robotics, and autonomous systems where accurate depth perception and DoF control are critical.

Key Points
  • MODEST is the first ultra-high-resolution (20MP, 5472×3648px) stereo dataset with 20,000 images across 50 optical setups (focal lengths 28–70mm, apertures f/2.8–f/22).
  • Includes challenging scenes with reflective surfaces, glass, fine details, and multi-scale depth illusions, plus calibration data for real-camera systems.
  • Enables rigorous evaluation of shallow DoF rendering and defocus deblurring models, exposing SOTA weaknesses and bridging synthetic-to-real gaps.

Why It Matters

MODEST provides the missing benchmark for real-world stereo depth-of-field research, accelerating advancements in AR/VR, robotics, and computational photography.

📬 Get the top 10 AI stories daily