Image & Video

Polish researchers slash traffic camera errors by 52%

New method cuts road surveillance localization errors by over half using AI vision

Deep Dive

A two-stage geometry-aware pipeline using a YOLO26-based detector and a ResNet34 regression network can locate vehicles far more accurately from monocular surveillance camera imagery. Trained on synthetic CARLA data and fine-tuned on real-world DAIR-V2X footage, the method predicts the projected vehicle footprint on the road plane. On DAIR-V2X, mean image-space localization error dropped from 31.77 px to 15.30 px—a 51.8% improvement—with median error falling to 4.29 px. Median ground-plane error improved from 5.52 m to 0.90 m for medium-range vehicles and from 8.67 m to 1.84 m for far-range vehicles. The findings also show that context around the detector bounding box plays a key role, with the biggest gains seen for distant vehicles and cases with strong perspective distortion and parallax.

Key Points
  • New AI pipeline (YOLO26 + ResNet34) reduces median localization error by 51.8% compared to traditional bounding-box methods
  • System trained on synthetic CARLA data and fine-tuned on real DAIR-V2X footage from roadside cameras
  • Achieves 0.90m median ground-plane error for medium-range vehicles vs. previous 5.52m baseline

Why It Matters

This advancement enables more reliable traffic monitoring, autonomous vehicle positioning, and collision prediction systems using existing surveillance infrastructure.

📬 Get the top 10 AI stories daily