Image & Video

E-VLA boosts robot manipulation in dark/blurred scenes from 0% to 90%

Event-augmented VLA model sees in the dark where traditional cameras fail completely.

Deep Dive

Robotic vision-language-action (VLA) models have excelled at open-ended manipulation, but they collapse under sensing degradations like extreme low light, motion blur, or black clipping. A new paper accepted to ECCV 2026 introduces E-VLA (Event-Augmented Vision-Language-Action Model), a framework that directly integrates event streams—asynchronous pixel-level brightness changes—into VLA pipelines without reconstructing images. Built by Jiajun Zhai, Hao Shi, and collaborators from Zhejiang University, the system uses a DAVIS346 event camera and a custom teleoperation platform to collect a synchronized RGB-event-action manipulation dataset across diverse tasks and illumination levels.

The key innovation is treating event data as an additional perceptual modality rather than a reconstruction source. E-VLA explores lightweight, pretrained-compatible fusion strategies—from parameter-free overlays to learned event adapters. Results are striking: at 20 lux (roughly dim indoor lighting), a standard image-only VLA achieves 0% success on a Pick-Place task. Simply overlaying accumulated event maps onto RGB frames boosts success to 60%, and a dedicated event adapter reaches 90%. Under severe motion blur simulated with 1000 ms exposure, the image-only baseline again scores 0%, while the overlay fusion reaches 20–25%, and a more complex sorting task improves from 5% to 32.5%. These results provide systematic evidence that event-driven perception can be effectively fused into VLA models, pointing toward robust embodied intelligence beyond conventional frame-based imaging. Code and the new dataset will be released upon publication.

Key Points
  • At 20 lux illumination, image-only VLA fails entirely (0% Pick-Place success); E-VLA with event adapter achieves 90%.
  • Under severe motion blur (1000 ms exposure), E-VLA boosts Pick-Place from 0% to 25% and Sorting from 5% to 32.5%.
  • E-VLA uses a DAVIS346 event camera with lightweight fusion strategies (overlay, event adapter) and does not require image reconstruction.

Why It Matters

Event cameras let robots operate reliably in dark, blurry conditions where traditional vision fails, unlocking real-world warehouse, night-time, and high-speed tasks.

📬 Get the top 10 AI stories daily