Robotics

Shift & Drift benchmark shows RL outperforms imitation in AV planning

New zero-shot test evaluates planners under spatial shifts and dynamic noise.

Deep Dive

Shift & Drift is a novel benchmark designed to stress-test autonomous driving motion planners beyond in-distribution performance. It comprises two tracks: (1) The Semantic Shift Track transforms aerial imagery from the DeepScenario Open 3D dataset into the nuPlan simulation framework, enabling zero-shot evaluation on 1,182 scenarios across four German cities and San Francisco with dense pedestrian–cyclist interactions. (2) The State-Distribution Drift Track injects stochastic perturbations into the ego vehicle's dynamics to test robustness against compounding execution errors.

Evaluating diverse planning paradigms, the authors found that imitation learning (IL) methods perform well on in-distribution benchmarks but exhibit significant failures under semantic shift—especially in pedestrian-dense environments—and suffer persistent drift under correlated actuation noise. In contrast, reinforcement-learning-based planners demonstrate more graceful degradation, maintaining higher safety and progress metrics. The findings highlight an empirical trade-off between imitation fidelity and closed-loop resilience, offering a rigorous benchmark for progress toward reliable autonomous driving deployment.

Key Points
  • Semantic Shift Track uses 1,182 scenarios from 4 German cities and San Francisco via a novel aerial-to-simulation conversion pipeline.
  • State-Distribution Drift Track injects stochastic perturbations into ego vehicle dynamics to test robustness against execution noise.
  • Reinforcement learning planners degrade gracefully under shifts, while imitation learning fails in pedestrian-dense environments and under drift.

Why It Matters

Highlights that RL may be crucial for safe autonomous driving in unfamiliar or perturbed conditions.

📬 Get the top 10 AI stories daily