MulRobBench benchmark reveals UAV AI agents fail safety compliance
Best model scores just 0.16 accuracy on security-policy tests across 3,024 scenarios.
Researchers from multiple institutions have released MulRobBench, a decision-level benchmark designed to test whether multimodal UAV agents can follow operational security policies under realistic conditions. The benchmark evaluates vision-language-action (VLA) models across four stages: context understanding, evidence arbitration, degradation-aware reasoning, and risk-aware action planning. With 3,024 curated samples covering 17 task taxonomy nodes and 12 scoring dimensions, MulRobBench goes beyond perception and navigation to assess whether agents respect protocol constraints when making critical decisions.
Results expose major gaps in current models. The best semantic protocol-decision score reached only 0.5141, while the strict mean scoring-dimension accuracy was a dismal 0.1599. A controlled ablation study showed that changing just one modality (visual or text) altered 4-15 action selections per model, confirming both inputs matter but are poorly integrated. The primary causes of instability include modality-trust selection errors, difficulty extracting constraints from natural language, glare in visual feeds, missing data, and operator shorthand. These findings highlight that today's multimodal UAV agents are far from trustworthy for smart-city deployment.
- MulRobBench includes 3,024 samples, 17 task taxonomy nodes, and 12 scoring dimensions for UAV decision-making evaluation.
- Best model achieved only 0.5141 semantic compliance and 0.1599 strict accuracy, indicating poor safety adherence.
- A 20-anchor modality-ablation study revealed visual and textual inputs both influence decisions, with glare and missing data as top failure sources.
Why It Matters
As UAVs become autonomous decision-makers, this benchmark exposes critical safety gaps that must be fixed for smart-city deployment.