Research & Papers

Qwen-VL boosts robot proxemic danger detection, study finds limits

VLMs near random on proxemic danger until fine-tuned; Qwen-VL leads.

Deep Dive

A new arXiv paper (2608.12515) from Vladyslava Rudas and Dmytro Kuzmenko, accepted at the EMR 2026 workshop at ECCV 2026, evaluates whether open-source vision-language models (VLMs) can assess proxemic risk from a robot's egocentric perspective. The study tested three models—InternVL, Qwen-VL, and SmolVLM—on classifying robot-captured images into four danger levels using three prompting strategies and two rounds of QLoRA fine-tuning, benchmarked against a stratified random baseline.

Results show that without fine-tuning, all models perform near the baseline, while fine-tuning yields only modest overall improvements. Qwen-VL with an advanced prompt stands out, achieving substantially higher recall for high-danger cases than the other models. However, analysis of person localization reveals that correct danger classification does not correspond to better spatial grounding—a model can produce a useful safety label without actually attending to the relevant region. This suggests current VLMs remain limited in fine-grained proxemic reasoning and spatial grounding, though targeted prompting and fine-tuning can improve high-danger detection in select models.

Key Points
  • Qwen-VL with an advanced prompt achieved substantially higher high-danger recall than InternVL and SmolVLM.
  • Without fine-tuning, all three VLMs performed near the stratified random baseline on four-level proxemic risk classification.
  • Correct danger labels did not correlate with better spatial grounding, revealing a gap between classification and scene understanding.

Why It Matters

Robots need reliable proxemic risk assessment for safe human interaction; current VLMs still fall short, but targeted tuning helps.

📬 Get the top 10 AI stories daily