Study Finds Robots Need Simple Commands, Not Natural Speech
The way you phrase a request can decide whether a robot finds the right thing
Imagine asking a warehouse robot to grab "the dented blue box behind the forklift." That skill — spotting a specific object from a plain-language request, called open-vocabulary visual grounding — is what makes robots genuinely useful instead of pre-programmed. Until now, most tests used tidy web photos and one-word labels, which tells us little about how these systems behave in messy real places. So a research team built Pro-Bench: over 13,000 camera frames from tunnels, factories, streets, indoors and outdoors, with 74,500 hand-drawn labels on the objects.
Then they ran 16 different AI vision setups through it, using only real-world photos and no extra training. The headline result is unsettling for anyone hoping to just talk to machines. Most systems (10 of 16) scored highest when given short category words like "ladder" or "helmet." Only one did better with full descriptive sentences. In other words, how you phrase the request can matter as much as what the camera sees.
The second finding is subtler but important. Two systems can look equally accurate on average, yet differ wildly in whether they consistently find the same object when the request is reworded. A robot that finds your toolbox 70% of the time under one phrasing and 30% under another is a robot you cannot trust with a real job. Average scores hide that flakiness.
Why care? Robots are heading into warehouses, farms, hospitals and eventually homes, and the promise is that you describe a task in your own words. This study suggests the reality is closer to giving simple, fixed commands — at least for now. It also gives companies a shared yardstick to measure progress, which usually speeds up improvement. The catch: this is one benchmark, built by one team, testing today's models — not a final verdict on what robots can do next year. But it does explain why your voice assistant still asks you to rephrase.
- Pro-Bench uses 13,000+ real-world photos from tunnels, factories, streets and homes — not the polished web images most AI vision tests rely on.
- Ten of the 16 AI systems tested performed best with short labels like 'cup,' while only one did best with full natural sentences.
- Similar average accuracy scores can hide big differences in reliability, meaning a robot may find the same object easily with one phrasing and fail with another.
- This is lab testing, not a product launch — no robot or app is shipping because of it.
Why It Matters
If you plan to work alongside robots or use voice assistants, expect to speak in simple, fixed phrases for now.