Research & Papers

New Test Shows AI Still Gets Lost in Real Rooms

AI names objects in photos — but can't count or judge distance reliably.

Deep Dive

Artificial intelligence is very good at looking at a flat photo and telling you what's in it. Ask it to understand a whole room — how the furniture is arranged, which chair is closer, how many cups sit on the counter — and it starts to stumble. A team of researchers has now built a tougher test called SceneBench to measure exactly how badly it stumbles.

The test is made of 966 computer-generated rooms that look like photographs. The team used a technique called Gaussian Splatting (a way to build lifelike 3D scenes from photos) so the rooms keep real textures, printed labels, and materials instead of being stripped down to plain shapes. They then spent roughly 1,500 human hours labeling everything inside — 183,000 items in total — organized the way people actually think about space: the whole scene, then rooms, then functional zones like a kitchen counter, then groups of objects, then single items.

The results are telling. The best available AI models scored up to 85% when simply identifying objects. But on questions like counting, comparing sizes, estimating distance, and figuring out which direction something faces, accuracy slid to about 60%. Multi-step questions — "is the object on the table bigger than the one behind the sofa?" — were harder still. In other words, AI recognizes things but doesn't truly reason about where they are.

Why should you care? Because the next wave of technology depends on this exact skill. Warehouse robots need to count and grab the right box. Delivery drones need to judge distance. AR glasses that translate signs or give directions need to understand the physical world in front of you. Self-driving cars need to know a pedestrian is behind a parked van, not just that a van exists. SceneBench gives researchers a realistic yardstick to measure progress — meaning better, safer spatial AI is now easier to build and verify. The paper is a benchmark, not a product, so expect improvements over the next year or two.

Key Points
  • Researchers created 966 photorealistic 3D rooms as a test for spatial AI, with 183,000 labeled objects built over about 1,500 human hours.
  • Today's best AI spots objects about 85% of the time but drops to roughly 60% on counting, distance, and direction questions.
  • This matters for robots, delivery drones, AR glasses, and self-driving cars, which all need to understand where things are, not just what they are.

Why It Matters

Better spatial AI means smarter robots, safer self-driving cars, and AR glasses that actually understand your surroundings.

📬 Get the top 10 AI stories daily