Research & Papers

New Study: AI Misreads City Data When It Grabs the Biggest Number

Your AI assistant may mistake a busy airport for the strangest thing in town.

Deep Dive

A new benchmark asks whether tool-augmented language models can make baseline-relative comparisons — not just pick the bigger number. URBANCONTRASTIVEQA pairs two urban situations from public mobility data in NYC, Chicago, and Seattle, labeled by how far current activity deviates from that place's historical baseline. Twenty pickups in a quiet neighborhood can be more abnormal than 180 at an airport. Across six instruction-tuned models and five tool-output formats, models given only raw counts often picked the larger number even when it was less abnormal for its zone. Server-computed baseline scores and ordinal labels raised accuracy, but gains varied by model. The conclusion: for heterogeneous urban feeds, tool interfaces need to expose local baselines, not just activity volumes. The pair bank, labels, scoring scripts, and data card are released.

Key Points
  • AI models tend to grab the biggest number instead of the most surprising one, even when that leads to the wrong conclusion.
  • Researchers tested six AI models on real data from New York, Chicago, and Seattle, using 20 pickups versus 180 as an example.
  • Giving AI the local 'normal' for each place — not just raw totals — made answers noticeably more accurate.

Why It Matters

Cities, delivery apps, and insurers use AI to spot odd activity; bad comparisons mean missed problems and wasted money.

📬 Get the top 10 AI stories daily