Research & Papers

AI Phone Helpers Still Flunk Real-Life Tests

Your phone's AI assistant can't even book a flight without messing up

Deep Dive

A new benchmark called GMA puts AI mobile assistants through 300 real-world tasks across seven open-source apps, ranging from simple actions to complex multi-step workflows. When eight top models were tested, performance dropped sharply as task complexity increased, and current agents remain far from reliably handling realistic user requirements. The study also shows that smart harness design can meaningfully improve performance on demanding tasks—though the best approach varies by model.

Key Points
  • AI phone helpers fail 90% of complex real-life tasks like booking trips or comparing flights
  • Researchers created 300 tough tasks to test AI assistants, and even the best ones only got 10% right
  • Better AI design could help, but for now, don’t rely on AI to plan your vacation

Why It Matters

Your phone's AI helper can't reliably book a hotel or compare flights—so you'll still need to double-check (or do it yourself)

📬 Get the top 10 AI stories daily