AI Phone Helpers Still Flunk Real-Life Tests
Your phone's AI assistant can't even book a flight without messing up
A new benchmark called GMA puts AI mobile assistants through 300 real-world tasks across seven open-source apps, ranging from simple actions to complex multi-step workflows. When eight top models were tested, performance dropped sharply as task complexity increased, and current agents remain far from reliably handling realistic user requirements. The study also shows that smart harness design can meaningfully improve performance on demanding tasks—though the best approach varies by model.
- AI phone helpers fail 90% of complex real-life tasks like booking trips or comparing flights
- Researchers created 300 tough tasks to test AI assistants, and even the best ones only got 10% right
- Better AI design could help, but for now, don’t rely on AI to plan your vacation
Why It Matters
Your phone's AI helper can't reliably book a hotel or compare flights—so you'll still need to double-check (or do it yourself)