AI Safety

Study finds AI nutrition lenses fail on local products, scoring 59.5% accuracy

AI agents hit 88.9% on global foods but drop to random guessing on Swedish products.

Deep Dive

Can AI agents reliably tell you which supermarket product is healthier? A new paper by Jose Berengueres, accepted to the IEEE EMBC 2026 conference, puts that question to the test. The study, titled "Can We Trust AI Agents in the Supermarket? Sugar Content Inference from Product Images," evaluates vision-capable conversational agents on a bounded, verifiable task: inferring which of two packaged foods contains less sugar using only front-of-pack images. Using a Two-Alternative Forced Choice game, the researcher benchmarked AI systems across four national markets: Sweden, the USA, Australia, and Kazakhstan.

The results show a stark performance divide. Across 132 comparisons, AI agents achieved 88.9% accuracy on global products, a result with p < 0.0001 — clearly above chance. But when tested on local Swedish products, accuracy plummeted to 59.5% (p = 0.29), making the AI's dietary guidance statistically indistinguishable from random guessing. This cross-market bias is likely tied to uneven training-data coverage: AI systems are well-fed on globally distributed brands but poorly calibrated for regional food ecosystems. The findings raise serious concerns about trust, equity, and accountability as consumers increasingly shift nutritional judgment from auditable public labels to proprietary inference pipelines.

Berengueres concludes that AI nutrition lens applications are better framed as assistive, educational tools — not replacements for regulated labeling. He argues for auditable datasets and evaluation benchmarks aligned with local food environments. In an era where "AI nutrition lens" apps are marketed as quick health advisors, this study serves as a sobering reminder: when the stakes are your daily sugar intake, a model's confident answer might just be a guess.

Key Points
  • AI agents scored 88.9% (p < 0.0001) on global supermarket products but only 59.5% (p = 0.29) on local Swedish products, equivalent to random guessing
  • Study tested 132 comparisons across four countries (Sweden, USA, Australia, Kazakhstan) using a Two-Alternative Forced Choice game on front-of-pack images
  • Paper concludes AI nutrition apps should be assistive educational tools, not replacements for regulated labels, citing uneven training-data coverage

Why It Matters

Shows AI nutrition advice can be dangerously unreliable outside training markets, undermining trust in AI-powered health guidance globally.

📬 Get the top 10 AI stories daily