Research & Papers

AI Shopping Helpers Still Stumble When Your Request Is Vague

⚡AI can call the tools, but vague requests still stump it — a warning for shoppers.

Deep Dive

Researchers propose RecToolBench, a Model Context Protocol-based benchmark for evaluating tool-using recommender agents under fuzzy user instructions. It contains more than 1,200 executable tasks across three recommendation domains, 13 MCP servers, and 32 tools, spanning single-tool calls, parallel calls, sequential chains, and hybrid orchestration. Built with a synthesize–fuzzify–judge pipeline, it evaluates agent trajectories using rule-based checks and rubric-based LLM evaluation. Experiments on representative LLMs show syntactically valid tool calls don't guarantee successful recommendations: models struggle with semantic parameter grounding, multi-step evidence integration, and grounded final recommendations as orchestration complexity rises. The authors identify tool orchestration under fuzzy user intent as a major bottleneck for agentic recommender systems.

Key Points
  • A new test of 1,200+ vague requests found AI recommendation helpers often fail even when they use their tools correctly.
  • The test spans 32 tools across 13 app connectors and three recommendation areas, so it reflects realistic multi-step requests.
  • Failures get worse as tasks require more steps — like looking up a store, checking stock, then comparing prices.

Why It Matters

Treat AI shopping and travel suggestions as a starting point, not an answer — for now, they often miss.

📬 Get the top 10 AI stories daily