AI Shopping Helpers Still Stumble When Your Request Is Vague
AI can call the tools, but vague requests still stump it — a warning for shoppers.
Researchers propose RecToolBench, a Model Context Protocol-based benchmark for evaluating tool-using recommender agents under fuzzy user instructions. It contains more than 1,200 executable tasks across three recommendation domains, 13 MCP servers, and 32 tools, spanning single-tool calls, parallel calls, sequential chains, and hybrid orchestration. Built with a synthesize–fuzzify–judge pipeline, it evaluates agent trajectories using rule-based checks and rubric-based LLM evaluation. Experiments on representative LLMs show syntactically valid tool calls don't guarantee successful recommendations: models struggle with semantic parameter grounding, multi-step evidence integration, and grounded final recommendations as orchestration complexity rises. The authors identify tool orchestration under fuzzy user intent as a major bottleneck for agentic recommender systems.
- A new test of 1,200+ vague requests found AI recommendation helpers often fail even when they use their tools correctly.
- The test spans 32 tools across 13 app connectors and three recommendation areas, so it reflects realistic multi-step requests.
- Failures get worse as tasks require more steps — like looking up a store, checking stock, then comparing prices.
Why It Matters
Treat AI shopping and travel suggestions as a starting point, not an answer — for now, they often miss.