HoosierHelp benchmark shows LLM agents fail at social service navigation
LLM agents flop on 3,971 Indiana social services in new HoosierHelp benchmark
A team of researchers led by Yiyang Li and colleagues introduced HoosierHelp, a new interactive benchmark designed to evaluate how well LLM agents navigate social service resources. Unlike prior benchmarks, HoosierHelp grounds agents in a real dataset of 3,971 Indiana public social service listings and simulates users with realistic, varied needs and behavioral quirks—including impatience, rambling, unsupported requests, and self-contradictions. Agents must issue structured resource-search calls, handle non-ideal interactions, and choose final resources based on both needs and constraints.
In experiments across 240 samples and seven LLMs, current agents proved substantially unreliable. Performance dropped sharply on conversations requiring fallback strategies or dealing with contradictory user statements, with many agents failing to anchor to constraints during multi-turn dialogue. The authors argue these findings expose a critical gap: social service navigation requires agents to ground decisions in concrete, often changing user constraints, which today's models struggle to do consistently. This could leave vulnerable individuals without the right support.
- HoosierHelp is built on 3,971 real Indiana public social service resources, making it more grounded than synthetic benchmarks.
- Testing covered 7 LLMs across 240 samples, with simulated users exhibiting impatience, rambling, unsupported requests, and self-contradiction.
- LLM accuracy dropped sharply on fallback-required and self-contradictory conversations, showing current agents are unreliable for real-world resource navigation.
Why It Matters
As AI chatbots take on public service roles, unreliable agents could misdirect vulnerable people relying on social services.