New study: LLMs fail to grasp unspoken beliefs like humans do
LLMs can't read between the lines as well as we thought—new benchmark reveals gaps.
Researchers created the first expert-annotated dataset, crowdsourced for human judgments, to test how LLMs handle implicature—unspoken implied meanings—and their cancellation. They found LLM belief‑update understanding lags behind humans, especially in more naturally‑occurring scenarios. Control experiments suggest successes may stem partly from a reliance on prior beliefs, and failures depend on the type and form of belief update. Overall, current LLMs have not reached human‑level understanding of unspoken beliefs and belief updates.
- First expert-annotated implicature cancellation dataset created for LLM evaluation
- LLMs significantly underperform humans, especially in natural conversational contexts
- Successful belief updates by LLMs often rely on prior beliefs rather than true pragmatic understanding
Why It Matters
For LLMs to be trusted in high-stakes conversations, they need to handle implied meaning—this shows we're not there yet.