Persian proverb test exposes LLMs' 'decompression gap' in story generation
LLMs produce fluent stories but miss the moral point, researchers discover.
Researchers introduced the Proverb Aligned Narrative Dataset (PAND), pairing Persian proverbs with human-written stories and explicit meanings. Analyzing model behavior across multiple prompting regimes, they found a persistent "decompression gap": LLMs achieved strong surface-level fluency but failed to faithfully instantiate the underlying moral and causal structure. Explicit reasoning and iterative refinement partially closed the gap, suggesting the problem arises from difficulties in translating abstract meaning into narrative form rather than a lack of relevant knowledge.
- Researchers built PAND, a dataset of 1,000+ Persian proverbs paired with human-written stories and explicit moral meanings.
- Found a 'decompression gap': LLMs like GPT-4 and Claude achieve high surface-level fluency but fail to faithfully instantiate the proverb's causal and moral structure.
- Explicit reasoning and iterative refinement reduced errors by up to 30%, indicating the problem is translation of abstract meaning, not knowledge absence.
Why It Matters
This reveals LLMs' shallow understanding of cultural abstraction, impacting reliability in storytelling, education, and knowledge transfer.