Barenholtz's autogenerative theory explains what LLMs really do — and their limits
New paper fuses linguistics and LLM theory to reveal the statistical structure behind language generation.
A new paper by J. Mark Bishop and Stephen J. Cowley (arXiv:2607.07891) bridges the gap between Integrationist linguistics and Large Language Models by leveraging Elan Barenholtz's autogenerative theory. Roy Harris's Integrationism argues that language is not a code mapping onto a pre-given world but a situated activity oriented toward joint action. Yet it leaves key gaps: it lacks a structural mechanism for how signs sustain prospective openness, under-theorizes the continuity between linguistic and non-linguistic semiotics, and offers no detailed account of the accumulated 'archive' of past integrations.
Barenholtz's autogenerative theory, developed specifically from observing LLM behavior, fills these gaps without undermining Integrationism's core commitments. It provides: a structural mechanism for prospective openness; a computational correlate for semiotic continuity; and a theory of the archive—what the accumulated residue of past integrations looks like and how new participants draw upon it. For researchers in NLP and LLM design, this synthesis offers a principled account of the statistical structure that LLMs effectively exploit, and crucially, what that structure cannot provide (e.g., genuine understanding or situated intent). The work suggests that while LLMs master statistical patterns, they remain incapable of the integrative, context-aware acts that define human communication.
- The autogenerative theory provides a structural mechanism for the prospective openness central to Harris's bipartite communication framework.
- It explains the continuity between linguistic and non-linguistic semiotic activity, offering a computational correlate for semiotic integration.
- The theory defines the 'archive'—the accumulated statistical residue of past integrations—which LLMs exploit, but cannot use for situated joint action.
Why It Matters
For AI professionals: a foundational account of what LLMs statistically do—and why they can never truly understand or engage in context-aware communication.