Active Inference Model Reveals Computational Roots of Speech Dysfluency
A POMDP-based model simulates stuttering by reducing precision in word prediction
A new theoretical paper by Thomas Parr, Birtan Demirel, Youssuf Saleh, and Sanjay Manohar proposes a computational model of speech production and auditory segmentation grounded in active inference. The model uses Partially Observable Markov Decision Processes (POMDPs) to represent how the brain might generate and perceive speech in a dynamic social environment. Key innovations include a reciprocal mapping from words to sequences of varying numbers of phonemes, a shared state representation for both speaking and listening, and a social inference component that determines whether the speaker is self or other. This setup allows the model to simulate the cycles of discourse—taking turns, producing fluent sequences, and parsing incoming sounds.
The authors then probe how specific computational failures lead to speech dysfluency. They show that reducing the precision (confidence) assigned to the next word during generation introduces pauses and start-of-word repetitions—hallmarks of stuttering. By adjusting precision in different model parameters, they produce patterns that mirror both developmental stuttering and progressive loss of fluency in neurodegenerative diseases like primary progressive aphasia. The paper also discusses how the underlying message passing could be mapped to neural substrates, offering hypotheses testable with functional imaging. This work bridges computational neuroscience and clinical speech pathology, providing a principled framework for understanding and potentially treating disorders of fluency.
- Model uses Partially Observable Markov Decision Processes (POMDPs) to simulate speech production and auditory segmentation
- Reducing precision in the next word induces pauses and start-of-word repetitions analogous to stuttering
- Includes a social inference component to determine who is speaking (self or other)
Why It Matters
This model offers a computational testbed for diagnosing and treating stuttering and neurodegenerative speech disorders.