Multi-LLM Council Maps Predictive Coding Neuroscience with Quantified Disagreement
Ten local LLMs score 31 studies, revealing structured disagreement across experimental paradigms.
A new paper from Hamed Nejat, Alexander Maier, Jesse Spencer-Smith, and André M. Bastos introduces an ontology-constrained multi-LLM pipeline for synthesizing fragmented scientific literature. The team manually defined a predictive-coding glossary of 36 concepts grouped into three hypotheses (predictive suppression, feedforward error propagation, and ubiquity). They then assembled a council of ten local language models—no cloud APIs required—to read 31 studies, extract evidence including figure descriptions, and score each study's agreement or disagreement with each glossary factor across local and global oddball contexts. This allowed pairwise study-agreement analysis, cross-model comparison, and three-dimensional hypothesis-space mapping. The system found high agreement for some hypotheses but weaker for others, revealing structured disagreement particularly between local and global oddball paradigms.
The authors also define "hypothesis-space temperature," a geometric dispersion metric measuring how compactly studies occupy the hypothesis space. Temperature was lower for local oddball contexts and higher for global oddball contexts, indicating greater spread in the latter. The scoring geometry additionally enabled estimation of change vectors between experimental contexts. This framework generalizes to any interdisciplinary field where conventional meta-analysis fails due to lack of a common comparison space. The results demonstrate that local multi-LLM councils can produce auditable, quantitative disagreement measurements, effectively turning messy literature into a structured evidence space that researchers can explore and trust.
- Ten local LLMs scored 31 studies against a manually curated glossary of 36 predictive-coding concepts across 3 hypotheses.
- The pipeline introduced 'hypothesis-space temperature' as a dispersion metric—global oddball contexts showed 2x higher dispersion than local ones.
- Pairwise agreement analysis revealed structured disagreement between local and global oddball paradigms, auditable through cross-model comparison.
Why It Matters
Automates literature synthesis in fragmented fields, enabling auditable, quantitative evidence mapping with local LLMs.