Rozanova and Temerev's Voynich study: glyphs aren't letters, spaces aren't spaces
Voynich glyph entropy hits 2.7 bits—far below any known language's 3.5 bits.
The Voynich manuscript, Beinecke MS 408, has long been analyzed under three unstated assumptions: its glyphs are letters, the strings between blanks are words, and every blank is a word space. In a new arXiv paper, Liudmila Rozanova and Alexander Temerev put all three to the test using the Zandbergen-Landini transliteration, comparing against matched prose, cipher, and pseudo-text controls with quire-level resampling. Their results are stark: none of the assumptions hold.
Glyph regularity in Voynichese is far too strong for any one-to-one substitution cipher—conditional entropy is 2.7 bits versus about 3.5 bits for Latin, Italian, and English. Tokens form a plausible vocabulary, yet one token predicts the next by under 1% of token entropy, far below every matched control's 2-10%. Meanwhile, glyphs at token edges share 0.2 bits of mutual information, exceeding all prose controls. Blanks also split into two regimes: uncertain separators are physically narrower (AUC 0.905 from image coordinates), behave like word-internal junctures, and are crossed by learned units even when spaces are erased. Published Voynich imitators reproduce the low entropy and weak token order but fail to replicate the edge-glyph coupling or the hapax-rich vocabulary (70% singleton types vs 41-60% in controls). The authors conclude that any account of the manuscript must earn—not assume—the step from glyphs, tokens, and separators to letters, words, and word spaces.
- Conditional entropy of Voynich glyphs is 2.7 bits, vs ~3.5 bits for Latin, Italian, and English prose
- Token-to-token prediction is under 1% of token entropy, far below the 2-10% range of all prose controls
- Edge glyphs share 0.2 bits of mutual information, and uncertain separators are physically narrower on the page (AUC 0.905)
Why It Matters
Voynich researchers must abandon letter/word assumptions; new metrics reveal the manuscript's true structure at token edges.