Research & Papers

Rozanova and Temerev's Voynich study: glyphs aren't letters, spaces aren't spaces

⚡Voynich glyph entropy hits 2.7 bits—far below any known language's 3.5 bits.

Deep Dive

The Voynich manuscript, Beinecke MS 408, has long been analyzed under three unstated assumptions: its glyphs are letters, the strings between blanks are words, and every blank is a word space. In a new arXiv paper, Liudmila Rozanova and Alexander Temerev put all three to the test using the Zandbergen-Landini transliteration, comparing against matched prose, cipher, and pseudo-text controls with quire-level resampling. Their results are stark: none of the assumptions hold.

Glyph regularity in Voynichese is far too strong for any one-to-one substitution cipher—conditional entropy is 2.7 bits versus about 3.5 bits for Latin, Italian, and English. Tokens form a plausible vocabulary, yet one token predicts the next by under 1% of token entropy, far below every matched control's 2-10%. Meanwhile, glyphs at token edges share 0.2 bits of mutual information, exceeding all prose controls. Blanks also split into two regimes: uncertain separators are physically narrower (AUC 0.905 from image coordinates), behave like word-internal junctures, and are crossed by learned units even when spaces are erased. Published Voynich imitators reproduce the low entropy and weak token order but fail to replicate the edge-glyph coupling or the hapax-rich vocabulary (70% singleton types vs 41-60% in controls). The authors conclude that any account of the manuscript must earn—not assume—the step from glyphs, tokens, and separators to letters, words, and word spaces.

Key Points
  • Conditional entropy of Voynich glyphs is 2.7 bits, vs ~3.5 bits for Latin, Italian, and English prose
  • Token-to-token prediction is under 1% of token entropy, far below the 2-10% range of all prose controls
  • Edge glyphs share 0.2 bits of mutual information, and uncertain separators are physically narrower on the page (AUC 0.905)

Why It Matters

Voynich researchers must abandon letter/word assumptions; new metrics reveal the manuscript's true structure at token edges.

📬 Get the top 10 AI stories daily