Hiding a Message in a Voynich-like Book
The manuscript’s writing rules, measured against languages and medieval scribes, and reproduced by a steganographic generator
PDF, 38 pages, 347 KB, in English. CC BY 4.0.
The research record, with every preregistration, program, result and the laboratory notebook: github.com/AndreottiVIII/voynich-research. Earlier version: Version 1 (PDF, October 2026).
Abstract
We studied how the text of the Voynich manuscript (Beinecke MS 408) is written, in 716 logged experiments, 670 of them preregistered (hypothesis and criterion fixed before the run, code tested first on synthetic texts), and compared each finding with the same measurement in 24 natural languages, constructed languages, volunteers’ gibberish, published generators and ciphers and, for scribal habits, 79 medieval manuscripts.
The measurements quantify regularities described before (word boundaries marked redundantly by word form, a dependency between adjacent words that stops at the line break, line-edge forms, paragraph-initial gallows, copying from the line above) and add two that are, to our knowledge, new: the dependency between adjacent words mirrors the glyph sequences inside words, and immediate repetition is not avoided, unlike in every language and manuscript compared. Two properties described before for parts of the manuscript, the avoidance of starting a line like the line above and the drift of marked variants along the line, are measured across the whole text and exceed anything found in the medieval scribes.
From eighteen of these properties we built the voynichizer, which hides any text of up to about 20,000 characters in a Voynich-like book of 207 pages: the text is compressed, encrypted and authenticated with a key, and its bits decide only how many times each recurring word appears on each page. On 24 keys, two classifiers retrained on each book separate its pages from the real ones with AUC 0.55 and 0.60 (0.5 is chance), and the books pass 16 of the 17 attainable properties.
That a Voynich-like text can carry a real message had been shown at the level of words, most fully by the Naibbe cipher (Greshko, 2025). We show it again, with a modern construction and two stricter conditions: the book also follows most of the line, paragraph and page rules, and a book with a message cannot be told from one without (a classifier trained to separate them scores AUC 0.51), while the key recovers the text exactly. The regularities we measured therefore describe how the manuscript was written, but they cannot tell whether it contains a message: a text that follows most of them can carry one or not, and the two cases look the same.
The question
A century of proposed readings has not produced one that scholars accept, and the debate has narrowed to a question that statistics seems well placed to answer: is the text a meaningful message (a language, perhaps enciphered) or the meaningless product of a procedure? The paper asks whether regularities, however many and however strong, can decide it at all. Steganography gives the answer in principle: if texts with and without a message are samples of the same distribution, no test on the text can tell them apart (Cachin, 1998). The question then becomes constructive. For the statistics of words it had already been answered: hand ciphers turn real Latin into Voynich-like text that can be read back (Greshko, 2025). The paper asks it for the whole set of measured rules: can one build a book that follows the writing rules of the Voynich at the level of the line, the paragraph and the page as well as the word, and that carries a message or not, with the two cases indistinguishable by construction?
1. The generator: a Voynich-like book that carries a message
The voynichizer, which this site runs, turns any text of up to about 20,000 characters into a 207-page book in EVA transliteration and gives it back exactly with the key. The text is compressed, encrypted and authenticated; its bits drive an arithmetic decoder that samples how many times each recurring Voynich word occurs on each page. Everything else (layout, invented words, word order) follows the manuscript’s statistics and carries no information. The method of feeding encrypted bits to an arithmetic decoder to sample from a model is that of Ziegler, Deng and Rush (2019); to our knowledge no earlier tool combines a cover model of a whole historical book, distribution-preserving embedding and authenticated encryption, evaluated against the measured properties of that book. How it works describes it step by step.
Version 21 was run with 24 keys on the same Latin text (the opening of Isidore’s Etymologiae XVII, 38,304 bits). Keys 1–12 were also used during development, to choose between variants; keys 13–24 were never used for tuning, only to confirm earlier versions. The project had fixed in advance what an AUC means: 0.6 or less “hard to tell apart”, 0.7 or more “distinguishable”.
| measure | keys 13–24 (fresh) | all 24 keys |
|---|---|---|
| judge 1 (AUC; 209 page features) | 0.554 | 0.554 ± 0.004 |
| judge 2 (AUC; 228 page features) | 0.606 | 0.601 ± 0.005 |
| scorecard properties passed (of 17 attainable) | 15.9 | 16.0 |
| extra properties passed (of 8) | 5.2 | 5.4 |
| line gate passed | 12 of 12 | 24 of 24 |
| text read back exactly; wrong key rejected | 12; 12 | 24; 24 |
To read these numbers, the real manuscript was degraded on purpose and judged in the same way. With its lines shuffled within each page it scores 0.48 and 0.60; with its words shuffled within each line, 0.98 and 0.99. For the second judge the generated books are therefore at the level of the real text without its line order, but the book of one key in two is above 0.60 (14 of 24). As controls, the same judges tell Currier’s languages A and B apart with AUC 1.000, and give 0.53 and 0.52 with random labels. The generated books always miss the page profile (how far each page departs from its section) and usually the spelling choices that vary from line to line, the register of the first lines of paragraphs and the spread of word lengths; their lines are too irregular in width. These judges do not look at the vocabulary: one that checked whether each word occurs in the manuscript would separate the books at once, since invented words never do. For reference, the published generators pass from 1 to 10 of the 18 properties of the same scorecard.
What the generator does not do. It writes only running paragraph text: no labels, no circular or radial text, no drawings (this site places the real ones next to the generated pages). It does not model the hands, and it writes each page on its own, so the bifolium effects described below are absent. It copies from the line above by the position of a word in the line, not by its physical place; the closure of the line at a drawing and the drift along the line were not targeted. And it hides that there is a message, not that the program was used: about 86% of the running words of a book are Voynich words, the rest are invented, and a book can be recognised as the program’s output.
With and without a message. For each key a second book was written with filler bits only. Against the manuscript the two kinds score the same (paired differences −0.009 ± 0.008 and −0.007 ± 0.010), and a classifier trained directly to separate their pages (228 features, 9,420 pages) scores AUC 0.506, against 0.511 for a negative control. This holds by construction, whatever the judge, as long as the cipher is sound: the message enters only as encrypted bits, which no efficient test can tell from the random filler of a book without a message. The security is therefore computational, resting on SHAKE-256 as a pseudorandom generator. Two books written with the same key share their layout and can be recognised as twins (a random value per book would remove this, and is not implemented), so each message needs its own key. The construction is modern (arithmetic coding, current cryptography) and says nothing about how a fifteenth-century writer could have hidden a message.
2. The writing rules, measured
Each property is measured with its 95% interval (by bifolium) and set against the same measure in the comparison texts. The paper says, property by property, what was known before; many of these regularities were. What the project adds is to quantify them with the same code against the same comparison texts, and a few properties not found in the literature.
- Word form marks the word boundaries. The next glyph is very predictable (conditional entropy 2.23 bits, against 2.54–4.65 in the 24 languages); spaces can be predicted from the two glyphs on either side (F1 0.86), and cutting the text at other places allowed by the rules mostly yields words attested elsewhere (50%, median 6.5% in languages); the vocabulary fills the most probable glyph sequences (43%, median 5%), and frequent words are the most probable ones (rank correlation 0.58, median 0.09). Version 1 read this as a sign that the space is a weak boundary inside a chain of glyphs. That reading does not hold: a null model that keeps the word grammar and removes every link between words reproduces all four properties. Boundaries are marked redundantly by the shape of words, as in any script with positional forms. Described before by Bennett (1976), Lindemann and Bowern (2020), Stolfi (2000), Zattera (2023), Feaster (2020), Rozanova and Temerev (2026).
- Adjacent words are linked, and the link mirrors the rules inside words. The last glyph of a word
says a lot about the first glyph of the next: 0.188 bits beyond the shuffle within the line (0.163–0.211), against a
median of 0.097 in the languages; only Sanskrit and Tagalog, which write phonological links between words, exceed
it. Volunteers’ gibberish has 0.014, published generators 0–0.08 (except voynich-fingerprint, 0.16–0.17).
The preferences across the space resemble those
between consecutive glyphs inside a word (rank correlation 0.47, at most 0.37 in the 24 languages), and this part
does not follow from word form. Two junction rules: before a gallows a word begins with
qo-after-y,-o,-d(59–64%) and witho-after-n,-r,-s,-m(qo-only 23–27%); a word ends in-rbeforea-(85%) and in-lbeforek-,t-,d-,l-,s-,q-(63–86%). They hold at each position in the line with enough cases (for-l/-ronly from the fourth word on, and the other positions go the same way), in the three main hands, in every section with enough data and in three transliterations. They have the shape of rules like English a/an or French liaison, and say nothing about the sounds behind them. Where two words copied from the line above would break a rule, the piece that changes is the one that a writer who knows the next word while finishing the current one would change (only 37 and 60 such cases; not robust to the strictest correction). Beyond the next word there is no link at all, whereas 22 of 24 languages have one. The link was observed by Currier (1976), Smith (2017) and Smith and Ponzi (2019), and the first junction rule was described in part by Currier and Smith. The mirroring of the rules inside words is added here. - The line is a closed unit. Between the last word of a line and the first of the next the link is
zero (−0.003 bits), against 0.19 inside the line. In the comparison languages the ratio of the two has a median of
0.49 and only Hebrew falls below 0.10, but their lines are those of the source editions, often ending at a clause or
a verse, not lines of manuscripts, so they are a weak reference here. The link also drops where a plant interrupts a
line (from 0.167 to 0.026 bits). Lines have edge forms: for the same rest of the word, at line end
-mgains 15 points and-rloses 12 (darinside a line,damat its end); at line starts-andy-gain 9–10 points. In the labels of the drawings, which stand alone,qo-before a gallows is almost absent (0–5%, against 30–68% in lines), as predicted from the junction rule before looking (the fact itself had been noticed by Zandbergen, 2026). A herbal printed one verse per line closes its lines too: a line that is a unit of content would do the same. Described before by Currier (1976), Smith (2015–2016), Feaster (2020–2023), Rozanova and Temerev (2026); measured here for the same word and against scribes. - The scribe copies from the line immediately above. A word has an identical or nearly identical word in the two lines above more often than the vocabulary of the paragraph predicts (z = 17.4), more than twice as much as in the generator of Timm and Schinner (2020). The copying comes almost entirely from the line just above (in languages recurrence spreads over three or four lines), neighbouring words are copied from neighbouring words, usually in the same order, a copied word takes the form that agrees with its new neighbour (effect +0.49 against +0.16), and copying stops at the paragraph and the page. Observed by Timm (2014); the shape of the copying is added here.
- Each line avoids starting like the line above. Identical starts are 0.51 of the number expected
(0.40–0.61), strongest for
qo-(z = −9.7) ando-(−7.0): if the line above starts withqo-, the chance that the next does drops by 52 points. Only the left margin matters: not where lines resume to the right of a drawing, not the line two above, not the right margin. The median language has 0.97 and one of 23 (Portuguese) falls below the Voynich, but here too languages without physical lines are a weak reference; gibberish has 1.32, generators 0.93–2.05, and none of the 79 medieval manuscripts avoids the margin. A forum study had reported vertical avoidance of the same glyph in the herbal pages of one scribe (tavie, Voynich Ninja, 2024); measured here across the manuscript and against comparison texts. - Choices between variant forms follow a short state. Many words come in pairs that differ by one
glyph (
k/t,sh/ch,-ey/-dy); treating them as variant forms of one word is a working hypothesis, not a result. Nearby words repeat the same choice more than each word’s own preferences predict: +0.132 between adjacent words (0.111–0.153), +0.091 at two and three words, halving in about three words. The agreement is almost the same between very different words as between similar ones, it is not an effect of position, it is separate for each choice and the same in hands 1–3, in languages A and B and in three transliterations. Agreement of this size exists in languages as grammatical agreement (Italian -o/-a), so it does not set the Voynich apart from a language with agreement; in the manuscripts compared it is weaker or absent, but not established as unique (below). A slower state that crosses the line break is normal in handwriting. Local clustering of similar forms was noted by Timm (2014) and Schinner (2007), and a switch afterch/shthat holds for whole folios by Parisel (2026); the short, word-independent state was not found described before. - Marked variants drift along the line. For the same word,
qo,k,shand-eybecome rarer towards the right: −0.020 to −0.034 per ten glyphs, about 8–13 points over a line of median length. It is clearest forsh/ch(z = −6.3); the other three go the same way but each alone does not pass the strictest multiplicity correction. The medieval scribes drift at most 0.008 per ten glyphs, in no common direction. The rightward indices of Feaster (2021–2023) describe the same tendency; measured here for the same word and against scribes. - Immediate repetition is not avoided. One adjacent pair in a hundred is the same word repeated
(
chol chol), 3.9% nearly identical. Against the composition of each line the Voynich repeats the previous word as often as chance (ratio 1.04), whereas every comparison language and manuscript avoids it: languages 0.01–0.92, Nordic scribes 0.02–0.07, German manuscripts at most 0.79, the Copiale copyist 0. Fifteenth-century German herbals, recipe books and astrology texts avoid it too (0.05–0.51), so it does not come from genre. Version 1 read this as repetition beyond chance; the corrected reading is that it is not avoided, and that repeated words cluster in the same line. The rate was measured by Boxer (2023) and by Bowern and Lindemann (2021). That it is not avoided, unlike in every language and manuscript compared, is added here. - The book has a writing structure. 83% of paragraphs open with a gallows (
pin 54%), a mark of position more than of the word; first and last lines have their own register;-eyrises by 12 points (7–18) from top to bottom of the page; each page departs from its section 2.8 times more than chance. The bifolium is a unit: a page shares its identity with the conjugate leaf, which in the bound book can lie far away, and not with the facing page (z = 6.6); the two halves share rare words and copy from each other; all 48 bifolia with enough labelled pages are in one Currier language, 94% of 52 in one hand. Languages A and B differ in vocabulary, not in glyphs: B is not A enciphered with another key. Paragraph gallows: Currier (1976), D’Imperio (1978) and others; hands following the bifolia in the herbal section: Davis (2025). The shared rare words and the copying between conjugate leaves are added here.
3. Against languages, medieval scribes, ciphers and generators
If these rules were ordinary scribal habits, real scribes copying real texts would have them. The same properties were measured, with the same code, in 79 manuscripts transcribed at the level of letter forms with their original lines and pages: six Nordic manuscripts of about 1200–1550 (Menota) and 73 German manuscripts of 1350–1500, the area and century of the Voynich (Reference Corpus of Early New High German).
| property | Voynich | 79 medieval manuscripts | reading |
|---|---|---|---|
| junction rules (effect) | 0.058 | Nordic: 2 of 11 choices at 0.013–0.015; German: 2 of 13 at 0.001–0.005; the rest zero | about 4 times the strongest Nordic choice, 12 times the strongest German |
| avoiding the start of the line above | 5 starts avoided | none in 72 Nordic tests; 3 of 72 German with one start, as expected by chance | not found |
| line-edge forms (gain of the most typical form) | end -m +0.151; start
s- +0.098 | +0.012 to +0.025 (abbreviations, line-break hyphen) | same kind, 4–8 times stronger in the Voynich |
| immediate repetition (observed / expected in the line) | 1.04 | Nordic 0.02–0.07; German at most 0.79; Copiale 0 | avoided in every manuscript |
| drift along the line (per ten glyphs) | −0.020 to −0.034 | at most 0.008, in no common direction | absent with power in 20 of 29 choices |
| short state of choices (2–3 words) | +0.09 | absent with power in 14 of 29 choices; power insufficient in 13; present in 2 (one uncertain); the Copiale copyist rotates homophones (−0.18) | weaker or absent; not established as unique |
| slow state between consecutive lines | +0.050 | 6 of 17 choices, 5 as large as the Voynich | normal in handwriting |
These are Old Norse and Early New High German manuscripts; fifteenth-century Latin or Italian manuscripts transcribed at the level of letter forms, the closest to the Voynich, are not in the battery.
Ciphers and readings. Every negative result here comes from a test that recovers the answer when it is there. Real texts were encoded in many ways (word-for-word codes, 450 verbose ciphers, the Naibbe cipher, abbreviations, nulls, transpositions, encoded lists and prayers): some take on the letters, the vocabulary or the predictability of the Voynich, none its repetitions, page homogeneity, junction and line properties. Homophonic substitution with spaces in 14 languages recovers 69–98% of the words of the controls and at most 36% of Voynich words, mostly the same three; a solver without spaces, recovering 91–100% of its positive controls, never reliably beats the negative controls on the Voynich. The published readings that could be tested (Bax 2014, Vatne 2021, Cheshire 2019, Schechter 2026, Gatta 2026) do not pass their controls; applying a reading to a text generated without content is proposed as a minimum test for any announced decipherment. Historical mechanisms were simulated one at a time, each test first checked on texts enciphered on purpose. Several of them are later than the vellum, so these are tests of a kind of mechanism, not historical hypotheses: homophones chosen by the copyist would leave a negative agreement between nearby words (the Voynich has a positive one); no frequent word or sign behaves like the index letters of Alberti’s cipher disk; a message in the binary choices, as in Bacon’s biliteral cipher, restarting at each page or line, is absent (a continuous one is beyond the reach of the test); Trithemius’ word tables leave no rhythm; an autokey restarting at each line closes the line like the Voynich, but destroys repetition, inflates the vocabulary and leaves no margin avoidance.
Generators. No constructed language has the junction rules, the line properties or an agreement
between adjacent word forms like the Voynich’s. Volunteers’ gibberish has no junction, no line-edge forms, no copying
from the line above and no margin avoidance. Among the published generators measured, none combines the junction, the
closed line and the margin avoidance. voynich-fingerprint (Sachak, 2026), which reproduces 44 statistics on
held-out folios, has the junction and its mirroring of the rules inside words and, in two books of three, a drift of
qo-/o- close to the manuscript’s; but its lines are not closed, and it has no margin
avoidance, almost no line-final forms, no copying from the line above and no short state.
4. What it means
Two claims must be kept apart. A book with a message and a book without one are samples of the same distribution: no test on the text can separate them without the key. A generated book and the real manuscript are not: the judges separate them a little, several properties are not reproduced, and a stronger or independently built judge would probably do better. Together they mean that a text written under the rules measured here can carry a message, and that whether it does cannot be read from its statistics. This does not show that the manuscript carries a message, any more than generators without a message show that it does not. It shows, with a construction, what several studies had suspected from the statistics they measured (Gaskell and Bowern 2023; Bowern and Gaskell 2023; Parisel 2026; Rozanova and Temerev 2026): the debate between “meaningless” and “enciphered” will not be settled by the statistics of the text alone. What is specific to the Voynich here is not the indistinguishability, which any cover would give with a good cipher, but the capacity at a given distance from the manuscript: about 83,000 bits per book at AUC 0.55 and 0.60 with these judges, a lower bound that a better model could raise.
Not supported, with the controls described: a natural language in an ordinary alphabet read with a simple or homophonic substitution, in the languages tested; the published readings tested; the historical mechanisms simulated, one at a time; a text invented by hand without rules; and the idea that the rules are the ordinary habits of the 79 Nordic and German manuscripts measured (with reservations for the short state).
Open, four scenarios the data do not separate: a language written with a very artificial writing system (the junction rules are compatible with a phonetically written language); a code or cipher of an untested kind, for example with a word repertory or a lost key book; no message, from a procedure that can be executed by hand, though no generator measured reproduces the junction rules, the closed line and the margin avoidance together; and a line that is a unit of content (a verse, an entry, a recipe), compatible with any of the other three.
What would change the conclusions: a reading that passes these controls (it works on the real text and not on shuffled or generated text, produces meaningful sentences on pages not used to build it, and explains the junction rules, the closed line, the copying and the margin avoidance); a medieval manuscript or historical cipher with the same rules, in particular a fifteenth-century Latin or Italian one; a simple hand procedure that reproduces them all together. A reading will need an external anchor that regularities cannot give: a text of which the manuscript is a copy, a secure identification of many plants, a key found in an archive.
Limitations
- The preregistrations are internal, dated files in a public version-controlled history, without third-party timestamps; the programme was exploratory, and confirmations on data not used to form a hypothesis are few.
- Everything rests on transliterations that share the line segmentation; the images were checked only on a sample.
- Comparison manuscripts are Old Norse and Early New High German only. They are 79 manuscripts, not 79 known scribes: the files do not mark changes of hand, and the 73 German ones were measured together for the state of the choices, so one scribe’s state could be diluted; for the short state the power is insufficient in 13 of 29 choices. The lines of the 24 languages are those of their source editions, not of manuscripts. Only one real enciphered manuscript (the Copiale) was available.
- Three planned tests could not be run because their measures did not separate the cases on synthetic texts.
- The judges and the scorecard are the project’s own, the scorecard bands were set after seeing the results (a specification, not a test), and the generator was tuned against them. An independently built judge has not been tried: the books of the 24 keys, the keys and the code are public, and one built by others would be welcome. Keys 1–12 were used during development. The books are not indistinguishable from the manuscript. The encryption has not been reviewed by a cryptographer, and the scrypt salt is fixed.
- Several effects are small in absolute terms (the short state is about 13 points of agreement), and the look-ahead
and the drift of single choices other than
sh/chdo not survive the strictest multiplicity correction.
The research record
Everything behind the paper is public at github.com/AndreottiVIII/voynich-research: the preregistrations, the programs, the results with their provenance records, the laboratory notebook (in Italian), the log of all 716 experiments with their outcomes, and the sources of the paper with the script that draws its figures. Seeds are fixed, so re-running an experiment with the same software versions is meant to give the same numbers. Code is under the MIT licence; texts, results and figures under CC BY 4.0. Copyright-protected corpora are not included and must be downloaded from their sources. The notebook also records the provisional readings that later checks corrected.
The work was carried out in a continuous conversation with Claude, a large language model by Anthropic, which proposed experiments, wrote the preregistrations, the analysis programs and the notebook entries, ran the experiments and drafted the text; the voynichizer was developed in the same way. The author set the questions and the direction of the research, chose among the avenues proposed, approved every data download and change of scope, and reviewed results and text (Section 4.5 of the paper).
How to cite
Caniatti, D. (2026). Hiding a Message in a Voynich-like Book: the manuscript’s writing rules, measured against languages and medieval scribes, and reproduced by a steganographic generator. Preprint, version 2. CC BY 4.0.