About
The Voynichizer takes any text and a key and writes a second Voynich manuscript: a new book of 207 pages, named like the folios of the real one, with the sections, Currier languages, vocabulary and line structure of Beinecke MS 408, in EVA, the alphabet used to transliterate it, and in a Voynich-like script. With the same key, the hidden text comes back exactly.
Colophon. Made by Davide Caniatti in 2026, alongside a statistical study of the Voynich manuscript described in the preprint Hiding a Message in a Voynich-like Book: first the research, then a generator built on eighteen of the properties it measured. The program, its script and this site are free software under the MIT licence. Write to davide.caniatti [at] gmail.com.
Your privacy
- Everything runs in your browser. Your text, your key and your manuscripts are never sent anywhere.
- To run the program, the page loads Python and its libraries (Pyodide) from
cdn.jsdelivr.net. These are program files: nothing of yours is sent with the request. - No cookies, no analytics, no accounts. The site is hosted on GitHub Pages, which, like any web server, sees the address of whoever visits.
How close it is to the Voynich
Measured by hiding the same Latin text (the opening of Isidore’s Etymologiae, book XVII) with 24 different keys. The “judges” are two classifiers (logistic regression, 10-fold cross-validation grouped by page) that try to tell generated pages from real ones by looking at 209 and 228 page statistics: 0.5 means they are guessing, 1 that they never fail. The details of every measure are in How it works.
| measure (version v21, 24 keys) | value |
|---|---|
| judge 1 (glyphs, glyph pairs, words, line edges, vertical similarity) | 0.554 ± 0.004 |
| judge 2 (plus: word pairs seen elsewhere, identical neighbours, the first and last words of lines, the first lines of paragraphs, how far a page’s glyphs depart from the book and between its two halves) | 0.601 ± 0.005 |
| the same two judges on keys 13–24, never used to tune the program or to choose between versions | 0.554; 0.606 |
| scorecard of 18 properties of the text (17 attainable: the Voynich itself fails one when measured the same way) | 16.0 |
| 8 further properties | 5.4 |
| line test (five conditions on how lines start, end and follow each other) | 24 keys out of 24 |
| lines much wider than the others on their page (over 1.5 times the median; Voynich 2.9%) | 2.8% |
| lines somewhat wider (over 1.25 times the median; Voynich 6.0%) | 11.6% |
| the text comes back exactly; wrong key rejected | 24 of 24; 24 of 24 |
A manuscript with a message and one without cannot be told apart: a classifier trained directly to separate their pages scores AUC 0.506 (and 0.511 on two sets of books that both lack a message). This holds by construction for any test, as long as the cipher is sound: the encrypted message only takes the place of the filler, a stream of pseudorandom bits, that decides how many times each known word appears on each page.
What this does not mean
It is not “indistinguishable from the Voynich”:
- the second judge still recognises it a little (14 keys of 24 give a manuscript above 0.60);
- some known properties are not reproduced: the page profile is always missed; the spelling choices that vary from line to line (23 keys of 24), the spread of word lengths (20 of 24) and the first lines of paragraphs as a register of their own (18 of 24) usually are; and lines are too irregular in width (see the table);
- it writes only running paragraph text: no labels and no circular or radial text; it does not imitate the different hands, and it writes each page on its own, so what the two halves of a bifolium share in the manuscript is absent;
- the program was developed against these same instruments (its weights were tuned with seeds and keys of their own, and keys 1–12 were used to choose between versions); a judge built independently has not been tried;
- about 86% of the running words are Voynich words; the rest are invented and never occur in the manuscript. The judges do not look at the vocabulary: a test that checked whether each word occurs in the manuscript would spot the books at once, and anyone who knows the program can tell that a book was made with it. What they cannot tell is whether there is a message inside.
Nor does it decipher the Voynich manuscript, or say how it was made: the program imitates its statistics, it does not explain it.
Security
- The passphrase goes through scrypt (N = 215, r = 8, p = 1) to a master key. The compressed text is framed with its length and a 64-bit HMAC-SHA256 tag, and the frame is encrypted with a SHAKE-256 keystream: with the wrong key the program answers “wrong key”. The security is computational: it rests on SHAKE-256 behaving as a pseudorandom generator.
- These are standard building blocks, but the whole has not been reviewed by a cryptographer: do not use it for secrets that matter.
- The scrypt salt is fixed, so a list of guessed passphrases can be prepared once for all books: use a long random passphrase, six random words or twelve random characters.
- Every choice of the program derives from the key, so two books written with the same key share their layout and can be recognised as twins (there is no random value per book): use a different key for every book.
- The book must stay as it is: moving words within a page changes nothing, but adding, removing or changing a word breaks the reading, and there is no error correction.
Versions
- This site runs version v21 of the program (commit
4b1383c), unchanged, from its repository; the files and their SHA-256 are listed inengine/v21/manifest.json. Python and its libraries come from Pyodide 314.0.7. - v21 reads manuscripts written by v17 (the first published version) and v20: they differ only in the arrangement of the words, which carries no message. Future versions that change the word counts will keep the older versions available here for reading.
- The browser (Python compiled to WebAssembly) and the command-line program (CPython on Windows) use different mathematical libraries. In our tests the words on every page were the same and only their order within the lines could differ, and reading back worked in both directions. Other machines have not been tested, and the probabilities are computed in floating point.
- The browser lacks one function,
hashlib.scrypt: the site supplies it from noble-hashes 2.4.0, a standard implementation checked against the RFC 7914 test vector every time the page starts.
Sources and credits
- Text of the Voynich used by the program: from the ZL transliteration by René Zandbergen and Gabriel Landini, version 3b of 13 May 2025, published at voynich.nu, where the transliterations are stated to be in the public domain and are made available under the Creative Commons CC0 licence. Only the running paragraph text is included, with the readable words.
- The manuscript is kept at the Beinecke Rare Book and Manuscript Library, Yale University (MS 408). The drawings on these pages are details of the photographs that Yale University Library publishes under its open access policy for works in the public domain; only the background of the vellum has been evened out.
- The Voynich-like script was drawn for this project by a program, stroke by stroke; it imitates the shapes of the Voynich signs and is not a copy of any existing font.
- Built with Pyodide, NumPy, SciPy and scikit-learn; noble-hashes (MIT) and jsPDF (MIT); text set in EB Garamond (SIL Open Font Licence).
Licence and contact
The program, the script and this site are under the MIT licence; the Voynich text is in the public domain (CC0). Questions and remarks: the repository on GitHub.