Voynichizer
2r

How it works

The Voynichizer writes a second Voynich manuscript: a new book of 207 pages with the same folio names, sections and Currier languages as the real one, built from its vocabulary and imitating most of the regularities that statistical study has found in it, from the shape of the words to the structure of lines and pages (those it does not reproduce are listed under Known limits). Your text, compressed and encrypted with the key, decides only how many times each known word appears on each page; everything else comes from the key and from the statistics of the real manuscript, and carries no information.

This page describes version v21 exactly as it runs on this site and in the repository (commit 4b1383c). Names in monospace are the functions and files of the program, whose comments are in Italian. All numbers are measured, and where they come from is said.

A second Voynich: what it has in common with the real one

The aim is a book that could stand beside MS 408 as another copy of the same kind of text: not a translation, not a collage of real pages, but new pages written by the same rules. Each property below is a regularity of the real manuscript; the second column says how the program reproduces it, the third how close it comes (version v21; the measures are described below).

property of the Voynichhow the program reproduces itmeasured
207 pages with running text (folio sides such as f1r, f1v), in seven sections (herbal, recipes, biological, pharmaceutical, …) and two “languages”, Currier A and Bthe same 207 folio names, same section and language for each pageby construction
the alphabet: EVA glyphs, with the compound glyphs ch sh cth ckh cph cfhsame glyphs; the words are Voynich words or built from its glyphsglyph frequencies: judge group G1 at 0.49 (chance)
vocabulary: each section and language has its own words and frequencies; Zipf-like ranksknown words drawn from the lexicon of the same section and language, never from the page itselfZipf slope and distinct words per word within tolerance
each page has a “character”, a preference for some glyphsthe character of another real page of the same kind is borrowed, at strength κpage distance from the book (G9) at 0.47; but the page profile, how much each page departs from its section, is missed by every key (0.66 against 1.02)
13.7% of the running words occur only once, more often at line starts and in first linesinvented words, shaped like the one-off words, in the places where those occurone-off share within tolerance; words unique in the text slightly too many (0.154 against 0.145)
word shape: glyphs keep to fixed positions (q first, -dy, -aiin last…) known words are real; invented ones follow a model of the one-off words that predicts each glyph from the three before itshape of the words within tolerance; glyph pairs (G2) at 0.54
two neighbouring words, written together, often form another word of the book (about twice as often as by chance)some invented words are joins of two known words of the page, checked glyph by glyphattested joins within tolerance (2.13 against 1.96)
lines are units: special first and last words, links across word boundaries, rare repeats of a line’s first glyph from the line aboveplace model learned from the Voynich, edge links, Metropolis arrangement line edges (G4) 0.56, line words (G7) 0.56; line test passed by all 24 keys
first lines of paragraphs are different (more gallows p f t k, longer words)place classes distinguish first linesG8 0.61 over 24 keys: with the word statistics (G3, also 0.61), the group the second judge sees best; “first lines as a register” missed in 18 keys of 24
words two places apart resemble each other; the line above is echoedsimilarity terms in the arrangementA 1.025 (Voynich 1.047); vertical within tolerance
line widths vary in a characteristic wayexpected width per line (since v18; the measured curve since v19)lines over 1.5 × median 2.8% (Voynich 2.9%); over 1.25 × median 11.6% (Voynich 6.0%)
spelling choices (k/t, sh/ch, -ey/-dy…) vary from line to linepartly (agreement terms in the arrangement)not reproduced: missed in 23 keys of 24

Taken together: over 24 keys the judges score 0.554 and 0.601, against the 0.5 of pure chance. On the 12 keys used while the program was being developed both stay at or below the threshold of 0.6 that the research set in advance for “hard to tell apart” (0.554 and 0.596); on keys 13–24, never used to tune the program or to choose between versions, the second judge is just above it (0.606), and over all 24 keys it is above 0.6 with 14 of them. For comparison, the real manuscript with its lines shuffled within each page scores 0.60 with the second judge. The judges were built by us, they do not look at the vocabulary, and a few properties are still missed: the second Voynich is close, not indistinguishable, and this page says exactly where it differs.

Overview of the program

key scrypt master key k text zlib, length, tag ⊕ SHAKE-256(k) bits + filler page layout,word places weights of theknown words arithmetic decoder:bits → word counts inventedwords arrangement→ EVA file random choices: HMAC-SHA256(k, use, page) repeated for each of the 207 pages, in the order of the manuscript
  1. The key goes through scrypt and becomes a 256-bit master key k.
  2. The text is compressed (zlib), framed with its length and an authentication tag, and encrypted with a SHAKE-256 keystream derived from k. More keystream follows as filler, so the bit stream that enters the book looks random from start to end.
  3. For each of the 207 pages of the real manuscript that carry running text, in order, the program draws from k a layout (lines, words per line, paragraphs) and decides which places hold invented words. The other places hold known words: words that occur at least twice in the Voynich.
  4. Each page has a probability distribution over the known words of its section and language, taken from the other pages only. An arithmetic decoder, fed with the encrypted bits, draws from it how many times each known word appears on the page. This is the only step that carries the message.
  5. Invented words are made with the shape of the Voynich’s one-off words; then all the words of the page are placed in the lines by a sampler that imitates the Voynich’s line structure.
  6. Reading back recomputes the distributions from the key, counts the known words of each page, runs the arithmetic encoder on those counts to get the bits back, decrypts, checks the tag and decompresses.

What depends on what

part of the bookdrawn fromcarries the message
layout of each page (lines, words per line, paragraphs)key, statistics of the sectionno
which real page lends its “character” to each pagekeyno
which places hold invented wordskey, statistics of the one-off wordsno
how many times each known word appears on each pagekey, section lexicon, encrypted textyes
invented wordskey, model of the one-off words, the page’s known wordsno
order of the words in the lineskey, model of the linesno

The design principle

A message hidden in a text leaves a trace if it changes any statistic of the text. The program therefore never chooses anything because of the message: every page is a sample from a fixed, key-dependent model of the Voynich, and the encrypted message is used only as the source of randomness for one part of that sample. An arithmetic decoder fed with uniformly random bits produces an exact sample of its model, up to the 16-bit quantisation of the probabilities. No efficient test can tell the output of a good cipher from random bits. So a book that carries a message and a book that carries only filler are, to any such test, samples from the same distribution. This is what “cannot be told apart” means here, and the security is computational: it rests on SHAKE-256 behaving as a pseudorandom generator. It holds by construction, whatever the judge, as long as the cipher is sound, and the measurement agrees: a classifier trained to separate books with and without a message scores AUC 0.506 (below). The method is that of Ziegler, Deng and Rush (2019), the measure of security that of Cachin (1998).

Two consequences follow. The model, not the message, decides how close the book comes to the manuscript. And the capacity is fixed by the entropy of the part of the model that carries the bits: it cannot be raised without changing the statistics.

The earliest versions hid the bits in spelling choices (ch/sh, k/t, -dy/-ey and others), first directly (v0), then through the same arithmetic coding (v1). Spelling choices carry little information per place and are among the most studied features of the Voynich; from v10 on the message lives in the counts of the words of each page, which carry about 2.7 bits per known word: from about 46,000 bits per book with the spelling choices, the capacity nearly doubled.

The source: the Voynich text

The program carries its own copy of the Voynich text (voynich_zl3b.json): the ZL transliteration by René Zandbergen and Gabriel Landini, version 3b of 13 May 2025 (public domain, CC0), reduced to the running paragraph text, with the page, the paragraph starts, the section and the Currier language of each line. Words containing unreadable characters (?, *) are left out. Pages without running paragraph text (most of the zodiac and some diagrams) are not part of the model; labels and circular text are not used.

whatcount
pages with running text (the pages of every generated book)207
lines; of which open a paragraph4,130; 740
running words; distinct words34,863; 7,022
distinct words occurring at least twice (“known words”)2,246, covering 86.3% of the running words
words occurring once (“one-off words”, hapax legomena)4,776, 13.7% of the running words
lines per page; words per linemedian 14 (4–54); median 9 (mean 8.4)
section (ZL code)Currier ACurrier Bunassigned
herbal (H)95321
recipes, “stars” (S)223–
biological, “bathing” (B)–19–
pharmaceutical (P)16––
text only (T)15–
cosmological (C)–44
astronomical (A)––5

Glyphs are counted as EVA letters, with ch, sh, cth, ckh, cph, cfh as single glyphs (longest match first).

The key and the random choices

  • Master key: k = scrypt(key in UTF-8, salt = "voynichizzatore/canale-sacco/1", N = 215, r = 8, p = 1, length 32 bytes). scrypt costs 128·r·N = 32 MiB of memory per evaluation, which makes guessing keys slow and expensive.
  • Generators: every random choice has its own generator, one per page and per use: random.Random(int(first 16 bytes of HMAC-SHA256(k, "sacco|use|page"))), Python’s Mersenne Twister.
usedecidesneeded to read back
caratterewhich real page lends its character to the pageyes
impaginazionethe layout of the pageno
postiwhich places hold invented wordsno
ordinethe initial order of the known wordsno
nuovethe invented wordsno
disposizione (one for the book)the arrangement of all pagesno

Because each page and use has its own generator, reading back recomputes only what it needs (the distributions) and does not rebuild the book.

The message: compression and encryption

bit_cifrati, testo_dai_bit.

fieldsizecontent
length4 byteslength of the compressed text, big-endian
tag8 bytesfirst 8 bytes of HMAC-SHA256(k, "etichetta" ‖ compressed text)
datavariablethe text in UTF-8, compressed with zlib at level 9
  • The frame is XORed with the start of the keystream SHAKE-256(k ‖ "flusso"); the next 30,000 bytes of the same keystream follow as filler. Bits are taken most significant first.
  • A book with no message uses filler only. Filler and ciphertext are both keystream-like, so the counts of a book with a message and of one without come from the same distribution: without the key, and as long as the cipher is sound, nobody can tell where a message ends, or whether there is one (measured: AUC 0.506, see below).

The layout of each page

statistiche_gabbia, gabbia_statistica. The generated page keeps the name, section and language of the real page, but not its layout. From the pages of the same section and language (or of the same section, or of the whole book, when fewer than five pages are available) the program draws, with the generator impaginazione:

  • the number of lines of the page, and the page “width”, i.e. the median number of words per line of a real page;
  • paragraph lengths in lines, until the page is full (a leftover single line closes the previous paragraph, because the Voynich has no one-line paragraphs);
  • for each line, its number of words: the width times a ratio drawn from the real ratios of lines with the same role (first, middle or last line of a paragraph, or single line).

No page of a generated book has the layout of a real page. Books have about the same number of lines as the Voynich (4,048–4,179 in the books counted; Voynich 4,130) and about 4% more words (36,346 on average over 24 keys; Voynich 34,863).

Which places get invented words

Each place in a line belongs to one of eight classes: line opening a paragraph or not, times first, second, middle or last word of the line. In the Voynich the share of one-off words depends strongly on the class:

place in the lineother linesfirst line of a paragraph
first word0.1750.477
second word0.0900.200
middle0.0870.189
last word0.2120.308

With the generator posti, each place holds an invented word with probability min(0.95, share of its class × m), where m is the one-off rate of the page that lends its character (next section) divided by the book average. The remaining n places of the page take known words.

The known words of each page

Sacco.carattere, Sacco.pesi_carattere, canale_sacco.distribuzione.

  • Lexicon. The lexicon of page p is the list of all occurrences of known words on the other pages of the same section and language (falling back to the section, then to the book, if it has fewer than 500 occurrences). Its size is 283–2,246 distinct words, median 1,170. A page never draws from itself.
  • Character. Real pages differ in the glyphs they prefer. With the generator carattere the program picks another real page q of the same section and language (with at least 40 known words) and measures, for each glyph g, how much more or less q uses it than its section: δg = log[ (cq(g) + α f(g)) / (nq + α) / f(g) ] with α = 50, cq(g) the count of g in the known words of q, nq their total and f(g) the frequency of g in the section lexicon.
  • Weights. Each word w of the lexicon gets the weight c(w)γ · exp( κ · Σg ng(w) δg ) where c(w) is its count in the lexicon, ng(w) the number of glyphs g in it, γ = 0.9758 and κ = 0.6435 (fitted on the measurements; the parameters are in pezzi_parametri_v21.json).
  • Weights are turned into integers on a scale of 240 (at least 1) and sorted from the heaviest, ties by the word, so that writing and reading see exactly the same list on any computer.

For example, with one key the page f1r gets a lexicon of 598 words, an entropy of 8.25 bits per word and ol, shedy, chedy, shey, or as its most likely words; f33v gets 844 words with aiin, daiin, or, ar, okaiin at the top; f103r, a recipe page, 1,481 words with shedy, chedy, qokeey.

Hiding the bits in the word counts

conteggi_dai_bit, bit_dai_conteggi, _passi; the coder is in v1.py.

Given n places and integer weights w1 ≥ … ≥ wm (total W), the counts are drawn as a chain of binomials, which is exactly one multinomial draw:

c1 ~ Bin(n, w1/W),   c2 ~ Bin(n − c1, w2/(W − w1)), … ;  the last word takes what remains.

Each binomial is drawn as a sequence of yes/no decisions: having already placed j copies of the word, “stop here?” has the probability

hj = P(X = j) / P(X ≥ j),   X ~ Bin(remaining places, w/remaining weight)

computed in double precision and quantised to 16 bits (between 1 and 65,535 out of 65,536). Every decision is a binary symbol produced by an arithmetic decoder (the coder of Witten, Neal and Cleary, with 32-bit registers and underflow handling) whose input is the encrypted bit stream:

  • an arithmetic decoder turns uniformly random bits into symbols with exactly the requested probabilities, up to the 16-bit quantisation; since no efficient test can tell the encrypted bits from random ones, the counts are an honest sample of the page’s distribution, whatever the text;
  • the bits are consumed in reading order, page by page. A short message sits in the first few pages (1,160 bits in 4 pages), a long one runs through the book (76,888 bits in 200 pages); the Latin test text of the measurements, 38,304 bits, fills 124–131 of the 207 pages, depending on the key. Filler does the rest;
  • the capacity of the book is the number of bits the decoder has consumed when the last page is done, minus the 32 bits of its register: if the message needs more, the program stops with “text too long for this book”.

The known words that result are shuffled (generator ordine) and go to the arrangement together with the invented words.

The invented words

parole_nuove.FormeUniche. Invented words imitate the 4,776 words that occur only once in the Voynich:

  • for each of the eight place classes, a length L (in glyphs) is drawn from the lengths of the one-off words in that class;
  • a glyph sequence is drawn from a model of the one-off words that predicts each glyph from the three before it (1,044 such three-glyph contexts seen at least three times), backing off to the two before it (422 contexts), with the counts of the class plus a tenth of the counts of all classes. A candidate must have at least two glyphs, have length L (up to 400 tries, then the closest), and be neither a Voynich word nor a word already used in the book;
  • with probability 0.25, an invented word of six glyphs or more is instead the join of two known words of the same page whose lengths add up to L, accepted only if every glyph triple in it occurs in the one-off words and the two triples across the join have a conditional probability of at least 0.05. In the Voynich a third of the “attested joins” come from one-off words made of two words written together;
  • six candidates are drawn and one is chosen with weight exp( 0.37 · Σg profile(g) ) where the profile is the log-ratio of each common glyph (at least 1% of the book) on this page’s known words against the whole book, smoothed with α = 50: invented words follow the character of their page.

The arrangement of the words

disposizione.py, modello.py. The arrangement carries no information: reading back ignores it. Its purpose is to make the lines look like the Voynich’s, where the place of a word in the line, its neighbours and the line above are far from random. Two models are learned from the Voynich when the program starts:

  • Place affinity: a multinomial logistic regression (scikit-learn, C = 1) from 578 binary features of a word (first glyph, first two, last, last two, length up to 9, prefix and ending from a segmentation into the 150 most frequent pieces, presence of p and f) to the eight place classes. It is learned on features, not on the identity of the words, so it applies to invented words too.
  • Links between neighbours: observed/expected ratios, clipped to [0.2, 5], for the pairs (ending, prefix), (prefix, prefix) and (last glyph, first glyph) of consecutive words; whether the two words written together form a word of the book (log 0.0916/0.0485 if yes, log 0.9084/0.9515 if not: 9.16% of neighbouring pairs in the Voynich against 4.85% after shuffling); whether the pair occurs on other pages.

The score of an arrangement is a weighted sum (weights of v21, the same as v20):

termweight
place affinity of every word1
(ending, prefix) and (prefix, prefix) links between neighbours1.003
neighbours that form an attested word when joined0.979
(last glyph, first glyph) link between neighbours0.742
neighbouring pair seen on other pages−0.366
identical neighbours−0.417
similarity with the word in the same position in the line above (1 − normalised Levenshtein)0.152
first glyph of a line equal to the first glyph of the line above (the Voynich avoids it)−0.948
similarity of words two places apart in a line0.300
spelling choices agreeing within a line; with the line above (the counts start from the lines of the page; in v12–v20 they started from zero by mistake, which made the choices vary too much from line to line: 1.200 against 1.103 in the Voynich; v21 gives 1.127)0.044; 0.048
difference in glyph content between the two halves of the page (Σ of squared differences / glyphs)1.103
line widths in characters: squared deviation from the expected width (since v18)−0.01

The expected width of a line (since v19) comes from a curve measured on the Voynich: eleven points relating the logarithm of the line’s words over the page median to the logarithm of its width in characters over the page median, linear between the points; the expected widths are scaled so that they add up to the characters of the page, and a word is as wide as its EVA letters plus one space.

Starting from a random order, the sampler proposes 60 × N swaps of two random words of the page (N = words on the page) and accepts each with the Metropolis rule at temperature 1: always if the score does not decrease, otherwise with probability eΔ. This step takes about four fifths of the writing time.

The manuscript file

UTF-8 text; lines starting with # are comments; then one line per line of the manuscript, with the folio and line number, the words separated by dots, and @ in front of the lines that open a paragraph:

@<f1r.1> shocphol.shol.daiin.shdal.daiin.ckhed.ofy.qopchedy.cfhol.shey.dam
<f1r.2> yfars.orom.cheey.dal.olsheol.sheodaiin.dair.oraiin.otaldain.chedy.kchody.oram

A book is about 260 KB. Folio names are those of the real manuscript (f1r … f116r, including sub-pages such as f67r1 and the rosette foldout fRos).

Reading back, and what breaks it

  1. Derive k from the key; for each page, recompute the distribution of its known words (lexicon, borrowed character, weights): this needs only the generator carattere.
  2. Count how many times each word of that distribution occurs on the page; every other word and the order of the words are ignored.
  3. Run the arithmetic encoder on the same yes/no decisions: it gives back the bits.
  4. XOR with the keystream, read the length, check the HMAC tag, decompress.
change to a manuscripteffect on reading
moving words within a page, or between lines of the same pagenone
changing an invented word into another word that is not in the page’s lexiconnone
adding, removing or changing a known word, or moving one to another pagethe decisions change from that point: “wrong key, or altered manuscript”
wrong key“wrong key, or manuscript without a message” (length check) or “altered manuscript” (tag)
the PDFcannot be read: there is no program from the script back to EVA

The tag is 64 bits: a wrong key passes it with probability 2−64.

Capacity, time, memory

measurevalue
capacity of a book (bits of the encrypted frame)81,100–85,900 over the 24 measurement keys (mean 83,400); 80,344–84,678 in the five test books of this site. It depends on the key, and slightly on the message, because the decoder consumes a number of bits that depends on the bits themselves
the same, per known wordabout 2.7 bits
frame overhead12 bytes
examples174 characters → 1,160 bits; 1,027 characters of Italian → 4,936 bits; 32,014 characters of very repetitive English → 76,888 bits
writing, desktop computerabout 105 s: model 9 s, word counts 7 s, arrangement 87 s, check 2 s
writing, this site (browser)about 3½–4½ minutes
reading backabout 5 s on a computer, 12 s in the browser
memory in the browserabout 400 MB to write, 480 MB with the PDF
download on first useabout 27 MB (Python and its libraries), 10 MB more for the PDF

How many characters fit depends on how well the text compresses: ordinary prose fits about 20,000 characters, a repetitive text more. The site measures the text before writing: a text that cannot fit is refused at once; a text near the limit is checked as soon as the capacity of the book is known, about a tenth of the way through.

Determinism, browser and computer

  • On one computer the program is deterministic: the same text and key give the same file, byte for byte, also across Python hash seeds (tested with seeds 0 and 12345).
  • This site runs the program’s files, unchanged, in the browser with Pyodide 314.0.7 (Python 3.14.2, NumPy 2.4.6, SciPy 1.18.0, scikit-learn 1.8.0); our reference computer used Python 3.12.10, NumPy 2.5.3, SciPy 1.18.1, scikit-learn 1.9.1. In our tests the words of every page came out identical: the floating-point quantities behind the counts (word weights, binomial probabilities) are rounded to integers before use, weights on a scale of 240 and probabilities on 16 bits, so a difference in the last digit between two machines could change them only if it fell exactly on a rounding boundary, which never happened in our tests. Their order in the lines may differ, because the arrangement compares floating-point scores of a fitted model at every step. In our tests a manuscript written here read back with the command-line program, and vice versa. This was tested in both directions between CPython on Windows and Pyodide in the browser, which use different mathematical libraries; other machines have not been tested.
  • The browser’s Python has no hashlib.scrypt; the site supplies it from noble-hashes 2.4.0, checked against the RFC 7914 test vector every time the page starts. Progress is read by wrapping, from outside, the two functions the program calls once per page; what they compute is not changed. The files are listed with their SHA-256 in engine/v21/manifest.json.
  • v21 differs from v17 and v20 only in the arrangement, so it reads manuscripts written by both; the word counts of every page are identical in the three versions.

The acceptance test of this site

Five cases, each written both by the command-line program on a computer and by this site in the browser, then read back crosswise (4 October 2026, version v20; v21 changes only the arrangement, see below):

casetextbrowser → computercomputer → browserv17 book → v20wrong keysame words on every page
c1174 characters of Englishexactexactexactrejected207 of 207
c21,027 characters of Italian with accents, quotes, blank lines, double spaces; key with accentsexactexactexactrejected207 of 207
c3Greek, Russian, Chinese, Japanese, Korean, Arabic, Hebrew, emoji, combining marks, tabs, a Windows line end; key with emojiexactexactexactrejected207 of 207
c432,014 characters, 76,888 bits: 91% of that book’s capacity; a 264-character keyexactexactexactrejected207 of 207
c5a book with no message“no message”, as it should“no message”“no message”rejected207 of 207

The capacity of each book was the same on both sides (80,344–84,678 bits), and the case c1 written twice in the browser gave the same file byte for byte. Writing took 216–350 s in the browser, on a computer busy with other work.

The same checks were repeated for v21, the version the site runs now: the browser reads back exactly the books written on the computer by v17 and v20 (cases c1, c3 and c4); a book it writes reads back with the command-line v21; and its words are the same as those of the v20 book on all 207 pages (writing took 157 s).

The script and the illustrated book

carattere.py. The Voynich-like script is a TrueType font (version 1.1) drawn by a program, stroke by stroke; it imitates the shapes of the Voynich glyphs and is not a copy of any existing font. Simple glyphs sit on the EVA letters; the compound glyphs ch sh cth ckh cph cfh are separate glyphs at private code points from U+E000, and also ligatures of their letters, so that plain EVA text shows correctly. Every glyph is always the same: it does not imitate the variations of a pen.

The illustrated book (PDF, and the preview on the Write page) is made by this site, not by the program: it is a way of looking at the manuscript, and it cannot be read back. The program writes only running text; the drawings are real, taken from the Voynich, and are placed by the site. Each page of the generated manuscript is composed on a sheet of 6 × 8.6 inches:

  • The vellum is the grain of the blank flyleaf of MS 408, made seamless and isotropic (its spectrum averaged over directions, with random phases), in the colour of the written leaves, darker towards the edges.
  • The drawing is the one of the real folio with the same name, lifted from the Yale photograph: the painted areas are found by their pigments (green, blue, saturated red), the box around them is cut out, the original writing inside it is erased (bands of small dark strokes, and small marks away from the paint) and the vellum is evened out to white, so that the drawing can be multiplied onto the new page. 153 pages have a drawing of their own; the others take one from their section (the sun of f67r for the astronomical pages, a rosette for the cosmological ones); recipe pages get, as in the manuscript, a small star beside each paragraph (eight stars taken from f105r–f112r); text-only pages get none. The labels and the circular text of the real drawings are not imitated: the new text runs in straight lines around the drawing.
  • The layout tries three arrangements (drawing on one side, across the top, in the middle with text above and below), each with the largest type size that fits all the lines of the folio and with the drawing shrunk if needed, and keeps the one that best combines readable text with a large drawing. Lines are never broken: a line too long for its place is set a little smaller.
  • The writing is set in the Voynich-like script in brown ink, each line with a slight slope and each word with a slightly different opacity and baseline; the folio number is written at the top right, as in the manuscript.
  • The PDF is made in the browser with jsPDF: the text is real text in the embedded font, the vellum is a single image shared by all pages, and each drawing is multiplied onto its piece of vellum before being placed. A book of 207 pages weighs about 8 MB and takes about half a minute.

How close it is: the measurements

Generated books are compared with the Voynich by two classifiers (“judges”), a scorecard of 18 properties, 8 further properties and a test of the line structure. They were built during the research and are not part of the program; their code is in the public research repository, and the paper reports them in full. The measurements hide the same Latin text (the opening of Isidore’s Etymologiae, book XVII, about 1,668 words, 38,304 bits once compressed and encrypted) with 24 keys: keys 1–12 were also used while choosing between versions of the program; keys 13–24 were never used to tune it or to choose between versions, only to confirm them (from v14 on), so the second set is the fairer estimate. The weights of the program were never tuned on these keys: they were tuned with seeds and keys of their own. Results are means over the keys, with their standard errors.

The two judges

  • Unit: a page with at least 40 legible words (202 pages). Each real page is paired with the generated page of the same name.
  • Classifier: standardisation followed by a logistic regression (L2, C = 1), fitted inside each fold.
  • Validation: 10-fold cross-validation stratified and grouped by page name, so that a real page and its twin always fall in the same fold; repeated with 5 seeds. The out-of-fold probabilities of each seed give one ROC AUC; the AUC of a key is the mean over the seeds. 0.5 means the judge is guessing, 1 that it never fails; the research set in advance ≤ 0.6 as “hard to tell apart” and ≥ 0.7 as “distinguishable”.
  • Controls: the same judge separates Currier A from B pages with AUC 1.000, and gives 0.52–0.53 with random labels (one permutation each; with about 400 pages, the standard error of an AUC near 0.5 is about 0.03). Two earlier generators of the project score 0.97 / 0.98 and 0.87 / 0.94.
  • What they do not look at: no feature tests whether a word belongs to the Voynich vocabulary. A judge that did would separate the books at once, since invented words never occur in the manuscript.
groupwhat it measures on each pagefeaturesAUC of the group alone, v21, keys 1–12
G1frequencies of the 30 most frequent glyphs300.49
G2frequencies of the 150 most frequent pairs of consecutive glyphs inside words1500.54
G3word length (mean, deviation), distinct words per word, words unique in the page, words unique in the text, share of the top 100 words60.60
G4line edges: first glyph and last glyph of the lines, similarity of neighbours, neighbours that form an attested word220.56
G5vertical similarity with the line above10.54
G6neighbouring pairs seen elsewhere, identical neighbours, similarity at distance two30.53
G7length of the first, second and last word of a line; first words in y/d/s; last words in m/n/l/r/y90.56
G8first lines of paragraphs: excess of p/f/t/k glyphs, word length50.62
G9distance (Jensen–Shannon) of the page’s glyphs from the book’s, and between the page’s two halves20.47

Judge 1 (e231) uses G1–G5, 209 features; judge 2 (e266) uses G1–G9, 228 features. Over all 24 keys the groups the second judge uses best are G3 and G8, at 0.61 each. The features the second judge leans on most (first key) are pairs of neighbouring words seen elsewhere (generated 0.236, Voynich 0.224), neighbours forming an attested word (0.084 against 0.092), a few glyph pairs (l+ch, d+y, r+ch), lines starting with t, and words unique in the text (0.154 against 0.145).

The scorecard

Eighteen properties of the text, each with a tolerance around the Voynich’s value, measured on the whole book cut into pseudo-pages of 29 lines (a convention fixed early; real and generated books are cut the same way). The bands were chosen after seeing the results of the experiments, so they are a specification, not a test. Measured this way the Voynich itself fails one of them (the “gradient”, 0.683 against a band of 0.70–0.97, set on a different page division), so 17 are attainable.

propertypasses if
uncertainty of the next glyph (h2)within ±0.15 of the Voynich
predictability of a space from the previous glyphwithin ±0.08
share of one-off words in 34,000 wordswithin ±0.08
distinct words per wordwithin ±0.04
word identical to the previous one (relative to the line)within ±0.2
similarity of the words of a line (“homogeneity”)within ±0.008
similarity six lines apart over similarity within a line (“gradient”)between 0.70 and 0.97
information between the last glyph of a word and the first of the nextwithin ±0.05
neighbours that form an attested word, over chancebetween 1.5 and 2.5
flat curve of new words (one-off share at 1,000 words minus at 34,000)at most 0.10
drift of the vocabulary between near and distant blocksat least 0.05
page profile: how much each page departs from its section, glyph position by glyph position in the wordR between 0.8 and 1.25, share within 0.5×–1.5×
correlation of the lengths of neighbouring wordsat least 0.08
Zipf slopewithin ±0.10
shape of the words (glyphs at first, second, second-last, last position)no farther than Voynich A is from B
vertical: same-position word in the line aboveat least 1.015
formulas: two- and three-word sequences recurring at least 10 pages apartwithin 0.5×–1.5×
line edges: glyphs at line start and endwithin 0.5×–2×

Eight further properties

propertypasses if
rare words per page (herbal pages)R at most 5 (Voynich 1.96)
distinct words per word, within a pagewithin ±0.03
words unique within a pagewithin ±0.04
spread of word lengthswithin ±0.04
first lines of paragraphs as a register of their ownz above 3
spelling choices that vary by lineat least 10 of 12 classes with z above 3
agreement of word endingswithin ±0.012
neighbouring pairs seen elsewherewithin ±0.02

The line test

Passes only if all five hold: a line rarely starts with the same glyph as the line above (S1 ≤ 0.7); the link between the last glyph of a line and the first of the next is weak compared with inside a line (R ≤ 0.1); words two places apart are at least as similar as neighbours (A ≥ 1.0); at least three of the five spelling choices vary by line beyond chance; consecutive lines correlate as in the Voynich (0.207 ± 0.07).

Results

measureVoynichv17, keys 1–12v20, all 24v21, keys 1–12v21, keys 13–24 (fresh)v21, all 24
judge 1 (e231), AUC–0.563 ± 0.0080.560 ± 0.0050.554 ± 0.0060.554 ± 0.0050.554 ± 0.004
judge 2 (e266), AUC–0.601 ± 0.0060.607 ± 0.0050.596 ± 0.0080.606 ± 0.0060.601 ± 0.005
keys with judge 2 above 0.60––15 of 246 of 128 of 1214 of 24
scorecard (17 attainable)1715.015.716.015.916.0
further properties (of 8)85.85.55.55.25.4
line test passedyes11 of 1224 of 2412 of 1212 of 1224 of 24
line test: S1; A; consecutive lines r0.531; 1.047; 0.2070.521; 1.009; 0.223–0.505; 1.025; 0.1920.525; 1.032; 0.188–
lines over 1.25 times the page median (characters)6.0%16.7%11.5% (keys 1–12)11.6%––
lines over 1.5 times the page median2.9%5.0%2.8% (keys 1–12)2.8%––
text back exactly; wrong key rejected–12; 1224; 2412; 1212; 1224; 24

Properties v21 misses, with the number of keys that miss them (keys 1–12 / 13–24): gradient (12 / 12, the Voynich misses it too), page profile (12 / 12), spelling choices by line (11 / 12), spread of word lengths (9 / 11), first lines as a register (9 / 9), homogeneity (0 / 1), pairs seen elsewhere (0 / 1), rare words per page (1 / 0).

The values of the single keys range from 0.517 to 0.580 for judge 1 and from 0.551 to 0.632 for judge 2. On keys 13–24 judge 2 is 0.010 higher than on the keys used during development (0.016 for v20), about one standard error of the difference: consistent with a small optimism of the development keys, not established. Between 0.3% and 0.5% of the word triples of a book (mean 0.43%) also occur in the manuscript. For reference, the published generators pass from 1 to 10 of the 18 scorecard properties.

How to read an AUC: the real manuscript, degraded on purpose

To give the judges’ numbers a scale, the real manuscript, with its true layout, was degraded one level at a time and judged like a generated book (e400, means of four seeds):

textjudge 1judge 2scorecard (of 18)
the manuscript as it is0.5000.50017.0
lines shuffled within each page0.4810.59517.0
words shuffled within each line0.9840.99011.5
words shuffled within the page0.9890.99811.2
words drawn from other pages of the same section and language0.9981.0006.2
words drawn from other pages of the book1.0001.0005.0

The second judge’s floor for any text without the real line order is therefore about 0.60: the generated books (0.601) are at the level of the real manuscript with its lines shuffled, and far from the manuscript with its words shuffled within the lines.

With and without a message

For each of the 24 keys a second book was written with filler bits only, with the same parameters and key (e419). Against the manuscript the two kinds score the same: the paired differences (with minus without) are −0.009 ± 0.008 for judge 1 and −0.007 ± 0.010 for judge 2, the scorecard differs by −0.04 ± 0.04, and all 48 books pass the line test. A classifier trained directly to tell the pages of books with a message from those of books without one (the 228 features of judge 2, folds by key, 9,420 pages) scores AUC 0.506, and no group of features exceeds 0.51; as a negative control, the same classifier separating books of odd and even keys, all without a message, scores 0.511. These were the outcomes predicted in the preregistration. One prediction failed: the capacity of a book is not exactly the same with and without a message, because the decoder consumes a number of bits that depends on the bits themselves; capacity cannot be observed without the key.

Security

  • The building blocks are standard (scrypt, SHAKE-256, HMAC-SHA256, zlib), but the whole has not been reviewed by a cryptographer. Do not use it for secrets that matter.
  • The security is computational: it rests on SHAKE-256 behaving as a pseudorandom generator.
  • The scrypt salt is fixed: the same key gives the same master key for everybody, so a dictionary of guessed keys can be prepared once and tried on any book. Use a long random passphrase: six random words, or twelve random characters.
  • The keystream depends only on the key; there is no random value (nonce) per book. Two books written with the same key share the keystream and, because every random choice derives from the key, also their layout, the pages that lend each page its character and the places of invented words: they can be recognised as twins. A random nonce per book, sent through the decoder as its first bits, would prevent this at a cost of 128 bits; it is not implemented. Use a different key for every book.
  • The program hides that there is a message, not that the program was used: about 86% of the running words of a book are Voynich words and the invented words never are, so a book is recognisable as made with the Voynichizer, and a judge that checked each word against the Voynich vocabulary would separate the books at once.

Known limits

  • Not indistinguishable from the Voynich: the second judge still recognises the pages a little (0.601 over 24 keys, above 0.60 with 14 of them), and several properties are not reproduced (above): the page profile always, the spelling choices by line, the spread of word lengths and the first lines as a register usually.
  • The judges and the scorecard are ours, and the program was developed against them; a judge built independently has not been tried. The books of the 24 keys, the keys and the code are public: a judge built by others would be welcome.
  • Still too many lines somewhat wider than the others (11.6% over 1.25 times the median, Voynich 6.0%).
  • Only the running paragraph text is imitated: no labels, no circular or radial text. The drawings of the illustrated book are the real ones, lifted from MS 408, not generated; the script has one fixed shape per glyph.
  • It does not imitate the scribes’ hands, and it writes each page on its own, so what the two halves of a bifolium share in the manuscript is absent. Copying from the line above is modelled by the position of a word in the line, whereas the manuscript copies by physical position. Two rules are not targeted: the closure of the line at the gap left by a drawing, and the drift of marked spelling variants along the line; the short-lived agreement of the spelling choices is targeted only as agreement within a line.
  • There is no error correction: changing, adding or removing a known word breaks the reading.
  • The program imitates the statistics of the Voynich; it does not decipher it and says nothing about how it was made.

How the program came about

The program grew out of a statistical study of the Voynich (see Research), by successive versions, each measured on the same bench before being kept. The weights were never tuned on the measurement keys: they were tuned with seeds and keys of their own. Keys 1–12 were used to choose between versions; from v14 on each version was also measured on keys 13–24. The third column gives the values of the paper (with the keys they were measured on), the fourth those of keys 1–12, used during development; the early v8 benchmark used one key with three seeds and is optimistic compared with later replications.

versionchangejudges (AUC), as in the paperkeys 1–12
v0message forced into the five spelling choices of a text from an earlier generator0.958 / 0.980 (1 key)–
v1arithmetic coding over a model of the spelling choices; about 46,000 bits0.889 / 0.957 (1 key)–
v5best “copy and modify” generator–0.816 / 0.917 (3 keys)
v8generator “by pieces”: page lexicon, borrowed character, invented words, learned arrangement0.604 / 0.715 (3 seeds)–
v10message moved into the word counts0.572 / 0.690 (12 keys)0.572 / 0.690
v11–v12line test passed; the two halves of a page made different–0.576 / 0.619
v13–v14word bag refitted (v13); joins accepted only with a probable junction (v14)v14: 0.575 / 0.600 (24 keys)0.566 / 0.591
v15–v17authenticated encryption (scrypt, SHAKE-256, HMAC); page layout drawn from statistics; first public release (v17)0.559 / 0.597 (24 keys)0.563 / 0.601
v18–v20line widths: the arrangement is pushed towards the width expected for each line (v18–v19); the weight of similarity with the word two places away re-tuned (v20); published0.560 / 0.607 (24 keys)0.554 / 0.599
v21v20 with the correction of an initialisation defect in the arrangement (the counts of the spelling choices), found by reading the public code; published0.554 / 0.601 (24 keys)0.554 / 0.596

Map of the source code

filerole
voynichizzatore.pycommand line: encode, decode, empty, pdf
canale_sacco.pykey, encryption, page loop, layout, binomial chain, reading back
v1.pybinary arithmetic coder (Nasconditore hides, Rilettore reads back)
sacco.pysection lexicons, character of the page, weights, places of invented words
parole_nuove.pymodel of the one-off words, invented words, joins
disposizione.py, modello.pyplace model, links between neighbours, line widths, Metropolis arrangement
pezzi.py, pezzi_parametri_v21.jsonassembly of the pieces; the fitted parameters of v21
trascrizione.py, voynich_zl3b.jsonthe Voynich text
misure.py, e135_stato_riga.py, e145_abitudini.pyglyph segmentation, distances, spelling choices
v0.pyfile format (salva, carica), bit helpers
carattere.py, VoynichizzatoreEVA.ttfthe script
pagine.pythe text-only PDF of the command line (pdf); the site makes its own illustrated book
prova.pyread-back test on another computer
A column of painted jars, from the pharmaceutical section of the Voynich manuscript
Beinecke MS 408, f. 88r