[09] How it could be cracked
[09] · D'AGAPEYEFF INVESTIGATION

How it could be cracked

A strategic assessment: why 222 million rejections are informative, the flaw shared by all of them, and the one objective that has never been tried.

The argument from unicity — why failure is informative

Before asking how to crack it, ask whether it is crackable in principle. A keyword-filled square costs perhaps 20–40 bits of real key (it is a dictionary word, not a random permutation of 25 letters). A columnar key of width 12 adds log₂(12!) ≈ 29 bits. Call the whole thing 50–70 bits.

English carries about 3.2 bits of redundancy per letter, so 196 letters carry roughly 627 bits of it. Unicity distance — the length past which only one key can produce sensible plaintext — is around 22 letters. We have 196.

So a correct solution, if one exists, is uniquely identifiable nine times over. There is no hiding place where several keys all produce plausible English. Which means the 222 million rejections are not bad luck or insufficient compute. If the answer were inside the space searched, it would have stood up and separated, the way it does on every control. One of the assumptions underneath the search must be wrong.

The flaw shared by all 222 million attempts

Every attack in Section 04 — without exception — had to guess a substitution in order to score a key. Undo a transposition and you hold 196 meaningless two-digit units; to ask "is this English?" you must first decide which letter each unit stands for. That single requirement did two kinds of damage.

It made every candidate expensive. A hill-climb over 18 symbols costs a millisecond or two. Multiply by a key space and the ceiling arrives fast: two seven-letter keys was as far as double transposition could reach.

And it welded the search to English. Every score came from a 147,626-letter English corpus. A French, Latin or Russian plaintext would have been rejected at every step, silently, by machinery that could not see it.

Both problems have the same root, and both have the same fix.

The fix — score without the substitution, and without the language

E7 established that the count of recurring n-grams is invariant under any one-to-one substitution. That was used there as a negative — to prove no substitution could explain the cipher. Turned around, it is the objective this investigation has been missing.

Undo a candidate transposition and simply count how often the unit sequence repeats itself. Pairs, triples, quads. No substitution required, because the count does not depend on one. No language model required, because every natural language repeats itself and random orderings do not. A correct key restores the repeats; a wrong key does not.

Measured on controls — English through a square and a real columnar key, 40 trials per width, 3,000 random keys compared against the true one each time:

widthtrue keyrandom keysseparationrandom keys beating it
9191.764.0 ± 13.49.5 σ1.48%
11198.160.1 ± 11.711.7 σ0.29%
13210.762.4 ± 12.411.7 σ0.49%

Nine to twelve sigma. The true key is not merely detectable, it is unmistakable — and it was detected without anyone deciding what a single unit means.

What that unlocks

A candidate now costs about ten microseconds instead of two milliseconds, and the score is deterministic so no restarts are needed. On sixteen cores:

exhaustive searchcandidatesold waynew way
single columnar, every key at w=12479,001,60016.6 h5 min
single columnar, every key at w=136,227,020,8009 d1.1 h
single columnar, every key at w=1487,178,291,2000.3 y15 h
single columnar, every key at w=151,307,674,368,0005.2 y9.5 d
double, two keys both ≤ 81,625,702,40056 h17 min
double, two keys both ≤ 9131,681,894,4000.5 y23 h
double, two keys both ≤ 1013,168,189,440,00052 y0.3 y

The double-transposition barrier moves out by two full key letters — from two seven-letter keys to two nine-letter keys for the same effort, and two ten-letter keys become a four-month run rather than a fifty-year one. That is the single largest gain available, and it lands squarely on the family the 2017 finding points at.

It also makes the single columnar searches genuinely exhaustive for the first time. The dynamic program in Section 04 is exact only given a substitution; it was iterated from seeds, so gaps were possible. Enumerating every key at w ≤ 13 leaves none.

And it tests every language at once, for free. If the plaintext is French or Russian, this objective finds the key anyway — the identification of the language comes afterwards, from the recovered order, not before.

The untested cipher family — a turning grille

Everything tried has been columnar, or a fixed geometric route, or two of those in series. One classical family has never been touched, and the arithmetic is uncomfortable.

A turning grille is a square card with holes, laid over a grid, written through, then rotated ninety degrees and written through again, four times in all. On a 14 × 14 board a valid grille has exactly 49 holes, and four rotations write exactly 196 letters. Not approximately. Exactly.

The fit goes further. A turning grille reorders every letter independently — which is precisely the fingerprint E8 measured and could not account for with any columnar arrangement. And d'Agapeyeff knew grilles: he devotes pages 113–115 to one, telling a long story about a Russian professor breaking a grille message in the Baltic states after 1918.

The argument against: the grille he describes is Cardan's — steganography, hiding words in an innocent letter — not a turning grille used as a transposition. He never teaches the turning variety. So this would be him reaching past his own book, which he does not obviously do elsewhere.

Why it is still worth doing: a grille is 49 independent four-way choices, about 98 bits, which is under the unicity bound and therefore recoverable in principle. And it decomposes: each orbit of four cells can be scored almost independently, so hill-climbing with the repeat objective is tractable where brute force (449) never would be. Nobody has tried it. It is the largest genuinely unexplored hypothesis on the board.

The path that is not a search at all

Two published analyses are cited everywhere and have been read by nobody in this investigation, this page included. David Shulman in The Cryptogram, April/May 1952 and March/April 1959; and Wayne G. Barker in Cryptologia, 1978. They are quoted only at second hand, from bibliographies.

Shulman's first article appeared in 1952 — the same year the challenge was quietly cut from the reprint. That is either coincidence or it is the whole story, and the article would say which. An American Cryptogram Association member can pull both from the archive in an afternoon.

Beyond that: Oxford University Press production files for the 1939 Meridian series would show what changed between proof and print, which is where a shifting key would leave its fingerprint. D'Agapeyeff's own papers — he was a working cartographer with an institutional life — may survive somewhere. Any of these could contain the answer outright, and all of them are cheaper than another hundred million keys.

Honest odds

routechance it cracks itreasoning
Exhaustive single columnar, repeat objective, w≤14 ~10% Closes the gaps the seeded dynamic program may have left, and tests every language at once. Low odds only because this family has already been hammered hardest.
Double transposition, two keys ≤ 9–10 ~20% The family the 2017 finding points at, now reachable. The best-motivated hypothesis meeting the first method that can actually search it.
Turning grille ~10% 196 = 4 × 49 is a hard structural coincidence and the reordering signature matches exactly. Discounted because it is outside what the book teaches.
Non-English plaintext ~10% Comes free with the new objective rather than as a separate effort. Real but modest prior.
Archival — Shulman, Barker, OUP files ~15% The only route that could produce a definitive answer rather than a candidate, and the only one that costs no compute at all.
Nothing works, because the message is broken ~40% The published consensus, the conclusion of Section 05, and the reading most consistent with nine separate methods that provably solve controls all returning flat noise here.

Those do not sum to 100 because they are not exclusive — the last line is the background against which the others are bets. Taken together: roughly a 50–60% chance this cipher is recoverable at all, and if it is, the combination most likely to do it is the repeat objective pointed at double transposition. That is one program, a few days of compute, and it has never been run.

If I had one week and this machine

Day 1. Build the repeat objective into the solver as a first-class scorer. It is perhaps eighty lines. Re-run every single columnar width exhaustively to w=13 as a regression test — it should reproduce every negative in Section 04, and if it does not, something in this investigation is wrong and that matters more than the cipher.

Days 2–3. Double transposition, both keys to width 9, exhaustive. 132 billion pairs, about a day. This is the run that has the best claim on being the one that works.

Day 4. Turning grille. Hill-climb 49 orbits under the repeat objective, many restarts. Cheap, unexplored, and structurally motivated by 196 = 4 × 49.

Days 5–7. Two keys to width 10 in the background while writing to the ACA for the Shulman articles. The letter costs nothing and might end the whole thing.

And if all of that comes back flat: that is a result too. It would make the broken-message reading close to the only one standing, and this page should say so rather than keep proposing searches.