What to do next
Ranked by expected value, not by how interesting it sounds.
Still open — ranked by expected value
| # | attack | why it is worth doing | cost |
|---|---|---|---|
| 1 | A non-English plaintext | Every scorer used in this investigation is English. D'Agapeyeff was of Russian parentage and
quotes French cryptograms throughout the book. The solver takes its corpus from a file: swap
corpus.txt for French, Latin or transliterated Russian and rerun everything. Cheapest
untried idea on the list by a wide margin. |
hours |
| 2 | Key phrases, not key words | 6,616 single words and dates have been tried. A phrase has not — nor has a sentence cut in proof and therefore absent from the printed book, which is exactly the failure mode Section 07 describes. Generate every n-gram of the book's sentences up to length 20 and use each as a key. | days |
| 3 | A syllabary under the transposition | Ruled out as a solution on its own, not as the layer beneath. Needs an n-gram model over syllables rather than letters, which nothing in this project has. | weeks |
| 4 | Double transposition, two realistic keys | Two independent ten-letter keys is 1.3 × 1013 combinations, about eleven years of one desktop. Two twelve-letter keys is 2.3 × 1017. Not a waiting problem, and with crib dragging now closed there is nothing left to make it tractable. | years |
Closed — run, and negative
| was # | attack | what came back | best |
|---|---|---|---|
| 1 | Score only the opening | The highest-value idea on the original list, and the one most likely to have recovered a message broken part-way through. Four-gram score over the first 24 to 196 letters at every width. No width holds a plateau; at N=24 the fits beat real English, which is overfitting rather than signal. There is no cliff because there is no plateau. | see 04 |
| 2 | Double transposition, all four directions | The 2017 lead's own prediction. UU, UA, AU and AA at every width pair from 2 to 7 — 139,806,976 key pairs across 144 combinations. The three directions nobody had searched are no better than the one that had been. | −1.34 |
| 1 | Crib dragging | Built, validated, and negative. A crib's letter-repetition pattern is a hard structural constraint — checkable without the substitution, impossible to overfit — which is why it was worth building after the head-scored sweep failed for the opposite reason. On a control it returns the true plaintext ranked first at −0.7914 with the runner-up 0.52 behind. On the real cipher: 1,495,884 placements, 143,948 pattern-consistent, and the top fifteen span a band 0.021 wide. When the answer is there it separates. Nothing separates. | −1.41 |
| 3 | The midpoint split | Each half swept alone, halves swapped, each half reversed, the anomalous unit deleted. | −1.15 |
| 4 | Other unit counts | 192 to 197 units, in case the tail is not three nulls or a group was lost from the end. | −1.32 |
| 8 | Wide keys that divide 392 | Widths 28, 49 and 56, where the rectangle is complete and only one long-column pattern exists. Width 49 returns the highest raw number anywhere in this investigation, −1.19 — and it means nothing, because at width 49 each column holds four units and ordering 49 fragments will fit almost anything. | −1.19 |
Where that leaves it
The two ideas with the strongest theoretical backing — scoring only the opening, and the four-direction double transposition the 2017 finding predicts — have both now been run, and both came back empty. That matters more than another negative usually would, because between them they were the two ways this message was most likely to still be recoverable by searching.
What survives is narrower and less comfortable. Either the plaintext is not English, or the key is a phrase rather than a word, or the encipherment is internally inconsistent in a way no key can undo. The first two are cheap and remain untested. The third cannot be tested at all — only outlived.
Crib dragging was the last idea with real structural leverage, and it has now been built, proved on a control, and run. It is negative. That closes the recommendation this page made one revision ago, and it should be said plainly: there is no longer a search-shaped idea on the table with a strong prior behind it.
What remains is cheap and speculative rather than expensive and promising. Swap the corpus for French or Russian — a five-minute change nobody has made. Try key phrases instead of key words. Build a syllable model. None of these carry the weight that scoring the opening or crib dragging did, and all three could be wrong without telling us anything.
The honest reading is the one Section 05 keeps arriving at and the published record already reached: the message is most likely not internally consistent, because the man who made it could not sort a nine-letter keyword correctly in print. If that is true, no search will ever find it — and the pattern of these results, where every method that provably works on a control returns a flat band of noise here, is what that looks like from the outside.