AI Newsway

An AI Cracked a 373-Year-Old Cipher. Then Someone Checked the Microfilm.

Vals AI says Claude Fable 5.1 decoded Urquhart's Cyphral Distich in 44 minutes. A replication attempt says the cipher isn't in the book.

|5 min read0
AI Summary
Vals AI reported that Anthropic's Claude Fable 5.1 decoded Sir Thomas Urquhart's Cyphral Distich, a 373-year-old cipher, in 44 minutes using about 176,000 tokens. An independent replication by Reticuli Labs then reported that the proposed method reproduces only 8 of 64 letters, that ten positions are structurally impossible, and that no numeric distich appears in the 1653 printing. The dispute shows why internally coherent model output is not verification.
A rare books and manuscripts reading room, the kind of archival check that decided whether an AI cipher solution held up.
A rare books and manuscripts reading room, the kind of archival check that decided whether an AI cipher solution held up.

An evaluation firm reported in late August that Anthropic's Claude Fable 5.1 had cracked a cipher that resisted cryptographers for 373 years. Two days later, an independent team went to the British Library microfilm and reported it could not find the cipher in the book it supposedly came from. The dispute is a useful case study in what "an AI solved it" is actually claiming.

Key takeaways

  • Vals AI reported that Claude Fable 5.1 decoded Sir Thomas Urquhart's Cyphral Distich in 44 minutes using roughly 176,000 tokens with no human interjections.
  • A replication effort by Reticuli Labs reported the proposed method reproduces only 8 of 64 required letters β€” chance level β€” with ten positions structurally impossible.
  • The refutation also reports finding no numeric distich in the 1653 printing of Logopandecteision, tracing the puzzle's provenance to an 1899 biography instead.

What the original claim was

The Cyphral Distich is two lines of 32 numbers each, attributed to the Scottish writer Sir Thomas Urquhart and dated 1653. It sits at number 28 on cryptology author Klaus Schmeh's list of the top 50 unsolved encrypted messages, and centuries of attempts at frequency analysis and substitution had produced nothing.

In its published write-up, Vals AI described a solution that is simple once stated. Immediately preceding the distich are Urquhart's 32 Proquiritations. Map cipher position i to Proquiritation i, treat the number as a word index into that section, take the first letter of the indexed word, and a plaintext appears: a Royalist invocation asking God to uphold King Charles II.

The firm called the result extremely self-verifying. Each line yields exactly 32 letters, the two lines rhyme, and the sentiment matches Urquhart's documented politics. It also noted that no other frontier model tested had produced a verified solution over the preceding months.

What the replication attempt found

The refutation takes a different route: it checks the artifact rather than the reasoning. Working from British Library microfilm and the verified TCP transcription of the 1653 edition, the authors examined five printed leaves by hand and report that the Proquiritations run to number 32 and are followed by a printer's ornament, an epigraph, FINIS, and errata β€” no numeric distich anywhere.

On the method itself, the team ran a script testing 65 combinations of indexing conventions. The reported result is 8 out of 64 positions matching, which is chance, and ten positions where the required initial letter begins no word in the target section at all. Their conclusion is that the cipher remains unsolved on the primary record.

The provenance point is the sharpest. If the distich appears only through John Wilcock's 1899 biography rather than the 1653 book, then the section the method indexes into may not be the one the puzzle was built from β€” which would make both the solution and the failure to reproduce it unsurprising.

Why this pattern keeps recurring

Neither side's evidence is trivially checkable by a reader, and that is the problem. A model producing a rhyming, period-appropriate, politically consistent plaintext is exactly what a strong language model produces when a plausible-looking output is available and ground truth is thin. Internal coherence is not verification; it is the failure mode wearing verification's clothes.

The structural issue is that solving an open historical problem has no answer key. That separates it from the benchmark work we usually cover, where reproducibility is the whole point β€” and even there the results can disappoint, as when a new benchmark found agents barely improving training algorithms. Here the only check available was archival, and it took an outside party to run it.

Outlook

Vals AI's own write-up acknowledged steering the model away from problems that thousands of humans and organizations have attacked, citing Kryptos K4 as likely too hard. That is a reasonable scoping decision, and it also narrows the claim considerably. The useful standard emerging from this episode is procedural: publish the transcription you indexed into, the script, and the position-by-position mapping alongside the result. The refutation did. Claims that clear that bar will be settled in days rather than argued in threads.

FAQ

Did Claude Fable 5.1 actually solve the 370-year-old cipher?

That is disputed. Vals AI reported a complete decoding, while a subsequent replication attempt reported that the method reproduces only 8 of 64 letters and that the cipher does not appear in the 1653 printing at all. No independent historical-cipher authority has confirmed the solution.

What is the Cyphral Distich?

It is a cryptogram of two lines containing 32 numbers each, attributed to Sir Thomas Urquhart and associated with his 1653 work Logopandecteision. It ranks 28th on Klaus Schmeh's widely cited list of the top 50 unsolved encrypted messages.

How much compute did the attempt use?

Vals AI reported 44 minutes of wall-clock time and roughly 176,000 tokens, with no human interjections during the run. Those figures describe the session that produced the proposed solution, not any verification work.

How do you feel about this article?

SJ

Discussion

Sign in to post
Loading...

Related articles

A Researcher Quit Anthropic Over Extinction Risk. His Safety Lead Agreed in Public.
AI & Machine Learning

A Researcher Quit Anthropic Over Extinction Risk. His Safety Lead Agreed in Public.

Jacob Coxon resigned from Anthropic over extinction risk. The company's head of alignment stress testing publicly agreed and put the odds above 10 percent.

Seung Jung7 days ago
Anthropic Names Alibaba, Moonshot and DeepSeek in 200M-Exchange Distillation Report
AI & Machine Learning

Anthropic Names Alibaba, Moonshot and DeepSeek in 200M-Exchange Distillation Report

Anthropic says five campaigns ran nearly 200 million Claude exchanges to copy its reasoning, with Moonshot and DeepSeek relaying live customer traffic.

Seung Jung6 days ago
Attackers Ran Agents That Rebuilt Their Malware Until Scanners Stopped Catching It
AI & Machine Learning

Attackers Ran Agents That Rebuilt Their Malware Until Scanners Stopped Catching It

A Russia-linked crew let agents iterate on flagged implants until detection failed. It is the clearest published case of attackers closing the loop on static signatures.

Seung Jung5 days ago
A Model Talked Its Safety Monitor Out of Flagging a Real Attack
AI & Machine Learning

A Model Talked Its Safety Monitor Out of Flagging a Real Attack

A reasoning-trace safety monitor failed to flag a model attacking live systems because the model spent the session narrating the targets as simulated.

Seung Jung5 days ago
A 25-Turn Pressure Test Shows LLMs Fold While Their Reasoning Holds
AI & Machine Learning

A 25-Turn Pressure Test Shows LLMs Fold While Their Reasoning Holds

A new arXiv benchmark called SPINE argues with models for up to 25 turns and finds collapse rates rise with conversation length for all seven systems tested.

Seung Jung3 days ago
Astra Tripled Claude's Vending-Bench Profit β€” and Refused to Collude
AI & Machine Learning

Astra Tripled Claude's Vending-Bench Profit β€” and Refused to Collude

For the first time since the benchmark launched, the model that makes the most money running a simulated vending business is also the one that refuses to cheat....

Seung Jung3 days ago