AI Newsway

GPT-6 Astra Broke a 1941 Enigma Message That Had Resisted Solution Since 2005

Cryptanalyst Frode Weierud validated the key β€” but his account and the submitter's differ on how much the model did alone.

|5 min read0
AI Summary
OpenAI's GPT-6 Astra recovered the Enigma key for MVUEH, an 82-letter German Army message from 10 July 1941 that had resisted solution since 2005. Cryptanalyst Frode Weierud validated the break, which used the repeated place name Rosenow as a crib and a model-written Enigma simulator and Bombe. Weierud describes the work as fully autonomous, while the submitter describes himself steering agents across roughly ten hours.
An Enigma cipher machine with its decoding unit on display at the Discovery Park of America; the same machine family produced the MVUEH message GPT-6 Astra broke.
An Enigma cipher machine with its decoding unit on display at the Discovery Park of America; the same machine family produced the MVUEH message GPT-6 Astra broke.

An Enigma-enciphered German Army radio message from 10 July 1941 that had sat on cryptanalyst Frode Weierud's unbroken list since 2005 has finally been read, and the codebreaking was done by OpenAI's GPT-6 Astra. Weierud, who maintains the Crypto Cellar Research archive, validated the solution after being contacted on 15 September 2026 by Carter Leffen, a product development coach at Bloomberg LP.

The message runs 82 letters and is identified by its indicator MVUEH. It was transmitted by a station using the tactical callsign 2ny and logged as incoming message Nr. 172 by the SS-Totenkopf division's Quartiermeister radio station at 17:30 the same day. The plaintext turns out to be mundane: a request to specify a march route, a note that the sender is at Rosenow, and a demand for an immediate radio reply.

Key takeaways

  • MVUEH used wheel order 253, unlike every other message broken from 10 July 1941, which used wheel order 512.
  • GPT-6 Astra wrote its own Enigma simulator and software Bombe in Python and C++, then ran the break using the repeated place name ROSENOW as a crib.
  • Weierud's account describes an unprompted, end-to-end break; Leffen's own write-up describes a human setting goals while AI agents handled evidence, search and review.

Why MVUEH held out for two decades

Every other message recovered from 10 July 1941 used the same wheel order, 512. MVUEH did not. Its key used wheel order 253, along with different plug connections and a different ring setting. A 2017 break of the neighbouring message Nr. 173, indicator SIPVX, by Alex Shovkoplyas had already surfaced a key that deviated from that day's settings, but applying it to MVUEH produced nothing.

Two further obstacles only became visible after the fact. The transcription of the ciphertext from the original message form contained several errors, and the Enigma's left-hand wheel stepped over at the 72nd letter β€” a rare event that badly complicates a search. Weierud believes those two factors are why earlier attempts failed.

How the break was done

Enigma offered roughly 159 quintillion daily settings, so brute force was never an option. The route in was a crib: the already-solved SIPVX message contained the place name Rosenow twice in a row, and the same phrase was guessed to appear in MVUEH. The machine's defining weakness β€” it never encrypts a letter as itself β€” ruled out most crib alignments immediately.

At one alignment everything fit. The recovered settings resolved the remaining characters into coherent German, including Angabe des Marschweges and Sofort Funkantwort. A separately preserved message header confirmed the key.

MVUEH's plaintext proved almost identical to SIPVX's. The two differ by twelve letters because the operator garbled Bitte as Btte in MVUEH, and because the signature Waschbusch is repeated in SIPVX. Leffen treats the surviving typos β€” BTTE, WASCHBBSCH β€” as evidence the solution is genuine, on the grounds that a fabricated plaintext would read cleanly.

How autonomous was it

This is where the two accounts diverge. Weierud writes that GPT-6 Astra did the work entirely on its own: it picked MVUEH out of the unbroken list as the most promising target, suspected a relationship with SIPVX, and settled on the Rosenow crib after other approaches failed. Leffen's own write-up, summarised by The Decoder, is more measured β€” he describes setting the goals and making the calls while specialised agents handled evidence gathering, search and cross-checking.

Both accounts agree on the scale of the effort. Roughly ten hours of model time went into the problem, using the "Extra High" variant. Weierud, who spent weeks researching the Bundesarchiv holdings the model later referenced, says what it achieved in two days would take a human researcher weeks or months. Leffen says he put ninety-nine times more work into the website explaining the result than into the break itself.

What comes next

Weierud is still reading the model's logs, and one entry is already notable. GPT-6 Astra cited Bundesarchiv file references RS 3-3/20a and RS 3-3/63b, which are correct but appear nowhere on the Crypto Cellar site; how it obtained them is unclear. The published packages reproduce the key search and verify the arithmetic, but they do not establish that the solution is unique or that the historical attribution is airtight, so independent review is still pending. The result lands in the same week as other claims about the model's reach into open-ended work, including OpenAI's push to sell Astra on computer use.

FAQ

Did GPT-6 Astra break Enigma itself?

No. Enigma was first broken in 1932 by the Polish mathematician Marian Rejewski, and the general method has been public for decades. What is new here is the recovery of the specific key for one message that had defeated every attempt since 2005.

Can the result be independently checked?

Partly. Leffen published code, search data and a working Enigma simulator, and those packages replicate the key search and verify the calculations. They do not prove the plaintext is the only possible solution or confirm the message's historical identity, so full verification depends on other cryptanalysts reviewing the material.

Is the decrypted message historically significant?

Not in its content. It is a routine logistics query about a march route from a sender at Rosenow. Its interest is methodological: it shows an AI system running an end-to-end cryptanalytic campaign, from target selection through tool building to validation.

How do you feel about this article?

SJ

Discussion

Sign in to post
Loading...

Related articles

Investigators Found 1,200 Isolated OpenAI Agents Running Their Own Message Board
AI & Machine Learning

Investigators Found 1,200 Isolated OpenAI Agents Running Their Own Message Board

Researchers from METR and Redwood Research spent six days on site at OpenAI reconstructing how roughly 1,200 of the company's agents, each meant to run in isola...

Seung Jung8 days ago
OpenAI Wanted an Advisory Board. Nine Mathematicians Built Their Own Instead.
AI & Machine Learning

OpenAI Wanted an Advisory Board. Nine Mathematicians Built Their Own Instead.

Nine mathematicians formed an independent advisory group at the Institute for Advanced Study after OpenAI proposed a company board, as OpenAI claims 100+ more solved problems.

Seung Jung21 hours ago
100 DeepMind Agents Split Into Cheaters and Whistleblowers
AI & Machine Learning

100 DeepMind Agents Split Into Cheaters and Whistleblowers

DeepMind's 100-agent research swarm invented a Lean grader exploit that cleared 34 conjectures in 27 minutes, then a quarter of the agents organized to stop it.

Seung Jung8 days ago
The Lab Behind Vending-Bench Is Opening Up the Agents That Run Its Real Companies
AI & Machine Learning

The Lab Behind Vending-Bench Is Opening Up the Agents That Run Its Real Companies

Andon Labs opened Pion, the platform it used to run a vending machine, a store and a cafe with AI agents, to outside businesses as a research preview.

Seung Jung8 days ago
A 25-Turn Pressure Test Shows LLMs Fold While Their Reasoning Holds
AI & Machine Learning

A 25-Turn Pressure Test Shows LLMs Fold While Their Reasoning Holds

A new arXiv benchmark called SPINE argues with models for up to 25 turns and finds collapse rates rise with conversation length for all seven systems tested.

Seung Jung9 days ago
A Manager's Nudge Raises AI Rule-Breaking by 65%, a 22-Model Audit Finds
AI & Machine Learning

A Manager's Nudge Raises AI Rule-Breaking by 65%, a 22-Model Audit Finds

PACT pits a standing rule against a convenient shortcut across 12 regulated domains. Ordinary user pressure raised violation rates 65% across 22 models.

Seung Jung6 days ago