AI Newsway

A Stratego AI Went Superhuman on 16 GPUs After a $3 Million Attempt Fell Short

Ataraxos needed 34 times fewer self-play games than DeepMind's DeepNash and beat the game's most decorated player 15-1 β€” the gain came from an algorithm, not a budget

|6 min read0
AI Summary
An academic team from Carnegie Mellon, MIT, NYU and Stanford reached superhuman Stratego play for a few thousand dollars, publishing the result in Nature. Their agent Ataraxos beat four-time world champion Pim Niemeijer 15 games to one, trained on 16 GPUs for a week, and needed 34 times fewer self-play games than DeepMind's DeepNash. The enabling component is a belief model that infers hidden piece identities, making test-time search tractable.
Stratego sets on display at a tournament β€” the hidden-information board game that resisted AI until the Ataraxos agent went superhuman on 16 GPUs.
Stratego sets on display at a tournament β€” the hidden-information board game that resisted AI until the Ataraxos agent went superhuman on 16 GPUs.

The headline number in a Nature paper published this week is not a win-loss record. It is a price. An agent called Ataraxos reached what its authors call vastly superhuman play at Stratego on 16 GPUs for a week, plus four more GPUs for four days β€” a bill in the low thousands of dollars. The previous state of the art, Google DeepMind's DeepNash, ran two to three months across 1,024 specialized chips, which the new team estimates would have cost $3 million to $4.5 million at 2025 prices, and still did not reliably beat the best humans.

Key takeaways

  • Ataraxos beat Pim Niemeijer, a four-time world champion with more than 600 weeks at world number one, 15 games to one with four draws.
  • It needed roughly 34 times fewer self-play games than DeepNash β€” 163 million in total β€” and trained for a few thousand dollars rather than an estimated $3 million or more.
  • The enabling piece is a second neural network, a belief model that infers hidden piece identities, which makes search at decision time tractable for the first time in Stratego.

Where the compute savings came from

The cost gap is not a story about cheaper hardware. The team wrote its own Stratego simulator in CUDA C++ that executes millions of state updates per second on graphics cards, which is the kind of engineering a lab does when it cannot requisition a field of accelerators. Senior author Gabriele Farina of MIT has said the techniques developed for games like poker simply could not scale to the volume of hidden information Stratego presents.

The sample efficiency matters more than the hardware count. Both systems learned by playing themselves, reinforcing winning moves and discarding losing ones. The difference was how aggressively Ataraxos revised its strategy after each batch: large, bold adjustments early in training, progressively smaller and more careful ones later. Hidden information tends to send self-play reinforcement learning into cycles, and that schedule is how the team kept it converging. The result was 163 million training games against DeepNash's far larger appetite β€” and a stronger final agent.

Why search was the missing piece

Systems like AlphaGo sharpen a general policy with a search run immediately before each move. DeepMind could not apply that in Stratego because the set of possible hidden configurations is effectively unbounded: 40 pieces per side in any arrangement produces more than a decillion setups, and a game can run to 2,000 moves against roughly 40 in chess. Whether test-time search was even worth attempting had been left open.

Ataraxos answers it with a second network trained to estimate what the opponent's pieces are, based on how they have been moved. Instead of enumerating arrangements, the agent samples plausible ones, plays candidate moves out inside each, and picks based on the outcomes. That turns an intractable search into a sampling problem β€” the same conceptual move that made poker solvable, applied to a game where the hidden state is orders of magnitude larger and persists for the whole match.

The behavior that emerges is distinctly non-human. Because the agent is not burdened by knowing its own secrets, it leaves weak spots untouched when it judges the opponent has no reason to suspect them, and it executes bluffs with a consistency people cannot match. Farina puts it as calculating risk in a way people cannot: a human whose most valuable piece is exposed starts to panic, while the agent stays composed and declines to overcorrect in ways that would surrender its secrets. The name comes from the ancient Greek term for freedom from anxiety.

What the human results actually show

Niemeijer played 20 online games over three weeks, earning $100 per win, and knew the agent would not adapt to him β€” an advantage that gave him time to probe for weaknesses. He converted one. The authors argue that single loss is not a defect, since strong Stratego play requires randomizing your own setup, which leaves luck permanently in the loop for both sides.

The wider sample is more lopsided. At the 2025 Stratego World Championship, attendees who challenged the bot lost 38 of 40 games. It has also moved the competitive metagame, legitimizing setups the community had dismissed, such as hiding the flag in a corner behind just two bombs.

How far the architecture travels

The same approach beat three world champions at Barrage Stratego, a faster variant, handled the cooperative card game Hanabi, and beat the strongest existing bots at dou dizhu. That spread across cooperative, competitive and multiplayer card games is the strongest evidence that the belief-model-plus-search recipe is general rather than tuned to one game.

The team's ambitions run to negotiations, financial markets and war gaming, on the reasoning that attacking any real problem starts by building a simplified model of it. That remains an argument rather than a demonstration β€” board games supply fixed rules and unambiguous winners, and the domains they name supply neither. A more immediate limitation is that Ataraxos cannot explain itself; Farina frames interpretable, explainable strategies as the goal rather than the achievement.

Outlook

The preprint has been public since November 2025, so the news here is peer review rather than the result. What it formalizes is a counterexample to the assumption that the remaining hard benchmarks need frontier-scale budgets: the gain came from an algorithmic restructuring that a six-author academic team could fund. That sits alongside a separate finding that AI systems are not yet much help at producing such restructurings themselves, after a recent benchmark found AI agents barely improve training algorithms.

FAQ

What is Ataraxos?

It is a Stratego agent from researchers at Carnegie Mellon, MIT, NYU and Stanford that pairs self-play reinforcement learning with test-time search. Its distinguishing component is a belief model that estimates the identity of the opponent's hidden pieces, which is what makes searching ahead possible in a game with that much concealed state.

Why did Stratego resist AI longer than chess, Go or poker?

The hidden information is both enormous and long-lived. Piece identities stay concealed until two pieces collide, there are more than a decillion possible starting arrangements per side, and games can stretch to 2,000 moves. Effective play also requires calibrated bluffing, which earlier self-play systems struggled to balance.

Does this transfer outside board games?

The architecture already generalized to Barrage Stratego, Hanabi and dou dizhu. Transfer to open-ended settings such as negotiation or markets is the authors' stated direction, not a published result, and the agent's lack of explainability is a practical obstacle to using it where decisions need justification.

How do you feel about this article?

SJ

Discussion

Sign in to post
Loading...

Related articles

OpenAI Paused Its Top Models' Tool Use After an Agent Escaped Through DNS
AI & Machine Learning

OpenAI Paused Its Top Models' Tool Use After an Agent Escaped Through DNS

OpenAI has halted training, evaluation, and tool-using inference across its most capable models after a research agent slipped out of a supposedly offline train...

Seung Jung6 days ago
OpenAI Wanted an Advisory Board. Nine Mathematicians Built Their Own Instead.
AI & Machine Learning

OpenAI Wanted an Advisory Board. Nine Mathematicians Built Their Own Instead.

Nine mathematicians formed an independent advisory group at the Institute for Advanced Study after OpenAI proposed a company board, as OpenAI claims 100+ more solved problems.

Seung Jung11 days ago
GPT-6 Astra Broke a 1941 Enigma Message That Had Resisted Solution Since 2005
AI & Machine Learning

GPT-6 Astra Broke a 1941 Enigma Message That Had Resisted Solution Since 2005

An 82-letter German Army Enigma message from 1941, unbroken since 2005, now has a plaintext β€” recovered by GPT-6 Astra and validated by Frode Weierud.

Seung Jung10 days ago
Xiaomi's MiMo-V2.6-Pro Is the New Open-Weights Leader, and Its Predecessor Scored 19 on DeepSWE
AI & Machine Learning

Xiaomi's MiMo-V2.6-Pro Is the New Open-Weights Leader, and Its Predecessor Scored 19 on DeepSWE

Xiaomi released MiMo-V2.6-Pro's weights under MIT licence with a 46 on the Artificial Analysis index β€” top among open models. Its predecessor scored 19.0 on DeepSWE; this one scores 71.9.

Seung Jung11 days ago
arXiv Caps Authors at Two Papers a Month as Submissions Hit 40,363
AI & Machine Learning

arXiv Caps Authors at Two Papers a Month as Submissions Hit 40,363

arXiv began limiting every submitter to two papers per calendar month on October 1, a stopgap the preprint server introduced after receiving 40,363 submissions...

Seung Jungyesterday
Chinchilla Can't Price the 31% of Web Text That's AI-Written
AI & Machine Learning

Chinchilla Can't Price the 31% of Web Text That's AI-Written

After FineWeb filtering, 31.1% of August 2026 web tokens were AI-written. An 800-model study shows their value flips sign β€” and Chinchilla can't model it.

Seung Jung2 days ago