The headline number in a Nature paper published this week is not a win-loss record. It is a price. An agent called Ataraxos reached what its authors call vastly superhuman play at Stratego on 16 GPUs for a week, plus four more GPUs for four days β a bill in the low thousands of dollars. The previous state of the art, Google DeepMind's DeepNash, ran two to three months across 1,024 specialized chips, which the new team estimates would have cost $3 million to $4.5 million at 2025 prices, and still did not reliably beat the best humans.
Key takeaways
- Ataraxos beat Pim Niemeijer, a four-time world champion with more than 600 weeks at world number one, 15 games to one with four draws.
- It needed roughly 34 times fewer self-play games than DeepNash β 163 million in total β and trained for a few thousand dollars rather than an estimated $3 million or more.
- The enabling piece is a second neural network, a belief model that infers hidden piece identities, which makes search at decision time tractable for the first time in Stratego.
Where the compute savings came from
The cost gap is not a story about cheaper hardware. The team wrote its own Stratego simulator in CUDA C++ that executes millions of state updates per second on graphics cards, which is the kind of engineering a lab does when it cannot requisition a field of accelerators. Senior author Gabriele Farina of MIT has said the techniques developed for games like poker simply could not scale to the volume of hidden information Stratego presents.
The sample efficiency matters more than the hardware count. Both systems learned by playing themselves, reinforcing winning moves and discarding losing ones. The difference was how aggressively Ataraxos revised its strategy after each batch: large, bold adjustments early in training, progressively smaller and more careful ones later. Hidden information tends to send self-play reinforcement learning into cycles, and that schedule is how the team kept it converging. The result was 163 million training games against DeepNash's far larger appetite β and a stronger final agent.
Why search was the missing piece
Systems like AlphaGo sharpen a general policy with a search run immediately before each move. DeepMind could not apply that in Stratego because the set of possible hidden configurations is effectively unbounded: 40 pieces per side in any arrangement produces more than a decillion setups, and a game can run to 2,000 moves against roughly 40 in chess. Whether test-time search was even worth attempting had been left open.
Ataraxos answers it with a second network trained to estimate what the opponent's pieces are, based on how they have been moved. Instead of enumerating arrangements, the agent samples plausible ones, plays candidate moves out inside each, and picks based on the outcomes. That turns an intractable search into a sampling problem β the same conceptual move that made poker solvable, applied to a game where the hidden state is orders of magnitude larger and persists for the whole match.
The behavior that emerges is distinctly non-human. Because the agent is not burdened by knowing its own secrets, it leaves weak spots untouched when it judges the opponent has no reason to suspect them, and it executes bluffs with a consistency people cannot match. Farina puts it as calculating risk in a way people cannot: a human whose most valuable piece is exposed starts to panic, while the agent stays composed and declines to overcorrect in ways that would surrender its secrets. The name comes from the ancient Greek term for freedom from anxiety.
What the human results actually show
Niemeijer played 20 online games over three weeks, earning $100 per win, and knew the agent would not adapt to him β an advantage that gave him time to probe for weaknesses. He converted one. The authors argue that single loss is not a defect, since strong Stratego play requires randomizing your own setup, which leaves luck permanently in the loop for both sides.
The wider sample is more lopsided. At the 2025 Stratego World Championship, attendees who challenged the bot lost 38 of 40 games. It has also moved the competitive metagame, legitimizing setups the community had dismissed, such as hiding the flag in a corner behind just two bombs.
How far the architecture travels
The same approach beat three world champions at Barrage Stratego, a faster variant, handled the cooperative card game Hanabi, and beat the strongest existing bots at dou dizhu. That spread across cooperative, competitive and multiplayer card games is the strongest evidence that the belief-model-plus-search recipe is general rather than tuned to one game.
The team's ambitions run to negotiations, financial markets and war gaming, on the reasoning that attacking any real problem starts by building a simplified model of it. That remains an argument rather than a demonstration β board games supply fixed rules and unambiguous winners, and the domains they name supply neither. A more immediate limitation is that Ataraxos cannot explain itself; Farina frames interpretable, explainable strategies as the goal rather than the achievement.
Outlook
The preprint has been public since November 2025, so the news here is peer review rather than the result. What it formalizes is a counterexample to the assumption that the remaining hard benchmarks need frontier-scale budgets: the gain came from an algorithmic restructuring that a six-author academic team could fund. That sits alongside a separate finding that AI systems are not yet much help at producing such restructurings themselves, after a recent benchmark found AI agents barely improve training algorithms.
FAQ
What is Ataraxos?
It is a Stratego agent from researchers at Carnegie Mellon, MIT, NYU and Stanford that pairs self-play reinforcement learning with test-time search. Its distinguishing component is a belief model that estimates the identity of the opponent's hidden pieces, which is what makes searching ahead possible in a game with that much concealed state.
Why did Stratego resist AI longer than chess, Go or poker?
The hidden information is both enormous and long-lived. Piece identities stay concealed until two pieces collide, there are more than a decillion possible starting arrangements per side, and games can stretch to 2,000 moves. Effective play also requires calibrated bluffing, which earlier self-play systems struggled to balance.
Does this transfer outside board games?
The architecture already generalized to Barrage Stratego, Hanabi and dou dizhu. Transfer to open-ended settings such as negotiation or markets is the authors' stated direction, not a published result, and the agent's lack of explainability is a practical obstacle to using it where decisions need justification.






