Concepts
Evaluation
In July 2025 Ataraxos played a 20-game series against Pim Niemeijer, who holds four world championships, fifteen Dutch national titles, two online world championships and more than 600 weeks as the world's top-ranked player. The series was spread over three weeks so that he could prepare between games, and he was told that the bot would not adapt to him, so any weakness he found could be exploited again. Ataraxos won 15, drew 4 and lost 1, an 85% effective win rate. The bot ran on one H100 at 1.26 seconds per move under 15+3 time controls on the Strategus platform.
Ataraxos against humans
| Row | Value | Low | High | n |
|---|---|---|---|---|
| Series wins | 15 | 15 | 15 | 20 |
| Series draws | 4 | 4 | 4 | 20 |
| Series losses | 1 | 1 | 1 | 20 |
| Demo wins | 38 | 38 | 38 | 40 |
| Demo draws | 0 | 0 | 0 | 40 |
| Demo losses | 2 | 2 | 2 | 40 |
At the 2025 Stratego World Championship, August 1 to 3, attendees played 40 games against it: 38 wins, 2 losses, no draws, a 95% effective win rate.
Why 20 games
Stratego has enough randomness that competent players beat top players more often than in chess, so a short match says little. Twenty games reduce the variance and give the human a chance to adapt. Under the assumption that the games were independent, the one-sided p-value for the bot being more likely to win than lose is below 0.00026. The authors note the games were far from independent, since Pim adapted and human strategies vary from game to game, and offer a second argument: a bettor who wagered on Ataraxos with a modified Kelly rule, starting two thirds on a win and trending toward the empirical rate, would have finished with more than 2,000 times the wealth of the best time-invariant bettor who bet evenly or against it.
Elo and its compression
The paper reports Elo against a fixed reference for its training curve and for the search settings, and cautions that Elo in Stratego is useful but flawed: randomisation compresses margins, so two players' head-to-head result can be closer than their scores against a third player predict. A ladder built from Elo differences would inherit that; this program uses Elo only to order the standings and decides promotion from paired matches.
Contextualising DeepNash
The Ataraxos authors argue that DeepNash's Gravon result is weaker evidence than it appears: by April 2022 the platform's player base had largely moved on, only 25 players are listed in the final 2022 ranking, the opponents did not know they were facing a bot and had no reason to look for exploits, and DeepNash still did not reach the top of the site. DeepMind declined a direct match because the code no longer ran. None of that reduces what DeepNash showed about learning rules; it does mean the two systems have never been compared on the same games, which is exactly the comparison the DeepNash rung of this ladder sets up at small scale.
How the ladder is scored
Every rung is compared with the one below it in a preregistered match of paired seeds under the training rules, scored by effective win rate with a 95% interval, following the methodology. The endings histogram is always reported, because a rule-abiding way to be strong in self-play is to stall toward the battleless-move draw, and the paper's own self-play drew far more often than humans do. Time per move is reported beside strength; a rung that only wins with a hundredfold budget is recorded as such.
Research ladder standings
| Player | Elo | Interval | Games | W / D / L |
|---|---|---|---|---|
| Expectimaxexpectimax | 1535 | 1 to 1 | 460 | 343 / 7 / 110 |
| Greedygreedy | 1393 | 1 to 1 | 540 | 303 / 8 / 229 |
| Rolloutrollout | 1283 | 0 to 0 | 100 | 19 / 0 / 81 |
| Ataraxos smokeataraxos-smoke | 1090 | 0 to 0 | 80 | 18 / 26 / 36 |
| Randomrandom | 1000 | 0 to 0 | 700 | 127 / 219 / 354 |
| N-tuple leafntuple | 0 | 0 to 0 | 0 | 0 / 0 / 0 |
The standings fill in as records arrive. Nothing on the ladder above the greedy rung has been measured yet; the evidence boundary keeps the current line.