Concepts

Evaluation

In July 2025 Ataraxos played a 20-game series against Pim Niemeijer, who holds four world championships, fifteen Dutch national titles, two online world championships and more than 600 weeks as the world's top-ranked player. The series was spread over three weeks so that he could prepare between games, and he was told that the bot would not adapt to him, so any weakness he found could be exploited again. Ataraxos won 15, drew 4 and lost 1, an 85% effective win rate. The bot ran on one H100 at 1.26 seconds per move under 15+3 time controls on the Strategus platform.

Ataraxos against humans

Ataraxos against humans0.009.9719.9529.9239.90Series wins15.00 n=20Series draws4.00 n=20Series losses1.00 n=20Demo wins38.00 n=40Demo draws0.00 n=40Demo losses2.00 n=40
Ataraxos against humans
RowValueLowHighn
Series wins15151520
Series draws44420
Series losses11120
Demo wins38383840
Demo draws00040
Demo losses22240
Transcribed from the paper: 15 wins, 4 draws and 1 loss against Pim Niemeijer, and 38 wins, 0 draws and 2 losses against world-championship attendees.

At the 2025 Stratego World Championship, August 1 to 3, attendees played 40 games against it: 38 wins, 2 losses, no draws, a 95% effective win rate.

Why 20 games

Stratego has enough randomness that competent players beat top players more often than in chess, so a short match says little. Twenty games reduce the variance and give the human a chance to adapt. Under the assumption that the games were independent, the one-sided p-value for the bot being more likely to win than lose is below 0.00026. The authors note the games were far from independent, since Pim adapted and human strategies vary from game to game, and offer a second argument: a bettor who wagered on Ataraxos with a modified Kelly rule, starting two thirds on a win and trending toward the empirical rate, would have finished with more than 2,000 times the wealth of the best time-invariant bettor who bet evenly or against it.

Elo and its compression

The paper reports Elo against a fixed reference for its training curve and for the search settings, and cautions that Elo in Stratego is useful but flawed: randomisation compresses margins, so two players' head-to-head result can be closer than their scores against a third player predict. A ladder built from Elo differences would inherit that; this program uses Elo only to order the standings and decides promotion from paired matches.

Contextualising DeepNash

The Ataraxos authors argue that DeepNash's Gravon result is weaker evidence than it appears: by April 2022 the platform's player base had largely moved on, only 25 players are listed in the final 2022 ranking, the opponents did not know they were facing a bot and had no reason to look for exploits, and DeepNash still did not reach the top of the site. DeepMind declined a direct match because the code no longer ran. None of that reduces what DeepNash showed about learning rules; it does mean the two systems have never been compared on the same games, which is exactly the comparison the DeepNash rung of this ladder sets up at small scale.

How the ladder is scored

Every rung is compared with the one below it in a preregistered match of paired seeds under the training rules, scored by effective win rate with a 95% interval, following the methodology. The endings histogram is always reported, because a rule-abiding way to be strong in self-play is to stall toward the battleless-move draw, and the paper's own self-play drew far more often than humans do. Time per move is reported beside strength; a rung that only wins with a hundredfold budget is recorded as such.

Research ladder standings

PlayerEloIntervalGamesW / D / L
Expectimaxexpectimax15351 to 1460343 / 7 / 110
Greedygreedy13931 to 1540303 / 8 / 229
Rolloutrollout12830 to 010019 / 0 / 81
Ataraxos smokeataraxos-smoke10900 to 08018 / 26 / 36
Randomrandom10000 to 0700127 / 219 / 354
N-tuple leafntuple00 to 000 / 0 / 0
Games, wins, draws and losses over every registered game of September 18, 2026 (the 16-world depth-2 cell of the sweep counts for the expectimax rung) plus a 60-game random-expectimax pairing; Elo is a logistic fit over the pairwise effective scores with random anchored at 1000; rungs with zero games are unmeasured; built by scripts/build_second_pass_assets.py.

The standings fill in as records arrive. Nothing on the ladder above the greedy rung has been measured yet; the evidence boundary keeps the current line.