Reports · 2026-09-18
The self-play smoke run trains but does not yet beat random
Self-play smoke network after twenty iterations
| Row | Value | Low | High | n |
|---|---|---|---|---|
| network vs random | 60 | 44.6 | 73.7 | 40 |
| network vs greedy | 17.5 | 8.7 | 32 | 40 |
Entropy did not collapse
The registration named one failure mode: the paper's damping schedules annealing too fast on a tiny run and collapsing the policy. The opposite happened. Move entropy started at 2.70 nats, rose to 3.23 by iteration 7 and ended at 3.08; the lowest value was the first one, far above the half-of-initial floor. The importance-ratio clip rate stayed under 5% and the reverse KL to the magnet policy fell slowly from 0.135 to 0.093. The policy moved little in twenty iterations, which is what small updates under strong regularisation should do at the start of the paper's schedule.
Move entropy and clip rate over the smoke run
- move entropy
- clip rate
- KL to magnet
| Series | iteration | entropy (nats) or clip rate |
|---|---|---|
| move entropy | 1 | 2.7007 |
| move entropy | 2 | 3.0273 |
| move entropy | 3 | 3.1233 |
| move entropy | 4 | 3.1743 |
| move entropy | 5 | 3.1957 |
| move entropy | 6 | 3.2272 |
| move entropy | 7 | 3.2317 |
| move entropy | 8 | 3.2019 |
| move entropy | 9 | 3.2301 |
| move entropy | 10 | 3.2217 |
| move entropy | 11 | 3.1736 |
| move entropy | 12 | 3.1942 |
| move entropy | 13 | 3.1887 |
| move entropy | 14 | 3.1757 |
| move entropy | 15 | 3.1363 |
| move entropy | 16 | 3.1132 |
| move entropy | 17 | 3.124 |
| move entropy | 18 | 3.1339 |
| move entropy | 19 | 3.1634 |
| move entropy | 20 | 3.0788 |
| clip rate | 1 | 0.0241 |
| clip rate | 2 | 0.0438 |
| clip rate | 3 | 0.0128 |
| clip rate | 4 | 0.02 |
| clip rate | 5 | 0.0038 |
| clip rate | 6 | 0.0122 |
| clip rate | 7 | 0.0091 |
| clip rate | 8 | 0.0112 |
| clip rate | 9 | 0.0063 |
| clip rate | 10 | 0.0056 |
| clip rate | 11 | 0.0072 |
| clip rate | 12 | 0.0064 |
| clip rate | 13 | 0.0083 |
| clip rate | 14 | 0.0087 |
| clip rate | 15 | 0.0042 |
| clip rate | 16 | 0.0037 |
| clip rate | 17 | 0.0014 |
| clip rate | 18 | 0.0004 |
| clip rate | 19 | 0.0018 |
| clip rate | 20 | 0.0234 |
| KL to magnet | 1 | 0.1348 |
| KL to magnet | 2 | 0.1386 |
| KL to magnet | 3 | 0.1418 |
| KL to magnet | 4 | 0.1442 |
| KL to magnet | 5 | 0.1342 |
| KL to magnet | 6 | 0.1392 |
| KL to magnet | 7 | 0.1205 |
| KL to magnet | 8 | 0.1289 |
| KL to magnet | 9 | 0.12 |
| KL to magnet | 10 | 0.113 |
| KL to magnet | 11 | 0.115 |
| KL to magnet | 12 | 0.1208 |
| KL to magnet | 13 | 0.116 |
| KL to magnet | 14 | 0.1061 |
| KL to magnet | 15 | 0.1055 |
| KL to magnet | 16 | 0.1076 |
| KL to magnet | 17 | 0.1041 |
| KL to magnet | 18 | 0.0997 |
| KL to magnet | 19 | 0.0991 |
| KL to magnet | 20 | 0.0926 |
The network stalls rather than loses
Twenty-six of the forty games against random ended without a Flag capture, which under the training rules means the 100-move battleless limit. A network that has learned little about attacking plays quiet moves, random play rarely finds the Flag, and the game times out. The three losses and the seven wins against greedy show the network can lose material and can take a Flag; it has not learned to force either. The paper's own training used 100 battleless moves for the same reason it gives: to discourage dawdling. On this scale the counter plane alone did not do that in twenty iterations.
28%11/40, 16% to 43%Deviations recorded
The evaluator alternates the network's seat by game parity with one fresh
seed per game, so the forty games are seat-balanced but not paired on
identical setups; the network's own setups came from its setup network
rather than being uniform, because the run saved one. Neither changes
the outcome. The next registration scales the run (more iterations, the
small preset) and adds the draw-counter diagnostic before any claim
about the Ataraxos rung.
Reproduction
Registration experiments/db0047e7-dd11-4f65-8928-fcee3f1c0a6b.toml. Commands, from the research
directory with the virtualenv from just setup and the bindings from
just build-env:
.venv/bin/python -m armies_train train --preset smoke --iterations 20 --out runs/db0047e7-dd11-4f65-8928-fcee3f1c0a6b/smoke --seed 0
.venv/bin/python -m armies_train evaluate --run runs/db0047e7-dd11-4f65-8928-fcee3f1c0a6b/smoke --opponent random --games 40 --seed 0
.venv/bin/python -m armies_train evaluate --run runs/db0047e7-dd11-4f65-8928-fcee3f1c0a6b/smoke --opponent greedy --games 40 --seed 0
python3 scripts/build_experiment_assets.pyThe record carries the SHA-256 of the history, configuration and evaluation files; the weights are kept out of git.