docs(training): the stage-5 side imbalance was seed variance, not an asymmetry

The Hard-tier promotion noted a 17% physical side imbalance (physical teams
0-1 = 29-46) and flagged it as worth investigating, possibly in the arena or in
ship_observations.gd's team-1 mirroring. Testing it directly shows that was
wrong.

Ran hard.json against itself — self-play, so any team_0/team_1 split is purely
positional and cannot be a strength difference — over 10 independent seeds at
30 episodes each. Pooled: 113-125 across 300 episodes, 4.0% imbalance, sign
test p = 0.48, team 1 ahead in only 3 of 10 seeds. Per-seed imbalance ranged
0.0% to 43.3%, so swings larger than the original observation happen by chance
at this episode count.

The underlying mistake is worth recording, and is now in TRAINING.md:
evaluate.py --seed defaults to 1, so the two measurements that appeared to
agree were the same paired starting-state sequence rather than independent
samples, and seed 1 happens to favour team 1. Same reason the
physical_side_imbalance_ceiling gate in generation5.py is a single-seed
catastrophe check, not evidence about side balance.
This commit is contained in:
Josh Creek
2026-08-24 08:54:26 +01:00
parent e1f512c94e
commit 08eb9f5842
+20 -5
View File
@@ -201,11 +201,26 @@ direct 100-episode head-to-head between the two finished 36-39 with 25 draws —
a dead heat, so the wider indirect margin does not reflect a real strength a dead heat, so the wider indirect margin does not reflect a real strength
difference. Attempt 3 was taken on the tiebreakers: it is the later checkpoint difference. Attempt 3 was taken on the tiebreakers: it is the later checkpoint
(it resumed from attempt 2) and edges every telemetry metric. That head-to-head (it resumed from attempt 2) and edges every telemetry metric. That head-to-head
also measured a 17% physical side imbalance (physical teams 0-1 = 29-46, with also measured a 17% physical side imbalance (physical teams 0-1 = 29-46), which
`retry1` going 15-25 as team 0 but 21-14 as team 1) — inside the 20% bar used looked worth investigating as a possible asymmetry in the arena or in
elsewhere, but large enough to be worth understanding rather than assuming it `ship_observations.gd`'s team-1 mirroring. **It is not — it is seed variance.**
is noise, since it affects both models equally and may point at an asymmetry in A follow-up ran `hard.json` against *itself* (self-play, so any split is purely
the arena or in `ship_observations.gd`'s team-1 mirroring. positional and cannot be a strength difference) over 10 independent seeds at 30
episodes each: pooled 113-125 across 300 episodes, a 4.0% imbalance, sign test
p = 0.48, with team 1 ahead in only 3 of the 10 seeds. Per-seed imbalance
ranged from 0.0% to 43.3%, so swings far larger than the original observation
occur by chance at these episode counts.
The trap worth remembering: `evaluate.py --seed` defaults to 1, so every
evaluation in this file that did not pass `--seed` shares one paired
starting-state sequence, and seed 1 happens to favour team 1 (7-19 in the
self-play run above, 29-46 in the 100-episode head-to-head — same direction
because it is the same seed, not because it replicates). Two such runs are one
observation sampled twice, not independent confirmation. Vary the seed before
concluding anything from a side split. This also means the
`physical_side_imbalance_ceiling` gate in `generation5.py` is a single-seed
measurement and should be read as a coarse catastrophe check, not evidence
about side balance either way.
Every tier runs at full trained cadence (`reaction_ticks=8`, `action_noise=0`) Every tier runs at full trained cadence (`reaction_ticks=8`, `action_noise=0`)
— the game does not manufacture difficulty gaps by handicapping a model. When — the game does not manufacture difficulty gaps by handicapping a model. When