mirror of
https://github.com/jcreek/CosmicClash.git
synced 2026-09-10 16:04:04 +00:00
docs(training): the stage-5 side imbalance was seed variance, not an asymmetry
The Hard-tier promotion noted a 17% physical side imbalance (physical teams 0-1 = 29-46) and flagged it as worth investigating, possibly in the arena or in ship_observations.gd's team-1 mirroring. Testing it directly shows that was wrong. Ran hard.json against itself — self-play, so any team_0/team_1 split is purely positional and cannot be a strength difference — over 10 independent seeds at 30 episodes each. Pooled: 113-125 across 300 episodes, 4.0% imbalance, sign test p = 0.48, team 1 ahead in only 3 of 10 seeds. Per-seed imbalance ranged 0.0% to 43.3%, so swings larger than the original observation happen by chance at this episode count. The underlying mistake is worth recording, and is now in TRAINING.md: evaluate.py --seed defaults to 1, so the two measurements that appeared to agree were the same paired starting-state sequence rather than independent samples, and seed 1 happens to favour team 1. Same reason the physical_side_imbalance_ceiling gate in generation5.py is a single-seed catastrophe check, not evidence about side balance.
This commit is contained in:
+20
-5
@@ -201,11 +201,26 @@ direct 100-episode head-to-head between the two finished 36-39 with 25 draws —
|
||||
a dead heat, so the wider indirect margin does not reflect a real strength
|
||||
difference. Attempt 3 was taken on the tiebreakers: it is the later checkpoint
|
||||
(it resumed from attempt 2) and edges every telemetry metric. That head-to-head
|
||||
also measured a 17% physical side imbalance (physical teams 0-1 = 29-46, with
|
||||
`retry1` going 15-25 as team 0 but 21-14 as team 1) — inside the 20% bar used
|
||||
elsewhere, but large enough to be worth understanding rather than assuming it
|
||||
is noise, since it affects both models equally and may point at an asymmetry in
|
||||
the arena or in `ship_observations.gd`'s team-1 mirroring.
|
||||
also measured a 17% physical side imbalance (physical teams 0-1 = 29-46), which
|
||||
looked worth investigating as a possible asymmetry in the arena or in
|
||||
`ship_observations.gd`'s team-1 mirroring. **It is not — it is seed variance.**
|
||||
A follow-up ran `hard.json` against *itself* (self-play, so any split is purely
|
||||
positional and cannot be a strength difference) over 10 independent seeds at 30
|
||||
episodes each: pooled 113-125 across 300 episodes, a 4.0% imbalance, sign test
|
||||
p = 0.48, with team 1 ahead in only 3 of the 10 seeds. Per-seed imbalance
|
||||
ranged from 0.0% to 43.3%, so swings far larger than the original observation
|
||||
occur by chance at these episode counts.
|
||||
|
||||
The trap worth remembering: `evaluate.py --seed` defaults to 1, so every
|
||||
evaluation in this file that did not pass `--seed` shares one paired
|
||||
starting-state sequence, and seed 1 happens to favour team 1 (7-19 in the
|
||||
self-play run above, 29-46 in the 100-episode head-to-head — same direction
|
||||
because it is the same seed, not because it replicates). Two such runs are one
|
||||
observation sampled twice, not independent confirmation. Vary the seed before
|
||||
concluding anything from a side split. This also means the
|
||||
`physical_side_imbalance_ceiling` gate in `generation5.py` is a single-seed
|
||||
measurement and should be read as a coarse catastrophe check, not evidence
|
||||
about side balance either way.
|
||||
|
||||
Every tier runs at full trained cadence (`reaction_ticks=8`, `action_noise=0`)
|
||||
— the game does not manufacture difficulty gaps by handicapping a model. When
|
||||
|
||||
Reference in New Issue
Block a user