mirror of
https://github.com/jcreek/CosmicClash.git
synced 2026-09-12 20:32:03 +00:00
docs(training): the stage-5 side imbalance was seed variance, not an asymmetry
The Hard-tier promotion noted a 17% physical side imbalance (physical teams 0-1 = 29-46) and flagged it as worth investigating, possibly in the arena or in ship_observations.gd's team-1 mirroring. Testing it directly shows that was wrong. Ran hard.json against itself — self-play, so any team_0/team_1 split is purely positional and cannot be a strength difference — over 10 independent seeds at 30 episodes each. Pooled: 113-125 across 300 episodes, 4.0% imbalance, sign test p = 0.48, team 1 ahead in only 3 of 10 seeds. Per-seed imbalance ranged 0.0% to 43.3%, so swings larger than the original observation happen by chance at this episode count. The underlying mistake is worth recording, and is now in TRAINING.md: evaluate.py --seed defaults to 1, so the two measurements that appeared to agree were the same paired starting-state sequence rather than independent samples, and seed 1 happens to favour team 1. Same reason the physical_side_imbalance_ceiling gate in generation5.py is a single-seed catastrophe check, not evidence about side balance.
This commit is contained in:
+20
-5
@@ -201,11 +201,26 @@ direct 100-episode head-to-head between the two finished 36-39 with 25 draws —
|
|||||||
a dead heat, so the wider indirect margin does not reflect a real strength
|
a dead heat, so the wider indirect margin does not reflect a real strength
|
||||||
difference. Attempt 3 was taken on the tiebreakers: it is the later checkpoint
|
difference. Attempt 3 was taken on the tiebreakers: it is the later checkpoint
|
||||||
(it resumed from attempt 2) and edges every telemetry metric. That head-to-head
|
(it resumed from attempt 2) and edges every telemetry metric. That head-to-head
|
||||||
also measured a 17% physical side imbalance (physical teams 0-1 = 29-46, with
|
also measured a 17% physical side imbalance (physical teams 0-1 = 29-46), which
|
||||||
`retry1` going 15-25 as team 0 but 21-14 as team 1) — inside the 20% bar used
|
looked worth investigating as a possible asymmetry in the arena or in
|
||||||
elsewhere, but large enough to be worth understanding rather than assuming it
|
`ship_observations.gd`'s team-1 mirroring. **It is not — it is seed variance.**
|
||||||
is noise, since it affects both models equally and may point at an asymmetry in
|
A follow-up ran `hard.json` against *itself* (self-play, so any split is purely
|
||||||
the arena or in `ship_observations.gd`'s team-1 mirroring.
|
positional and cannot be a strength difference) over 10 independent seeds at 30
|
||||||
|
episodes each: pooled 113-125 across 300 episodes, a 4.0% imbalance, sign test
|
||||||
|
p = 0.48, with team 1 ahead in only 3 of the 10 seeds. Per-seed imbalance
|
||||||
|
ranged from 0.0% to 43.3%, so swings far larger than the original observation
|
||||||
|
occur by chance at these episode counts.
|
||||||
|
|
||||||
|
The trap worth remembering: `evaluate.py --seed` defaults to 1, so every
|
||||||
|
evaluation in this file that did not pass `--seed` shares one paired
|
||||||
|
starting-state sequence, and seed 1 happens to favour team 1 (7-19 in the
|
||||||
|
self-play run above, 29-46 in the 100-episode head-to-head — same direction
|
||||||
|
because it is the same seed, not because it replicates). Two such runs are one
|
||||||
|
observation sampled twice, not independent confirmation. Vary the seed before
|
||||||
|
concluding anything from a side split. This also means the
|
||||||
|
`physical_side_imbalance_ceiling` gate in `generation5.py` is a single-seed
|
||||||
|
measurement and should be read as a coarse catastrophe check, not evidence
|
||||||
|
about side balance either way.
|
||||||
|
|
||||||
Every tier runs at full trained cadence (`reaction_ticks=8`, `action_noise=0`)
|
Every tier runs at full trained cadence (`reaction_ticks=8`, `action_noise=0`)
|
||||||
— the game does not manufacture difficulty gaps by handicapping a model. When
|
— the game does not manufacture difficulty gaps by handicapping a model. When
|
||||||
|
|||||||
Reference in New Issue
Block a user