fix(training): correct stage-3 eval (locomotion-mask bugfix) and add grounded aggression stage

Re-ran stage-3 (curric-s3-no_draws vs curric-s2-defend) and the missing
stage-4 gate now that the locomotion-mask inference bugfix is in. Both
reverse or contradict the pre-fix bookkeeping: curric-s2-defend (grounded)
beats curric-s3-no_draws 60-26 and curric-s4-mechanics 57-24 when fairly
evaluated, so lifting the locomotion mask in stage 3 was a real regression
in floor play, not the improvement the buggy eval reported.

Adds a stage-5 "aggression" curriculum entry that resumes from stage 2
directly (via new resume_from_experiment/reference_experiment stage-dict
overrides in curriculum.py) instead of compounding the regression through
stages 3-4, keeps the locomotion mask on, and retunes ball-pursuit reward
weights for much more aggressive floor play. Extends train.py with the
three new --velocity-to-ball-weight/--ball-distance-penalty/--ball-touch-reward
flags needed to forward that retune to Godot's existing SHIP_AI_OVERRIDES.

curriculum_state.json and TRAINING.md are corrected/annotated in place
rather than silently rewritten, so the regression stays visible in history.
This commit is contained in:
Josh Creek
2026-07-22 12:48:45 +01:00
parent cf4859e61c
commit fca6a46200
5 changed files with 149 additions and 20 deletions
+24 -6
View File
@@ -122,8 +122,8 @@ appends to `training/eval_history.json` — the long-term progress record.
Evaluate each new candidate against the previous promoted bot and a fixed
early reference to see absolute progress over time.
If a model was trained with the locomotion mask on (curriculum stages 1-2 —
see below), pass `--grounded-a`/`--grounded-b` for whichever side it's on.
If a model was trained with the locomotion mask on (curriculum stages 1, 2,
and 5 — see below), pass `--grounded-a`/`--grounded-b` for whichever side it's on.
The eval otherwise runs `AIShipController` fully unmasked regardless of how a
model was trained, so a grounded model's untrained vertical/pitch-roll output
reaches the ship as noise it never had to contend with during training —
@@ -159,22 +159,40 @@ just with different curriculum flags.
| 2 — defend too | `--opponent-mode self_play --no-allow-vertical --no-allow-pitch-roll` | Reintroduces a live opponent (self-play) and the default episode-start mix — the same near-goal state is now simultaneously a finishing chance for one side and a defensive save for the other. Locomotion stays grounded. |
| 3 — no draws | `--draw-penalty 5 --reset-std 0.3` | Training episodes are golden-goal (end at the *first* goal), so there's no in-episode goal-margin to penalize — `draw_penalty` is the closest available signal: a one-time penalty when an episode times out with no goal at all, on top of the existing per-tick `time_penalty`. Also lifts the locomotion mask (full 3D controls) by omitting `--allow-vertical`/`--allow-pitch-roll`; pair that with `--reset-std` since the policy never got a reward gradient on those axes before now, so expect a brief re-exploration wobble. |
| 4 — mechanics/refinement | *(no curriculum flags — plain `next_run.sh`)* | Stock self-play, full controls, default reward/start-state mix. This is what all runs before this feature already did. |
| 5 — aggression | `--opponent-mode self_play --no-allow-vertical --no-allow-pitch-roll --velocity-to-ball-weight 0.05 --ball-distance-penalty 0.006 --ball-touch-reward 0.5` | **Resumes from stage 2 (`curric-s2-defend`), not stage 4** — see the regression note below. Retunes ball-pursuit reward weights (up from 0.02/0.002/0.4) for much more aggressive, constantly-chasing floor play, deliberately keeping the locomotion mask on so it can't reopen the stage-3 regression. |
> **Stages 3-4 regressed and are parked.** The locomotion-mask inference bugfix
> (`8c15c46`) revealed that stage 3's evals up to that point had been running
> with an unfairly unmasked grounded reference. Re-evaluated fairly,
> `curric-s2-defend` (grounded) beats both `curric-s3-no_draws` (26-60) and
> `curric-s4-mechanics` (24-57) — lifting the locomotion mask to full 3D in
> stage 3 was a clear regression in floor play that self-play never earned
> back. Stage 5 sidesteps this by resuming and evaluating against stage 2
> directly (`curriculum.py`'s `resume_from_experiment`/`reference_experiment`
> stage-dict overrides) instead of chaining through stages 3-4. Full 3D
> flight is parked as a separate initiative — see TODO.md — that will need a
> redesigned unmasking approach (more timesteps and/or reward rebalancing) so
> it doesn't cost floor fundamentals again. See `curriculum_state.json`'s
> stage-2/stage-3 log entries for the full eval numbers.
All curriculum flags default to leaving Godot's own `@export` defaults
alone (`train.py` only forwards a flag when you pass it), so ordinary runs
are unaffected. Full flag list: `--opponent-mode {self_play,inert,frozen}`,
`--opponent-model <path>` (for `frozen`), `--draw-penalty`,
`--attack-goal-bias`, `--kickoff-chance`, `--near-goal-chance`,
`--allow-vertical`/`--no-allow-vertical`, `--allow-pitch-roll`/`--no-allow-pitch-roll`.
`--allow-vertical`/`--no-allow-vertical`, `--allow-pitch-roll`/`--no-allow-pitch-roll`,
`--velocity-to-ball-weight`, `--ball-distance-penalty`, `--ball-touch-reward`.
### Running it automatically
`training/curriculum.py` (started via `curriculum.sh`, same detached-tmux
pattern as `start_training.sh`) drives all four stages end to end: for each
pattern as `start_training.sh`) drives all stages end to end: for each
stage it runs `run_training.sh` (pull, train, export, commit+push) with that
stage's flags, then evaluates the resulting checkpoint against a reference
bot — the fixed `rookie.json` baseline for stage 1, or the previous stage's
promoted checkpoint for stages 2-4 — over 100 episodes.
bot over 100 episodes — the fixed `rookie.json` baseline for stage 1, the
previous stage's promoted checkpoint by default for stages 2+, or an
explicit `resume_from_experiment`/`reference_experiment` override in that
stage's dict when it deliberately skips a since-regressed branch (stage 5).
```bash
cd training