Files
CosmicClash/TRAINING.md
T
Josh Creek 1811e9333e feat(training): curriculum generation 4 — MultiDiscrete action space redesign
Three curriculum generations (2026-07-21 through 2026-08-04) all tried
gating *when* the policy could use vertical thrust/pitch-roll on top of a
continuous Gaussian action space, and all three failed the same way: PPO's
action-distribution std collapsed within ~10% of steps and never recovered,
landing at a 15-32% win rate vs the grounded reference regardless of
mechanism (hard mask, then a gradual ramp). Generation 3's final attempt
just landed at 24% — the worst of the three.

Root cause, verified against this project's own physics: hovering this ship
requires *holding* thrust.y ~= 0.408 continuously (mass 5.0, vertical_thrust
120, gravity 9.8). A collapsed near-zero-mean Gaussian can brush that value
but never sustain it long enough to earn the reward gradient that would
move the mean — no amount of gating *when* the axis acts fixes a problem in
*how* the policy represents a decision on it. This also independently found
and fixes a real bug: godot_rl never marks an episode timeout as a
truncation, so PPO was bootstrapping V(s)=0 on every 30s draw in every
generation to date.

- Game/scripts/ship_action_codec.gd (new): single source of truth for a
  per-axis MultiDiscrete action space (7 heads, nvec [5,5,5,5,5,5,2]) shared
  by training and in-game inference, replacing the continuous Gaussian.
  thrust_y's bins are deliberately asymmetric so a random policy drifts
  through the volume instead of floor-pinning. Legacy continuous decode
  (ai_ship_controller.gd's old logic) preserved verbatim so every
  pre-generation-4 export (e.g. Game/bots/promoted/easy.json) keeps working
  unchanged via an optional "action_space" JSON field.
- ship_observations.gd: append own contact state (SIZE 31 -> 35, append-only)
  so the value function can see what wall_contact_penalty fires on.
- ship_ai_controller.gd: action space/decode via the codec; drop the
  vertical_ramp/pitch_roll_ramp mechanism entirely; tilt_penalty default
  lowered 4x (aerial approaches require pitching); flight telemetry
  (airborne_fraction, mean_altitude, air_touch_fraction, vertical_thrust_mean)
  and truncation-snapshot fields on get_info().
- training_mode.gd: new air_drill_chance state-setter branch (ball spawned
  high, ships low, kept clear of walls) so aerial practice is forced by the
  environment instead of relying on reward-driven exploration alone; snapshot
  terminal observations before a timeout reset for the truncation fix.
- cosmic_env.py: remap ShipAIController's truncated/terminal_obs info into
  SB3's TimeLimit.truncated/terminal_observation keys.
- train.py: --reset-logits (+ --reset-logits-heads) replaces the
  now-meaningless --reset-std; new EntropyFloorCallback (a persistent
  per-rollout ent_coef controller replacing the one-shot std-reset shock)
  and per-head entropy logging; FlightTelemetryCallback; --air-drill-chance/
  --tilt-penalty flags; optional AbortIfCallback kill-criterion.
- export_policy.py: writes the action_space block for MultiDiscrete models;
  index-level parity check (argmax per head) instead of comparing floats.
- curriculum.py: full rewrite — 3 stages (bootstrap/selfplay/gauntlet), no
  grounded stage, full action space live from step 1; deletes generation
  1-3's checkpoint-lineage machinery (nothing to resume from); final report
  evaluates against both promoted/easy.json and the new
  promoted/reference-grounded.json (a copy of curric-s5-aggression, the
  strongest grounded-era artifact, kept as a fixed yardstick).
- run_training.sh/.gitignore: commit only final.zip, not the ~2400
  intermediate checkpoint files a single stage was writing (~500MB ->
  ~0.2MB per run); requirements.txt pinned (behaviour here now depends on
  specific library internals, not just public APIs).
- test_action_space.py (new): offline rung-0 check catching a head-order
  mismatch before it silently corrupts 24h of training.

Validated: GDScript compiles clean (Godot --headless --import + script
validation), free_play.tscn and training.tscn both boot headless without
errors, offline action-space assertions pass. Not yet run: the actual
smoke-training/A-B validation ladder steps in TRAINING.md's "Generation 4"
section, before committing to the full ~32h curriculum.

See TRAINING.md's "Generation 4" section for the full design writeup.
2026-08-04 23:27:57 +01:00

33 KiB
Raw Blame History

Training the AI bot

Cosmic Clash bots are trained with reinforcement learning (self-play PPO): two ships in a headless arena share one policy that learns by playing against itself. Training runs in Python (Godot RL Agents bridge + Stable-Baselines3); the trained policy is exported to a small JSON file and runs inside the game in pure GDScript — shipped bots need no Python, no .NET, no network.

How it fits together

  • Game/scenes/training.tscn + scripts/training_mode.gd — headless self-play environment: two RL ships, randomized episode starts, goal rewards. Contains the vendored godot_rl_agents Sync node that talks TCP to the trainer.
  • scripts/ship_ai_controller.gd — training-side bridge (observations, rewards, action mapping). scripts/ship_observations.gd is the shared observation builder — training and in-game inference must stay identical, so never fork it.
  • training/train.py — PPO trainer; launches N parallel headless Godot instances (2 agents each) — from source by default, or from a pre-built binary via --exported-binary (see TRAINING_LINUX.md's "Exported-binary training" section; training/export_linux.sh builds it from Game/export_presets.cfg's "Linux Training" preset).
  • training/export_policy.py — SB3 checkpoint → JSON policy for the game.
  • scripts/ai_ship_controller.gd + scripts/policy_network.gd — in-game inference (GDScript MLP forward pass).
  • training/evaluate.py — pits two exported policies against each other and appends to training/eval_history.json.

Hardware

The environment is our own headless Godot sim — fully cross-platform:

  • Any machine (e.g. the M4 Mac mini): fine for pipeline development, smoke runs, and short experiments. Env stepping is CPU-bound; the policy is a small MLP, so even CPU-only PPO updates are cheap.
  • Linux + NVIDIA GPU (e.g. the RTX 3090 box): recommended for real multi-hour/overnight runs. PyTorch CUDA works out of the box; more CPU cores also mean more parallel Godot instances (--n-parallel).

There is no hard GPU requirement (unlike Rocket League tooling) — a GPU mainly speeds up learning updates on long runs.

For the Linux/3090 remote-training workflow (setup, throughput tuning, auto-copying results back to the Mac, dashboard over the network), see TRAINING_LINUX.md.

Setup

Needs Python 3.10+ and a Godot 4.7 binary.

cd training
python3.12 -m venv .venv          # macOS: brew install python@3.12
.venv/bin/pip install -r requirements.txt

On Linux, download the Godot 4.7 Linux binary and point at it:

export GODOT_BIN=~/godot/Godot_v4.7.1-stable_linux.x86_64

(macOS default is /Applications/Godot.app/Contents/MacOS/Godot; override with GODOT_BIN or --godot_bin if yours lives elsewhere.)

Run a training session

cd training
.venv/bin/python train.py --experiment run01 --timesteps 20000000 --n-parallel 6 --speedup 16
  • Checkpoints land in training/checkpoints/run01/ every --checkpoint-every steps (default 100k), plus final.zip on exit (also written on Ctrl-C).
  • Resume with --resume checkpoints/run01/final.zip.
  • --n-parallel = Godot instances (2 agents each). Scale with CPU cores.
  • --speedup = in-engine physics speedup. Raise until CPU saturates.
  • --wandb mirrors logs to Weights & Biases (pip install wandb first).

Expect the smoke-run scale (~100k steps) to only learn crude ball-chasing; real behaviour needs tens of millions of steps (hours on the 3090 box).

Watch progress

.venv/bin/tensorboard --logdir training/logs

Key curves: rollout/ep_rew_mean (should trend up), rollout/ep_len_mean (should trend down from 225 as goals end episodes early — 225 action steps = the 30s episode timeout), rollout/goal_rate (fraction of recent episodes that ended in an actual goal rather than timing out as a draw — the live signal for "is the policy actually finishing more episodes by scoring", since ep_rew_mean mixes that with dense reward-shaping (ball chasing/ touching) and doesn't isolate it).

Reward/observation tuning

Reward weights are exported vars on ShipAIController (goal reward on TrainingMode) — tune in training.tscn/scripts without touching the trainer. If you change the observation layout (ship_observations.gd), old checkpoints/exports become incompatible: retrain, and bump a note in your experiment name.

Export a checkpoint into the game

cd training
.venv/bin/python export_policy.py checkpoints/run01/final.zip ../Game/bots/hard.json

The exporter runs a parity check (JSON forward pass vs SB3 prediction) before writing. Models live in Game/bots/.

Evaluate progress between checkpoints

TensorBoard shows learning, but "is the new checkpoint actually better?" needs head-to-head play:

.venv/bin/python export_policy.py checkpoints/run01/ppo_5000000_steps.zip /tmp/candidate.json
.venv/bin/python evaluate.py ../Game/bots/hard.json /tmp/candidate.json --episodes 40

Golden-goal episodes (first goal wins, timeout = draw), sides swapped halfway for fairness, using the exact inference path that ships in-game. Every run appends to training/eval_history.json — the long-term progress record. Evaluate each new candidate against the previous promoted bot and a fixed early reference to see absolute progress over time.

If a model was trained with the locomotion mask on (curriculum stages 1, 2, and 5 — see below), pass --grounded-a/--grounded-b for whichever side it's on. The eval otherwise runs AIShipController fully unmasked regardless of how a model was trained, so a grounded model's untrained vertical/pitch-roll output reaches the ship as noise it never had to contend with during training — this understates it, not a neutral comparison.

Difficulty tiers

A bot is (model, reaction_ticks, action_noise) — configured on the Match mode (bot_model_path, bot_reaction_ticks, bot_action_noise in match.tscn) or any AIShipController:

  • Model: the main lever. An early checkpoint is an easy bot — promote e.g. easy.json / medium.json / hard.json from different stages of one training run (verify the gaps with evaluate.py).
  • reaction_ticks (default 8 = training cadence): higher = slower reactions, easier.
  • action_noise: adds execution error, easier.

Promoted bots (Game/bots/promoted/)

Game/bots/*.json is a flat, ever-growing dump of every experiment/curriculum export — useful for evaluate.py and for A/B-ing arbitrary past checkpoints against each other in the in-game Spectate dropdown (main_menu.gd lists Game/bots/ non-recursively, so anything one directory deeper is invisible to it), but none of those filenames (run07.json, curric-s3-no_draws.json, ...) are meant to be the shipped bot — they get superseded constantly and the automated curriculum pipeline (run_training.sh) only ever writes new flat files there, never touching subdirectories.

Game/bots/promoted/<tier>.json is the small, curated, hand-maintained set actually referenced by the shipped game — currently easy.json (promoted 2026-07-24 from curric-s6-unmask, the strongest checkpoint at the time — note curric-s6-unmask was itself generation 1's failed unmask stage, so easy.json is weaker than reference-grounded.json below; a strong generation 4 result should promote a real replacement, plus medium.json/hard.json) and reference-grounded.json (added for generation 4 — a copy of generation 3's curric-s5-aggression, made before the flat Game/bots/ dump was scrapped for the redesign, kept as the strongest grounded-era artifact and the fixed yardstick generations 1-3 were all measured against; see "Generation 4"'s final report). match.tscn/ spectate.tscn point their bot_model_path exports here directly, so a promoted file is never touched by training scripts, never overwritten by a same-named future export, and never disturbed by pruning old experiment files from the flat dump.

To promote a new bot into a tier: copy the chosen Game/bots/<experiment>.json to Game/bots/promoted/<tier>.json (overwriting the old one), and note the source experiment + date in this section. Do this for medium.json/ hard.json as later curriculum stages clear the bar against easy.json in evaluate.py.

Curriculum training

Training from scratch with self-play alone hands the network every skill at once — finishing, defending, positioning, not stalling to a draw — off a sparse ±40 goal reward. train.py has a curriculum flag group that stages this the way you'd coach a human: score first, then also defend, then learn that a draw is still a failure, and only then spend compute polishing general movement. Each stage is a normal chained run — a new --experiment resumed via --resume checkpoints/<previous>/final.zip, same as any other run — just with different curriculum flags.

curriculum.py has run through four generations so far. Generation 1 (below) ran stages 1-6 to completion/block and is archived; generation 2 started a fresh stage 1 seeded from generation 1's last clean pass instead of continuing to retry a stage that kept getting worse, but also failed 3 attempts; generation 3 replaced generation 2's single all-or-nothing unmask stage with a gradual ramp, and also failed (worse, on its final attempt, than either prior generation); generation 4 (the one curriculum.py actually runs today) is a full redesign, not a further patch — see "Generation 4" below, and "Generation 3" for why a fourth attempt at gating when the policy could use full 3D controls was abandoned rather than retried again.

Generation 1 (archived — see curriculum_state_gen1.json)

Stage Flags What it teaches
1 — score --opponent-mode inert --attack-goal-bias 1.0 --no-allow-vertical --no-allow-pitch-roll Team 1 is a do-nothing placeholder ship parked at its spawn (an effectively empty net); near-goal resets always target the goal the trainee attacks; the ship can't fly or pitch/roll, only drive and yaw. Isolated finishing practice.
2 — defend too --opponent-mode self_play --no-allow-vertical --no-allow-pitch-roll Reintroduces a live opponent (self-play) and the default episode-start mix — the same near-goal state is now simultaneously a finishing chance for one side and a defensive save for the other. Locomotion stays grounded.
3 — no draws --draw-penalty 5 --reset-std 0.3 Training episodes are golden-goal (end at the first goal), so there's no in-episode goal-margin to penalize — draw_penalty is the closest available signal: a one-time penalty when an episode times out with no goal at all, on top of the existing per-tick time_penalty. Also lifts the locomotion mask (full 3D controls) by omitting --allow-vertical/--allow-pitch-roll; pair that with --reset-std since the policy never got a reward gradient on those axes before now, so expect a brief re-exploration wobble.
4 — mechanics/refinement (no curriculum flags — plain next_run.sh) Stock self-play, full controls, default reward/start-state mix. This is what all runs before this feature already did.
5 — aggression --opponent-mode self_play --no-allow-vertical --no-allow-pitch-roll --velocity-to-ball-weight 0.05 --ball-distance-penalty 0.006 --ball-touch-reward 0.5 Resumes from stage 2 (curric-s2-defend), not stage 4 — see the regression note below. Retunes ball-pursuit reward weights (up from 0.02/0.002/0.4) for much more aggressive, constantly-chasing floor play, deliberately keeping the locomotion mask on so it can't reopen the stage-3 regression. Passed 2026-07-22 (41-47 vs grounded stage 2 — close, not yet a clear win).
6 — unmask --opponent-mode self_play --velocity-to-ball-weight 0.05 --ball-distance-penalty 0.006 --ball-touch-reward 0.5 --airborne-penalty 0.003 Re-opens full 3D controls on top of the aggression retune — this is the same grounded-checkpoint-to-full-3D transition that regressed stage 3, but this time paired with airborne_penalty (dense, scaled by height above the floor — see ship_ai_controller.gd) so the policy learns to prefer staying grounded through incentives instead of a hard mask, and can still pick up genuinely useful aerial/wall plays instead of never touching those axes. Failed 3 attempts in a row (25% → 20% → 15% win rate vs curric-s5-aggression) and blocked — see "Generation 2" below for what replaced it.

Stages 3-4 regressed; stage 6 deliberately reopens the same transition with a mitigation. The locomotion-mask inference bugfix (8c15c46) revealed that stage 3's evals up to that point had been running with an unfairly unmasked grounded reference. Re-evaluated fairly, curric-s2-defend (grounded) beats both curric-s3-no_draws (26-60) and curric-s4-mechanics (24-57) — lifting the locomotion mask to full 3D in stage 3 was a clear regression in floor play that self-play never earned back in 20M steps. Stage 5 sidesteps this by resuming and evaluating against stage 2 directly (curriculum.py's resume_from_experiment/reference_experiment stage-dict overrides) instead of chaining through stages 3-4. Stage 6 is where full 3D flight comes back — not masked away this time, but discouraged via airborne_penalty and given ~12x the training time to settle. See curriculum_state_gen1.json's log for the full eval numbers.

Stage 6 (unmask, retry1, retry2) all used identical flags — curriculum.py always reuses STAGES[stage_index]["flags"] on retry, only the resume checkpoint changes — so continued training just drifted the same policy further rather than converging differently (25% → 20% → 15% win rate vs curric-s5-aggression). After 3 failed attempts the script blocked for human review; rather than pile up retry4, retry5, ... on a lineage that kept getting worse, generation 2 (below) replaces it with a fresh stage 1.

Generation 2 (archived — see curriculum_state_gen2.json)

curriculum.py's STAGES list contained a single stage, unmask (displays as stage 1 — curric-s1-unmask), which picks up exactly where generation 1's regression analysis left off. It resumes directly from FOUNDATION_EXPERIMENT (curric-s5-aggression's own checkpoint — the last stage that passed cleanly) via resume_from_experiment/reference_experiment overrides, rather than re-running stages 1-5 or continuing generation 1's drifted retry2:

--opponent-mode self_play --velocity-to-ball-weight 0.08 --ball-distance-penalty 0.01 --ball-touch-reward 0.7 --airborne-penalty 0.003 --ball-velocity-to-goal-weight 0.06 --goal-reward 80 --draw-penalty 5

Compared to generation 1's stage 6:

  • velocity_to_ball_weight (0.05→0.08) and ball_distance_penalty (0.006→0.01) — the actual ball-chasing terms, unchanged since stage 5 despite three failed attempts — plus ball_touch_reward (0.5→0.7).
  • Two scoring-specific terms newly exposed via train.py (they already existed as ship_ai_controller.gd/training_mode.gd @exports, just not as CLI flags): ball_velocity_to_goal_weight (0.004 default → 0.06) rewards the ball actually moving toward the goal, not just being chased/touched; goal_reward (40 default → 80) is the terminal reward for scoring itself.
  • draw_penalty 5 (proven effective in stage 3 against passivity), which generation 1's stage 6 had never set — previously all carrot for scoring, no stick for never scoring.
  • reset_retry_checkpoint: True on the stage dict, so if this stage itself fails and retries, resume_checkpoint() resets to FOUNDATION_EXPERIMENT again instead of drifting a failed attempt further — the specific bug that made generation 1's 3 retries monotonically worse instead of converging.

Deliberately not added: a cooldown/cap on ball_velocity_to_goal_weight to guard against a bot farming near-misses (bouncing the ball toward goal repeatedly without finishing) instead of actually scoring. Unlike the touch-farming bug (see ship_ai_controller.gd's ball_touch_reward comments) this term is already direction-scaled by construction (it's a velocity-toward-goal quantity, not an undirected contact count), so the risk is theoretical rather than demonstrated. If this stage's eval shows high ball_velocity_to_goal_weight accrual without a matching rise in actual goals scored, that's the signal to add one.

Every experiment name curriculum.py generates is now timestamped (YYYYMMDD-HHMM-<name>, e.g. 20260727-0930-curric-s1-unmask), applied once in run_stage_attempt — this keeps generation 2's names from colliding with generation 1's plain ones (both checkpoint directories and TensorBoard run names come straight from --experiment) and makes run order obvious in TensorBoard without cross-referencing curriculum_state.json.

Generation 2 also failed 3 attempts in a row, landing at a stable 32% / 28% / 31% win rate vs curric-s5-aggression each time — the second and third attempts each continued the same checkpoint lineage for another full 240M steps with zero improvement, ruling out both the reward retune above and "just needs more time" as fixes. Every attempt showed train/std collapsing from ~0.30 to ~0.13-0.15 within the first ~10% of steps and never recovering. See "Generation 3" below for the redesign this prompted.

Generation 3 (current)

Generation 2's failures point at the mechanism of the transition, not the reward weights: flipping allow_vertical/allow_pitch_roll from false to true in one step let PPO's action-distribution std collapse on those axes before the policy ever meaningfully explored them. Generation 3 replaces that boolean mask with a float ramp (vertical_ramp/pitch_roll_ramp on ShipAIController, 0.0-1.0, multiplying the axis's effect in set_action instead of gating it) and spreads the transition across 4 stages instead of 1:

Stage vertical-ramp/pitch-roll-ramp airborne-penalty timesteps gated
1 — unmask-ramp25 0.25 0.0 40M (~4h) No — trains, checkpoints, always advances
2 — unmask-ramp50 0.5 0.001 40M (~4h) No
3 — unmask-ramp75 0.75 0.002 40M (~4h) No
4 — unmask 1.0 0.003 240M (~24h) Yes — evaluated against curric-s5-aggression, same 15-point regression gate as every prior attempt

The 3 warmup stages are deliberately ungated: they're waypoints en route to the real, measured transition, not decisions in their own right, so curriculum.py's main() loop trains and checkpoints them and always advances (no eval call, no retry logic — there's nothing to fail against). Only the final unmask stage is evaluated, with the same reference bot, opponent mode (self_play, not frozen — kept identical to every prior attempt so a pass or fail cleanly isolates the ramp as the only variable), and 240M-step budget as all 3 failed all-or-nothing attempts, for a direct comparison. airborne_penalty ramps in step with the axes so it doesn't fight a still-mostly-inert axis early on.

All the reward-shaping flags from generation 2's stage (velocity-to-ball-weight, ball-distance-penalty, ball-touch-reward, ball-velocity-to-goal-weight, goal-reward, draw-penalty) are unchanged and identical across all 4 stages, so the ramp is the sole studied variable.

curriculum_state.json was reset (generation 2's log archived to curriculum_state_gen2.json) rather than continuing to log against a stage list whose stage 0 no longer means what it used to.

Open question, resolved 2026-08-04. The gated unmask stage failed all 3 attempts: 29% → 30% → 24% win rate vs curric-s5-aggression (the third, worst by then) — landing in the same ~28-32% band the section above flagged as "evidence the plateau isn't an exploration/collapse problem at all." train/std collapsed from ~0.30 to ~0.13-0.15 within the first ~10% of steps in every attempt of every generation regardless of hard-mask vs. gradual-ramp mechanism, so gating when the axes were allowed to act never addressed the actual cause. See "Generation 4" below for the redesign and root-cause diagnosis this prompted, and curriculum_state_gen3.json for the archived full log.

Generation 4 (current) — action space redesign, not a further ramp patch

Three generations spent ~2 weeks trying different ways to gate when the policy could use vertical thrust/pitch/roll on top of a continuous Gaussian action space, and all three converged on the same failure: PPO's action std collapsing within the first ~10% of steps and never recovering, regardless of mechanism. Research into how self-play PPO bots that have actually solved this class of problem (RLGym/RLBot's Necto/Nexto) approach it turned up a structural difference — they don't gate control authority at all; they train the full action space from step 1 using discrete/bucketed actions, not a continuous Gaussian, plus reward/state-setter curriculum instead of action masking.

Root cause, verified against this project's own physics (not assumed): flight in this game is a sustained set-point, not an impulse. Ship mass 5.0, vertical_thrust 120 (ship.gd/ship.tscn), default gravity 9.8 m/s² → hovering requires holding thrust.y ≈ 0.408 continuously. A Gaussian whose mean sits near 0 and whose σ has collapsed to ~0.13 samples thrust.y ∈ [-0.4, 0.4] — it can brush the hover value but can never hold it long enough to earn the reward gradient that would move the mean. That's a fixed point; no ramp on the axis's downstream effect (which is applied after PPO samples the action) moves it, exactly as the "open question" above speculated might be the case. Independently, this redesign also found and fixed a real, previously-unnoticed bug unrelated to the action space: godot_rl's godot_env.py never marks an episode timeout as a truncation (it returns the same done array for both term and trunc — see its own # TODO update API to term, trunc), so PPO was bootstrapping V(s_T)=0 on every 30s draw in every generation to date instead of correctly estimating the value of the state it timed out in.

Action space: switched to per-axis MultiDiscrete (7 heads, nvec = [5,5,5,5,5,5,2]) instead of continuous Box(7) — see Game/scripts/ship_action_codec.gd, the single source of truth for the layout/decode shared by training and in-game inference. Not a single lookup table (RLGym's approach for Rocket League's coupled car controls): Cosmic Clash's 7 axes are near-independent thruster/torque channels, so a curated combination table would throw away that factorization for no benefit. thrust_y's bins are deliberately asymmetric (-0.5, 0, 0.45, 0.75, 1.0, vs. the symmetric -1, -0.5, 0, 0.5, 1 on every other axis) — a uniform-random policy over those 5 bins averages 0.34, just below the 0.408 hover point, so a fresh policy drifts gently through the volume instead of pinning to the floor (symmetric bins) or sticking to the ceiling (ceiling_pull_strength 11.5 > gravity 9.8). This is the direct analogue of the RLGym/RLBot fix for the same failure mode ("add more jump actions to the discrete action parser"). godot_rl's ActionSpaceProcessor already emits MultiDiscrete with zero Python-side changes when every action entry is Discrete — the only reason this project's action space flattened to Box(7) before was that turbo (binary) was mixed with continuous entries.

Backward compatibility: every export before generation 4 (e.g. Game/bots/promoted/easy.json) has no "action_space" field in its JSON; absence means {"type": "continuous"} and decodes through the exact same path as before (ShipActionCodec.from_continuous, moved verbatim out of ai_ship_controller.gd). PolicyNetwork.gd's forward pass itself never changed — only the caller's decode branches on the model's declared type. export_policy.py's parity check is now index-level for a MultiDiscrete model (argmax per head's logit slice, compared against SB3's own deterministic=True chosen index) rather than comparing clipped floats, since a head-order mistake would otherwise train and export cleanly and only surface as silently wrong in-game behaviour.

No grounded stage. Full action space live from step 1 — no successful self-play RL bot in this problem class gates control authority, it's failed 9/9 attempts (3 generations × 3 attempts) here, and every prior generation's checkpoints are a different, incompatible action/observation shape anyway (nothing to resume from). 3 stages instead of a ramp:

Stage Opponent Timesteps Gated What it teaches
1 — bootstrap inert 40M (~4h) No Empty-net finishing from a random policy — no moving target, full action space from the start.
2 — selfplay self_play 160M (~16h) Yes, vs stage 1 Where essentially all the learning happens.
3 — gauntlet frozen = stage 2's own export 120M (~12h) Yes, vs stage 2 A stationary opponent for a low-variance measurement, and a check that self-play didn't converge to a fixed point that only beats itself.

An "air drill" state-setter branch (training_mode.gd's air_drill_chance, new — ball spawned high, both ships spawned low and lateral, unsolvable without climbing, kept clear of every wall so the RLGym-warned wall-bounce exploit has no wall nearby to bounce off) runs at a constant rate across all stages rather than being introduced late — gating when a skill gets drilled would reproduce the exact "gate what the policy can do" pattern that failed 3 generations running.

Observations: ShipObservations.SIZE grew 31 → 35 (own contact normal

  • an in_contact flag, appended — never inserted, see that file's append-only invariant) so the value function can actually see the condition wall_contact_penalty fires on, instead of predicting a reward with no supporting signal.

Reward shaping: mostly unchanged — a farmability check on the existing weights (velocity_to_ball_weight's term telescopes to ~2.7 over a 20m approach, well under goal_reward=80; not gameable) argues generation 2/3's tuning was never the actual problem. Two changes: airborne_penalty is no longer passed by any stage (previously ramped up in lockstep with the axis generation 3 was trying to teach — directly adversarial to the goal of genuine aerial play), and tilt_penalty dropped 4x (0.002 → 0.0005 default) since an aerial approach to a high ball requires pitching. Deliberately not added: a standalone air-touch reward — that's the exact exploit RLGym warns about ("hits the ball off a wall high up instead of doing a real aerial"); the air-drill state setter already makes aerial skill instrumentally necessary to earn the existing ball-directed rewards.

Exploration: --reset-std (meaningless under MultiDiscrete — no log_std) is replaced by --reset-logits <scale> (multiplies action_net's weights/bias, optionally scoped to specific heads via --reset-logits-heads) for a deliberate post-diagnosis recovery, and more importantly by --entropy-floor (train.py's EntropyFloorCallback): a persistent per-rollout controller nudging ent_coef to hold policy entropy near a target that decays over the run, replacing the one-shot --reset-std shock that reliably decayed away within ~10% of steps in every prior generation with something that responds continuously instead of once. --ent-coef's default rose 0.0001 → 0.01 (tuned for MultiDiscrete's bounded ~10-nat entropy, not a Gaussian's unbounded differential entropy). Per-head entropy (train/entropy_head_<name>) replaces the old aggregate train/std scalar — it identifies which axis is collapsing instead of one number for all seven.

Validation before spending the full ~32h budget: see the ladder below — cheapest checks first (an offline action-space assertion, a headless Godot boot, a 100k-step smoke run, export parity + an in-game round trip against easy.json), then flight telemetry (rollout/airborne_fraction, mean_altitude, air_touch_fraction, vertical_thrust_mean — leading indicators visible from the first rollout instead of only in a win rate measured a full run later), then a short controlled A/B (MultiDiscrete vs. continuous, otherwise identical, ~20M steps each) before committing to the full curriculum — every past generation bet a full day on an unfalsifiable hypothesis, which is what made each failure expensive to diagnose.

  1. training/test_action_space.py — offline, seconds. Catches a head-order mismatch, the single most likely silent killer (trains "fine" for 24h, produces garbage — e.g. pitch commands driving strafe thrusters — with no error).
  2. godot --headless --path Game res://scenes/training.tscn with no trainer listening, 30s — catches class_name/observation-size regressions.
  3. .venv/bin/python train.py --experiment smoke --timesteps 100000 --n-parallel 2 — confirms the MultiDiscrete handshake and new callback metrics emit.
  4. export_policy.py on the smoke checkpoint (mandatory index-level parity check), then evaluate.py <smoke>.json ../Game/bots/promoted/easy.json --episodes 4 — exercises the real GDScript decode path.
  5. A short A/B: two 20M-step runs, identical except action space (MultiDiscrete vs. the old continuous Box(7)), comparing rollout/airborne_fraction. If discrete pulls meaningfully ahead, the 32h curriculum is a justified bet; if both stay near zero, the hypothesis above is wrong and reward/compute explanations move to the front — cheaper than a 4th blind multi-day generation either way.

All curriculum flags default to leaving Godot's own @export defaults alone (train.py only forwards a flag when you pass it), so ordinary runs are unaffected. Full flag list: --opponent-mode {self_play,inert,frozen}, --opponent-model <path> (for frozen), --draw-penalty, --attack-goal-bias, --kickoff-chance, --near-goal-chance, --air-drill-chance (generation 4's state-setter aerial curriculum), --velocity-to-ball-weight, --ball-distance-penalty, --ball-touch-reward, --airborne-penalty, --tilt-penalty, --ball-velocity-to-goal-weight, --goal-reward. (--vertical-ramp/--pitch-roll-ramp are gone — generation 4 has no locomotion mask/ramp to control.)

Running it automatically

training/curriculum.py (started via curriculum.sh, same detached-tmux pattern as start_training.sh) drives all stages end to end: for each stage it runs run_training.sh (pull, train, export, commit+push) with that stage's flags, then evaluates the resulting checkpoint against a reference bot over 100 episodes — the previous stage's own passing export (stage 1 is ungated, so this only applies to stages 2+). Once every stage passes, a final (non-gating) report evaluates the result against both Game/bots/promoted/easy.json (the shipped bot) and Game/bots/promoted/reference-grounded.json (a copy of generation 3's curric-s5-aggression, the strongest grounded-era artifact and the yardstick generations 1-3 were all measured against) — those two numbers are what actually answer "did generation 4 work?"

cd training
./curriculum.sh                                            # start/resume the curriculum
./curriculum.sh --seed-checkpoint checkpoints/some/final.zip # override stage 1's resume source for this run

The gate is deliberately lenient: it blocks a stage only on a clear regression (the reference beating the candidate by 15+ points of win rate), not "must show improvement." A 40-episode eval already misled us once in this project — run11 was the first model to deliberately score a goal, but its head-to-head eval read as a loss on sample noise alone. A strict gate would have retried that stage forever for the wrong reason; a loose one still catches a genuinely broken stage. Progress and every attempt's eval result are logged to curriculum_state.json (committed alongside eval_history.json after each attempt).

A stage gets up to 2 retries (3 attempts total) before the script stops and asks for a human look — it will not retry indefinitely or advance past a stage that keeps failing on its own. By default a retry resumes from that stage's own previous attempt (no reset_retry_checkpoint stage override is set in generation 4 — nothing yet suggests a retry needs to reset to a clean upstream checkpoint the way generation 3's single unmask stage did; add one if a stage's retries turn out to be drifting rather than converging). Once you've looked at why a block happened (more timesteps? a flag needs adjusting? the eval itself was misleading?), re-run with --force-retry to try again or --skip-to-next-stage if you judge the result good enough despite the gate.

Running a stage by hand (e.g. to experiment with flags before trusting the orchestrator) still works exactly as the table above describes — just call next_run.sh/run_training.sh directly with that stage's flags.

Self-play notes

By default both ships share the live policy (mirrored, team-relative observations — see ship_observations.gd), so training is against the current self. --opponent-mode inert/frozen (see Curriculum training above) replace that with a placeholder or a fixed exported policy for one side of a run; frozen is a single-fixed-model slice of full league play. Fixed-opponent training against a pool of past checkpoints sampled per episode (to avoid strategy collapse on long self-play runs) is still deferred — see TODO.md.