Files
CosmicClash/TRAINING.md
T
Josh Creek 1afdc301ab feat(training): add airborne_penalty and a stage-6 "unmask" curriculum run
Stage 5 (aggression) passed (41-47 vs grounded curric-s2-defend, within
the lenient gate but not yet a clear win). Rather than keep the locomotion
mask on indefinitely, stage 6 reopens full 3D controls on top of the
aggression retune and pairs it with a new dense airborne_penalty (scaled
by height above the floor) so the policy learns to prefer staying grounded
through incentives instead of a hard mask — same regime shift that
regressed stage 3, but this time with a mitigation and ~12x the training
time (~240M timesteps / ~24h vs ~20M / ~2h) to actually re-converge
instead of stalling mid-shift.

airborne_penalty follows the existing SHIP_AI_OVERRIDES pattern: default
0 (off) on ship_ai_controller.gd, exposed via train.py's new
--airborne-penalty flag, added to training_mode.gd's allow-list. Also adds
a per-stage timesteps override in curriculum.py (STAGES[n]["timesteps"])
since this is the first stage to need a different budget than the rest.
2026-07-22 21:22:27 +01:00

13 KiB

Training the AI bot

Cosmic Clash bots are trained with reinforcement learning (self-play PPO): two ships in a headless arena share one policy that learns by playing against itself. Training runs in Python (Godot RL Agents bridge + Stable-Baselines3); the trained policy is exported to a small JSON file and runs inside the game in pure GDScript — shipped bots need no Python, no .NET, no network.

How it fits together

  • Game/scenes/training.tscn + scripts/training_mode.gd — headless self-play environment: two RL ships, randomized episode starts, goal rewards. Contains the vendored godot_rl_agents Sync node that talks TCP to the trainer.
  • scripts/ship_ai_controller.gd — training-side bridge (observations, rewards, action mapping). scripts/ship_observations.gd is the shared observation builder — training and in-game inference must stay identical, so never fork it.
  • training/train.py — PPO trainer; launches N parallel headless Godot instances (2 agents each).
  • training/export_policy.py — SB3 checkpoint → JSON policy for the game.
  • scripts/ai_ship_controller.gd + scripts/policy_network.gd — in-game inference (GDScript MLP forward pass).
  • training/evaluate.py — pits two exported policies against each other and appends to training/eval_history.json.

Hardware

The environment is our own headless Godot sim — fully cross-platform:

  • Any machine (e.g. the M4 Mac mini): fine for pipeline development, smoke runs, and short experiments. Env stepping is CPU-bound; the policy is a small MLP, so even CPU-only PPO updates are cheap.
  • Linux + NVIDIA GPU (e.g. the RTX 3090 box): recommended for real multi-hour/overnight runs. PyTorch CUDA works out of the box; more CPU cores also mean more parallel Godot instances (--n-parallel).

There is no hard GPU requirement (unlike Rocket League tooling) — a GPU mainly speeds up learning updates on long runs.

For the Linux/3090 remote-training workflow (setup, throughput tuning, auto-copying results back to the Mac, dashboard over the network), see TRAINING_LINUX.md.

Setup

Needs Python 3.10+ and a Godot 4.7 binary.

cd training
python3.12 -m venv .venv          # macOS: brew install python@3.12
.venv/bin/pip install -r requirements.txt

On Linux, download the Godot 4.7 Linux binary and point at it:

export GODOT_BIN=~/godot/Godot_v4.7.1-stable_linux.x86_64

(macOS default is /Applications/Godot.app/Contents/MacOS/Godot; override with GODOT_BIN or --godot_bin if yours lives elsewhere.)

Run a training session

cd training
.venv/bin/python train.py --experiment run01 --timesteps 20000000 --n-parallel 6 --speedup 16
  • Checkpoints land in training/checkpoints/run01/ every --checkpoint-every steps (default 100k), plus final.zip on exit (also written on Ctrl-C).
  • Resume with --resume checkpoints/run01/final.zip.
  • --n-parallel = Godot instances (2 agents each). Scale with CPU cores.
  • --speedup = in-engine physics speedup. Raise until CPU saturates.
  • --wandb mirrors logs to Weights & Biases (pip install wandb first).

Expect the smoke-run scale (~100k steps) to only learn crude ball-chasing; real behaviour needs tens of millions of steps (hours on the 3090 box).

Watch progress

.venv/bin/tensorboard --logdir training/logs

Key curves: rollout/ep_rew_mean (should trend up), rollout/ep_len_mean (should trend down from 225 as goals end episodes early — 225 action steps = the 30s episode timeout).

Reward/observation tuning

Reward weights are exported vars on ShipAIController (goal reward on TrainingMode) — tune in training.tscn/scripts without touching the trainer. If you change the observation layout (ship_observations.gd), old checkpoints/exports become incompatible: retrain, and bump a note in your experiment name.

Export a checkpoint into the game

cd training
.venv/bin/python export_policy.py checkpoints/run01/final.zip ../Game/bots/hard.json

The exporter runs a parity check (JSON forward pass vs SB3 prediction) before writing. Models live in Game/bots/.

Evaluate progress between checkpoints

TensorBoard shows learning, but "is the new checkpoint actually better?" needs head-to-head play:

.venv/bin/python export_policy.py checkpoints/run01/ppo_5000000_steps.zip /tmp/candidate.json
.venv/bin/python evaluate.py ../Game/bots/hard.json /tmp/candidate.json --episodes 40

Golden-goal episodes (first goal wins, timeout = draw), sides swapped halfway for fairness, using the exact inference path that ships in-game. Every run appends to training/eval_history.json — the long-term progress record. Evaluate each new candidate against the previous promoted bot and a fixed early reference to see absolute progress over time.

If a model was trained with the locomotion mask on (curriculum stages 1, 2, and 5 — see below), pass --grounded-a/--grounded-b for whichever side it's on. The eval otherwise runs AIShipController fully unmasked regardless of how a model was trained, so a grounded model's untrained vertical/pitch-roll output reaches the ship as noise it never had to contend with during training — this understates it, not a neutral comparison.

Difficulty tiers

A bot is (model, reaction_ticks, action_noise) — configured on the Match mode (bot_model_path, bot_reaction_ticks, bot_action_noise in match.tscn) or any AIShipController:

  • Model: the main lever. An early checkpoint is an easy bot — promote e.g. easy.json / medium.json / hard.json from different stages of one training run (verify the gaps with evaluate.py).
  • reaction_ticks (default 8 = training cadence): higher = slower reactions, easier.
  • action_noise: adds execution error, easier.

Curriculum training

Training from scratch with self-play alone hands the network every skill at once — finishing, defending, positioning, not stalling to a draw — off a sparse ±40 goal reward. train.py has a curriculum flag group that stages this the way you'd coach a human: score first, then also defend, then learn that a draw is still a failure, and only then spend compute polishing general movement. Each stage is a normal chained run — a new --experiment resumed via --resume checkpoints/<previous>/final.zip, same as any other run — just with different curriculum flags.

Stage Flags What it teaches
1 — score --opponent-mode inert --attack-goal-bias 1.0 --no-allow-vertical --no-allow-pitch-roll Team 1 is a do-nothing placeholder ship parked at its spawn (an effectively empty net); near-goal resets always target the goal the trainee attacks; the ship can't fly or pitch/roll, only drive and yaw. Isolated finishing practice.
2 — defend too --opponent-mode self_play --no-allow-vertical --no-allow-pitch-roll Reintroduces a live opponent (self-play) and the default episode-start mix — the same near-goal state is now simultaneously a finishing chance for one side and a defensive save for the other. Locomotion stays grounded.
3 — no draws --draw-penalty 5 --reset-std 0.3 Training episodes are golden-goal (end at the first goal), so there's no in-episode goal-margin to penalize — draw_penalty is the closest available signal: a one-time penalty when an episode times out with no goal at all, on top of the existing per-tick time_penalty. Also lifts the locomotion mask (full 3D controls) by omitting --allow-vertical/--allow-pitch-roll; pair that with --reset-std since the policy never got a reward gradient on those axes before now, so expect a brief re-exploration wobble.
4 — mechanics/refinement (no curriculum flags — plain next_run.sh) Stock self-play, full controls, default reward/start-state mix. This is what all runs before this feature already did.
5 — aggression --opponent-mode self_play --no-allow-vertical --no-allow-pitch-roll --velocity-to-ball-weight 0.05 --ball-distance-penalty 0.006 --ball-touch-reward 0.5 Resumes from stage 2 (curric-s2-defend), not stage 4 — see the regression note below. Retunes ball-pursuit reward weights (up from 0.02/0.002/0.4) for much more aggressive, constantly-chasing floor play, deliberately keeping the locomotion mask on so it can't reopen the stage-3 regression. Passed 2026-07-22 (41-47 vs grounded stage 2 — close, not yet a clear win).
6 — unmask --opponent-mode self_play --velocity-to-ball-weight 0.05 --ball-distance-penalty 0.006 --ball-touch-reward 0.5 --airborne-penalty 0.003 Re-opens full 3D controls on top of the aggression retune — this is the same grounded-checkpoint-to-full-3D transition that regressed stage 3, but this time paired with airborne_penalty (dense, scaled by height above the floor — see ship_ai_controller.gd) so the policy learns to prefer staying grounded through incentives instead of a hard mask, and can still pick up genuinely useful aerial/wall plays instead of never touching those axes. Runs much longer (~240M timesteps / ~24h vs every prior stage's ~20M/~2h) to actually re-converge through the regime shift instead of stalling mid-way like stage 3 did in a fifth of the time.

Stages 3-4 regressed; stage 6 deliberately reopens the same transition with a mitigation. The locomotion-mask inference bugfix (8c15c46) revealed that stage 3's evals up to that point had been running with an unfairly unmasked grounded reference. Re-evaluated fairly, curric-s2-defend (grounded) beats both curric-s3-no_draws (26-60) and curric-s4-mechanics (24-57) — lifting the locomotion mask to full 3D in stage 3 was a clear regression in floor play that self-play never earned back in 20M steps. Stage 5 sidesteps this by resuming and evaluating against stage 2 directly (curriculum.py's resume_from_experiment/reference_experiment stage-dict overrides) instead of chaining through stages 3-4. Stage 6 is where full 3D flight comes back — not masked away this time, but discouraged via airborne_penalty and given ~12x the training time to settle. See curriculum_state.json's log for the full eval numbers.

All curriculum flags default to leaving Godot's own @export defaults alone (train.py only forwards a flag when you pass it), so ordinary runs are unaffected. Full flag list: --opponent-mode {self_play,inert,frozen}, --opponent-model <path> (for frozen), --draw-penalty, --attack-goal-bias, --kickoff-chance, --near-goal-chance, --allow-vertical/--no-allow-vertical, --allow-pitch-roll/--no-allow-pitch-roll, --velocity-to-ball-weight, --ball-distance-penalty, --ball-touch-reward, --airborne-penalty.

Running it automatically

training/curriculum.py (started via curriculum.sh, same detached-tmux pattern as start_training.sh) drives all stages end to end: for each stage it runs run_training.sh (pull, train, export, commit+push) with that stage's flags, then evaluates the resulting checkpoint against a reference bot over 100 episodes — the fixed rookie.json baseline for stage 1, the previous stage's promoted checkpoint by default for stages 2+, or an explicit resume_from_experiment/reference_experiment override in that stage's dict when it deliberately skips a since-regressed branch (stage 5).

cd training
./curriculum.sh                                              # start/resume the curriculum
./curriculum.sh --seed-checkpoint checkpoints/run11/final.zip # seed stage 1 instead of a fresh policy

The gate is deliberately lenient: it blocks a stage only on a clear regression (the reference beating the candidate by 15+ points of win rate), not "must show improvement." A 40-episode eval already misled us once in this project — run11 was the first model to deliberately score a goal, but its head-to-head eval read as a loss on sample noise alone. A strict gate would have retried that stage forever for the wrong reason; a loose one still catches a genuinely broken stage. Progress and every attempt's eval result are logged to curriculum_state.json (committed alongside eval_history.json after each attempt).

A stage gets up to 2 retries (3 attempts total, each resuming from that stage's own previous attempt with a fresh --reset-std) before the script stops and asks for a human look — it will not retry indefinitely or advance past a stage that keeps failing on its own. Once you've looked at why (more timesteps? a flag needs adjusting? the eval itself was misleading?), re-run with --force-retry to try again or --skip-to-next-stage if you judge the result good enough despite the gate.

Running a stage by hand (e.g. to experiment with flags before trusting the orchestrator) still works exactly as the table above describes — just call next_run.sh/run_training.sh directly with that stage's flags.

Self-play notes

By default both ships share the live policy (mirrored, team-relative observations — see ship_observations.gd), so training is against the current self. --opponent-mode inert/frozen (see Curriculum training above) replace that with a placeholder or a fixed exported policy for one side of a run; frozen is a single-fixed-model slice of full league play. Fixed-opponent training against a pool of past checkpoints sampled per episode (to avoid strategy collapse on long self-play runs) is still deferred — see TODO.md.