Files
CosmicClash/TRAINING.md
T

19 KiB

Training the AI bot

Cosmic Clash bots are trained with reinforcement learning (self-play PPO): two ships in a headless arena share one policy that learns by playing against itself. Training runs in Python (Godot RL Agents bridge + Stable-Baselines3); the trained policy is exported to a small JSON file and runs inside the game in pure GDScript — shipped bots need no Python, no .NET, no network.

How it fits together

  • Game/scenes/training.tscn + scripts/training_mode.gd — headless self-play environment: two RL ships, randomized episode starts, goal rewards. Contains the vendored godot_rl_agents Sync node that talks TCP to the trainer.
  • scripts/ship_ai_controller.gd — training-side bridge (observations, rewards, action mapping). scripts/ship_observations.gd is the shared observation builder — training and in-game inference must stay identical, so never fork it.
  • training/train.py — PPO trainer; launches N parallel headless Godot instances (2 agents each) — from source by default, or from a pre-built binary via --exported-binary (see TRAINING_LINUX.md's "Exported-binary training" section; training/export_linux.sh builds it from Game/export_presets.cfg's "Linux Training" preset).
  • training/export_policy.py — SB3 checkpoint → JSON policy for the game.
  • scripts/ai_ship_controller.gd + scripts/policy_network.gd — in-game inference (GDScript MLP forward pass).
  • training/evaluate.py — pits two exported policies against each other and appends to training/eval_history.json.

Hardware

The environment is our own headless Godot sim — fully cross-platform:

  • Any machine (e.g. the M4 Mac mini): fine for pipeline development, smoke runs, and short experiments. Env stepping is CPU-bound; the policy is a small MLP, so even CPU-only PPO updates are cheap.
  • Linux + NVIDIA GPU (e.g. the RTX 3090 box): recommended for real multi-hour/overnight runs. PyTorch CUDA works out of the box; more CPU cores also mean more parallel Godot instances (--n-parallel).

There is no hard GPU requirement (unlike Rocket League tooling) — a GPU mainly speeds up learning updates on long runs.

For the Linux/3090 remote-training workflow (setup, throughput tuning, auto-copying results back to the Mac, dashboard over the network), see TRAINING_LINUX.md.

Setup

Needs Python 3.10+ and a Godot 4.7 binary.

cd training
python3.12 -m venv .venv          # macOS: brew install python@3.12
.venv/bin/pip install -r requirements.txt

On Linux, download the Godot 4.7 Linux binary and point at it:

export GODOT_BIN=~/godot/Godot_v4.7.1-stable_linux.x86_64

(macOS default is /Applications/Godot.app/Contents/MacOS/Godot; override with GODOT_BIN or --godot_bin if yours lives elsewhere.)

Run a training session

cd training
.venv/bin/python train.py --experiment run01 --timesteps 20000000 --n-parallel 6 --speedup 16
  • Checkpoints land in training/checkpoints/run01/ every --checkpoint-every steps (default 100k), plus final.zip on exit (also written on Ctrl-C).
  • Resume with --resume checkpoints/run01/final.zip.
  • --n-parallel = Godot instances (2 agents each). Scale with CPU cores.
  • --speedup = in-engine physics speedup. Raise until CPU saturates.
  • --wandb mirrors logs to Weights & Biases (pip install wandb first).

Expect the smoke-run scale (~100k steps) to only learn crude ball-chasing; real behaviour needs tens of millions of steps (hours on the 3090 box).

Watch progress

.venv/bin/tensorboard --logdir training/logs

Key curves: rollout/ep_rew_mean (should trend up), rollout/ep_len_mean (should trend down from 225 as goals end episodes early — 225 action steps = the 30s episode timeout), rollout/goal_rate (fraction of recent episodes that ended in an actual goal rather than timing out as a draw — the live signal for "is the policy actually finishing more episodes by scoring", since ep_rew_mean mixes that with dense reward-shaping (ball chasing/ touching) and doesn't isolate it).

Reward/observation tuning

Reward weights are exported vars on ShipAIController (goal reward on TrainingMode) — tune in training.tscn/scripts without touching the trainer. If you change the observation layout (ship_observations.gd), old checkpoints/exports become incompatible: retrain, and bump a note in your experiment name.

Export a checkpoint into the game

cd training
.venv/bin/python export_policy.py checkpoints/run01/final.zip ../Game/bots/hard.json

The exporter runs a parity check (JSON forward pass vs SB3 prediction) before writing. Models live in Game/bots/.

Evaluate progress between checkpoints

TensorBoard shows learning, but "is the new checkpoint actually better?" needs head-to-head play:

.venv/bin/python export_policy.py checkpoints/run01/ppo_5000000_steps.zip /tmp/candidate.json
.venv/bin/python evaluate.py ../Game/bots/hard.json /tmp/candidate.json --episodes 40

Golden-goal episodes (first goal wins, timeout = draw), sides swapped halfway for fairness, using the exact inference path that ships in-game. Every run appends to training/eval_history.json — the long-term progress record. Evaluate each new candidate against the previous promoted bot and a fixed early reference to see absolute progress over time.

If a model was trained with the locomotion mask on (curriculum stages 1, 2, and 5 — see below), pass --grounded-a/--grounded-b for whichever side it's on. The eval otherwise runs AIShipController fully unmasked regardless of how a model was trained, so a grounded model's untrained vertical/pitch-roll output reaches the ship as noise it never had to contend with during training — this understates it, not a neutral comparison.

Difficulty tiers

A bot is (model, reaction_ticks, action_noise) — configured on the Match mode (bot_model_path, bot_reaction_ticks, bot_action_noise in match.tscn) or any AIShipController:

  • Model: the main lever. An early checkpoint is an easy bot — promote e.g. easy.json / medium.json / hard.json from different stages of one training run (verify the gaps with evaluate.py).
  • reaction_ticks (default 8 = training cadence): higher = slower reactions, easier.
  • action_noise: adds execution error, easier.

Promoted bots (Game/bots/promoted/)

Game/bots/*.json is a flat, ever-growing dump of every experiment/curriculum export — useful for evaluate.py and for A/B-ing arbitrary past checkpoints against each other in the in-game Spectate dropdown (main_menu.gd lists Game/bots/ non-recursively, so anything one directory deeper is invisible to it), but none of those filenames (run07.json, curric-s3-no_draws.json, ...) are meant to be the shipped bot — they get superseded constantly and the automated curriculum pipeline (run_training.sh) only ever writes new flat files there, never touching subdirectories.

Game/bots/promoted/<tier>.json is the small, curated, hand-maintained set actually referenced by the shipped game — currently just easy.json (promoted 2026-07-24 from curric-s6-unmask, the strongest checkpoint at the time). match.tscn/spectate.tscn point their bot_model_path exports here directly, so a promoted file is never touched by training scripts, never overwritten by a same-named future export, and never disturbed by pruning old experiment files from the flat dump.

To promote a new bot into a tier: copy the chosen Game/bots/<experiment>.json to Game/bots/promoted/<tier>.json (overwriting the old one), and note the source experiment + date in this section. Do this for medium.json/ hard.json as later curriculum stages clear the bar against easy.json in evaluate.py.

Curriculum training

Training from scratch with self-play alone hands the network every skill at once — finishing, defending, positioning, not stalling to a draw — off a sparse ±40 goal reward. train.py has a curriculum flag group that stages this the way you'd coach a human: score first, then also defend, then learn that a draw is still a failure, and only then spend compute polishing general movement. Each stage is a normal chained run — a new --experiment resumed via --resume checkpoints/<previous>/final.zip, same as any other run — just with different curriculum flags.

curriculum.py has run through two generations so far. Generation 1 (below) ran stages 1-6 to completion/block and is archived; generation 2 (the one curriculum.py actually runs today) starts a fresh stage 1 seeded from generation 1's last clean pass instead of continuing to retry a stage that kept getting worse — see "Generation 2" below.

Generation 1 (archived — see curriculum_state_gen1.json)

Stage Flags What it teaches
1 — score --opponent-mode inert --attack-goal-bias 1.0 --no-allow-vertical --no-allow-pitch-roll Team 1 is a do-nothing placeholder ship parked at its spawn (an effectively empty net); near-goal resets always target the goal the trainee attacks; the ship can't fly or pitch/roll, only drive and yaw. Isolated finishing practice.
2 — defend too --opponent-mode self_play --no-allow-vertical --no-allow-pitch-roll Reintroduces a live opponent (self-play) and the default episode-start mix — the same near-goal state is now simultaneously a finishing chance for one side and a defensive save for the other. Locomotion stays grounded.
3 — no draws --draw-penalty 5 --reset-std 0.3 Training episodes are golden-goal (end at the first goal), so there's no in-episode goal-margin to penalize — draw_penalty is the closest available signal: a one-time penalty when an episode times out with no goal at all, on top of the existing per-tick time_penalty. Also lifts the locomotion mask (full 3D controls) by omitting --allow-vertical/--allow-pitch-roll; pair that with --reset-std since the policy never got a reward gradient on those axes before now, so expect a brief re-exploration wobble.
4 — mechanics/refinement (no curriculum flags — plain next_run.sh) Stock self-play, full controls, default reward/start-state mix. This is what all runs before this feature already did.
5 — aggression --opponent-mode self_play --no-allow-vertical --no-allow-pitch-roll --velocity-to-ball-weight 0.05 --ball-distance-penalty 0.006 --ball-touch-reward 0.5 Resumes from stage 2 (curric-s2-defend), not stage 4 — see the regression note below. Retunes ball-pursuit reward weights (up from 0.02/0.002/0.4) for much more aggressive, constantly-chasing floor play, deliberately keeping the locomotion mask on so it can't reopen the stage-3 regression. Passed 2026-07-22 (41-47 vs grounded stage 2 — close, not yet a clear win).
6 — unmask --opponent-mode self_play --velocity-to-ball-weight 0.05 --ball-distance-penalty 0.006 --ball-touch-reward 0.5 --airborne-penalty 0.003 Re-opens full 3D controls on top of the aggression retune — this is the same grounded-checkpoint-to-full-3D transition that regressed stage 3, but this time paired with airborne_penalty (dense, scaled by height above the floor — see ship_ai_controller.gd) so the policy learns to prefer staying grounded through incentives instead of a hard mask, and can still pick up genuinely useful aerial/wall plays instead of never touching those axes. Failed 3 attempts in a row (25% → 20% → 15% win rate vs curric-s5-aggression) and blocked — see "Generation 2" below for what replaced it.

Stages 3-4 regressed; stage 6 deliberately reopens the same transition with a mitigation. The locomotion-mask inference bugfix (8c15c46) revealed that stage 3's evals up to that point had been running with an unfairly unmasked grounded reference. Re-evaluated fairly, curric-s2-defend (grounded) beats both curric-s3-no_draws (26-60) and curric-s4-mechanics (24-57) — lifting the locomotion mask to full 3D in stage 3 was a clear regression in floor play that self-play never earned back in 20M steps. Stage 5 sidesteps this by resuming and evaluating against stage 2 directly (curriculum.py's resume_from_experiment/reference_experiment stage-dict overrides) instead of chaining through stages 3-4. Stage 6 is where full 3D flight comes back — not masked away this time, but discouraged via airborne_penalty and given ~12x the training time to settle. See curriculum_state_gen1.json's log for the full eval numbers.

Stage 6 (unmask, retry1, retry2) all used identical flags — curriculum.py always reuses STAGES[stage_index]["flags"] on retry, only the resume checkpoint changes — so continued training just drifted the same policy further rather than converging differently (25% → 20% → 15% win rate vs curric-s5-aggression). After 3 failed attempts the script blocked for human review; rather than pile up retry4, retry5, ... on a lineage that kept getting worse, generation 2 (below) replaces it with a fresh stage 1.

Generation 2 (current)

curriculum.py's live STAGES list now contains a single stage, unmask (displays as stage 1 — curric-s1-unmask), which picks up exactly where generation 1's regression analysis left off. It resumes directly from FOUNDATION_EXPERIMENT (curric-s5-aggression's own checkpoint — the last stage that passed cleanly) via resume_from_experiment/reference_experiment overrides, rather than re-running stages 1-5 or continuing generation 1's drifted retry2:

--opponent-mode self_play --velocity-to-ball-weight 0.08 --ball-distance-penalty 0.01 --ball-touch-reward 0.7 --airborne-penalty 0.003 --ball-velocity-to-goal-weight 0.06 --goal-reward 80 --draw-penalty 5

Compared to generation 1's stage 6:

  • velocity_to_ball_weight (0.05→0.08) and ball_distance_penalty (0.006→0.01) — the actual ball-chasing terms, unchanged since stage 5 despite three failed attempts — plus ball_touch_reward (0.5→0.7).
  • Two scoring-specific terms newly exposed via train.py (they already existed as ship_ai_controller.gd/training_mode.gd @exports, just not as CLI flags): ball_velocity_to_goal_weight (0.004 default → 0.06) rewards the ball actually moving toward the goal, not just being chased/touched; goal_reward (40 default → 80) is the terminal reward for scoring itself.
  • draw_penalty 5 (proven effective in stage 3 against passivity), which generation 1's stage 6 had never set — previously all carrot for scoring, no stick for never scoring.
  • reset_retry_checkpoint: True on the stage dict, so if this stage itself fails and retries, resume_checkpoint() resets to FOUNDATION_EXPERIMENT again instead of drifting a failed attempt further — the specific bug that made generation 1's 3 retries monotonically worse instead of converging.

Deliberately not added: a cooldown/cap on ball_velocity_to_goal_weight to guard against a bot farming near-misses (bouncing the ball toward goal repeatedly without finishing) instead of actually scoring. Unlike the touch-farming bug (see ship_ai_controller.gd's ball_touch_reward comments) this term is already direction-scaled by construction (it's a velocity-toward-goal quantity, not an undirected contact count), so the risk is theoretical rather than demonstrated. If this stage's eval shows high ball_velocity_to_goal_weight accrual without a matching rise in actual goals scored, that's the signal to add one.

Every experiment name curriculum.py generates is now timestamped (YYYYMMDD-HHMM-<name>, e.g. 20260727-0930-curric-s1-unmask), applied once in run_stage_attempt — this keeps generation 2's names from colliding with generation 1's plain ones (both checkpoint directories and TensorBoard run names come straight from --experiment) and makes run order obvious in TensorBoard without cross-referencing curriculum_state.json.

All curriculum flags default to leaving Godot's own @export defaults alone (train.py only forwards a flag when you pass it), so ordinary runs are unaffected. Full flag list: --opponent-mode {self_play,inert,frozen}, --opponent-model <path> (for frozen), --draw-penalty, --attack-goal-bias, --kickoff-chance, --near-goal-chance, --allow-vertical/--no-allow-vertical, --allow-pitch-roll/--no-allow-pitch-roll, --velocity-to-ball-weight, --ball-distance-penalty, --ball-touch-reward, --airborne-penalty, --ball-velocity-to-goal-weight, --goal-reward.

Running it automatically

training/curriculum.py (started via curriculum.sh, same detached-tmux pattern as start_training.sh) drives all stages end to end: for each stage it runs run_training.sh (pull, train, export, commit+push) with that stage's flags, then evaluates the resulting checkpoint against a reference bot over 100 episodes — the fixed rookie.json baseline for a from-scratch stage 1 (no resume_from_experiment/reference_experiment override on STAGES[0]), the previous stage's promoted checkpoint by default for stages 2+, or an explicit override in that stage's dict when it deliberately skips a since-regressed branch (generation 1's stage 5) or seeds from a fixed foundation checkpoint (generation 2's stage 1 — see above).

cd training
./curriculum.sh                                              # start/resume the curriculum
./curriculum.sh --seed-checkpoint checkpoints/run11/final.zip # override stage 1's resume source for this run

The gate is deliberately lenient: it blocks a stage only on a clear regression (the reference beating the candidate by 15+ points of win rate), not "must show improvement." A 40-episode eval already misled us once in this project — run11 was the first model to deliberately score a goal, but its head-to-head eval read as a loss on sample noise alone. A strict gate would have retried that stage forever for the wrong reason; a loose one still catches a genuinely broken stage. Progress and every attempt's eval result are logged to curriculum_state.json (committed alongside eval_history.json after each attempt).

A stage gets up to 2 retries (3 attempts total) before the script stops and asks for a human look — it will not retry indefinitely or advance past a stage that keeps failing on its own. By default a retry resumes from that stage's own previous attempt with a fresh --reset-std; a stage can instead set reset_retry_checkpoint: True (generation 2's stage 1 does) to always reset to its normal resume source instead — see the generation 1 → 2 postmortem above for why blind same-checkpoint retries can make things monotonically worse. Once you've looked at why a block happened (more timesteps? a flag needs adjusting? the eval itself was misleading?), re-run with --force-retry to try again or --skip-to-next-stage if you judge the result good enough despite the gate.

Running a stage by hand (e.g. to experiment with flags before trusting the orchestrator) still works exactly as the table above describes — just call next_run.sh/run_training.sh directly with that stage's flags.

Self-play notes

By default both ships share the live policy (mirrored, team-relative observations — see ship_observations.gd), so training is against the current self. --opponent-mode inert/frozen (see Curriculum training above) replace that with a placeholder or a fixed exported policy for one side of a run; frozen is a single-fixed-model slice of full league play. Fixed-opponent training against a pool of past checkpoints sampled per episode (to avoid strategy collapse on long self-play runs) is still deferred — see TODO.md.