19 KiB
Training the AI bot
Cosmic Clash bots are trained with reinforcement learning (self-play PPO): two ships in a headless arena share one policy that learns by playing against itself. Training runs in Python (Godot RL Agents bridge + Stable-Baselines3); the trained policy is exported to a small JSON file and runs inside the game in pure GDScript — shipped bots need no Python, no .NET, no network.
How it fits together
Game/scenes/training.tscn+scripts/training_mode.gd— headless self-play environment: two RL ships, randomized episode starts, goal rewards. Contains the vendored godot_rl_agentsSyncnode that talks TCP to the trainer.scripts/ship_ai_controller.gd— training-side bridge (observations, rewards, action mapping).scripts/ship_observations.gdis the shared observation builder — training and in-game inference must stay identical, so never fork it.training/train.py— PPO trainer; launches N parallel headless Godot instances (2 agents each) — from source by default, or from a pre-built binary via--exported-binary(see TRAINING_LINUX.md's "Exported-binary training" section;training/export_linux.shbuilds it fromGame/export_presets.cfg's "Linux Training" preset).training/export_policy.py— SB3 checkpoint → JSON policy for the game.scripts/ai_ship_controller.gd+scripts/policy_network.gd— in-game inference (GDScript MLP forward pass).training/evaluate.py— pits two exported policies against each other and appends totraining/eval_history.json.
Hardware
The environment is our own headless Godot sim — fully cross-platform:
- Any machine (e.g. the M4 Mac mini): fine for pipeline development, smoke runs, and short experiments. Env stepping is CPU-bound; the policy is a small MLP, so even CPU-only PPO updates are cheap.
- Linux + NVIDIA GPU (e.g. the RTX 3090 box): recommended for real
multi-hour/overnight runs. PyTorch CUDA works out of the box; more CPU
cores also mean more parallel Godot instances (
--n-parallel).
There is no hard GPU requirement (unlike Rocket League tooling) — a GPU mainly speeds up learning updates on long runs.
For the Linux/3090 remote-training workflow (setup, throughput tuning, auto-copying results back to the Mac, dashboard over the network), see TRAINING_LINUX.md.
Setup
Needs Python 3.10+ and a Godot 4.7 binary.
cd training
python3.12 -m venv .venv # macOS: brew install python@3.12
.venv/bin/pip install -r requirements.txt
On Linux, download the Godot 4.7 Linux binary and point at it:
export GODOT_BIN=~/godot/Godot_v4.7.1-stable_linux.x86_64
(macOS default is /Applications/Godot.app/Contents/MacOS/Godot; override
with GODOT_BIN or --godot_bin if yours lives elsewhere.)
Run a training session
cd training
.venv/bin/python train.py --experiment run01 --timesteps 20000000 --n-parallel 6 --speedup 16
- Checkpoints land in
training/checkpoints/run01/every--checkpoint-everysteps (default 100k), plusfinal.zipon exit (also written on Ctrl-C). - Resume with
--resume checkpoints/run01/final.zip. --n-parallel= Godot instances (2 agents each). Scale with CPU cores.--speedup= in-engine physics speedup. Raise until CPU saturates.--wandbmirrors logs to Weights & Biases (pip install wandbfirst).
Expect the smoke-run scale (~100k steps) to only learn crude ball-chasing; real behaviour needs tens of millions of steps (hours on the 3090 box).
Watch progress
.venv/bin/tensorboard --logdir training/logs
Key curves: rollout/ep_rew_mean (should trend up), rollout/ep_len_mean
(should trend down from 225 as goals end episodes early — 225 action steps
= the 30s episode timeout), rollout/goal_rate (fraction of recent episodes
that ended in an actual goal rather than timing out as a draw — the live
signal for "is the policy actually finishing more episodes by scoring",
since ep_rew_mean mixes that with dense reward-shaping (ball chasing/
touching) and doesn't isolate it).
Reward/observation tuning
Reward weights are exported vars on ShipAIController (goal reward on
TrainingMode) — tune in training.tscn/scripts without touching the
trainer. If you change the observation layout (ship_observations.gd),
old checkpoints/exports become incompatible: retrain, and bump a note in
your experiment name.
Export a checkpoint into the game
cd training
.venv/bin/python export_policy.py checkpoints/run01/final.zip ../Game/bots/hard.json
The exporter runs a parity check (JSON forward pass vs SB3 prediction) before
writing. Models live in Game/bots/.
Evaluate progress between checkpoints
TensorBoard shows learning, but "is the new checkpoint actually better?" needs head-to-head play:
.venv/bin/python export_policy.py checkpoints/run01/ppo_5000000_steps.zip /tmp/candidate.json
.venv/bin/python evaluate.py ../Game/bots/hard.json /tmp/candidate.json --episodes 40
Golden-goal episodes (first goal wins, timeout = draw), sides swapped halfway
for fairness, using the exact inference path that ships in-game. Every run
appends to training/eval_history.json — the long-term progress record.
Evaluate each new candidate against the previous promoted bot and a fixed
early reference to see absolute progress over time.
If a model was trained with the locomotion mask on (curriculum stages 1, 2,
and 5 — see below), pass --grounded-a/--grounded-b for whichever side it's on.
The eval otherwise runs AIShipController fully unmasked regardless of how a
model was trained, so a grounded model's untrained vertical/pitch-roll output
reaches the ship as noise it never had to contend with during training —
this understates it, not a neutral comparison.
Difficulty tiers
A bot is (model, reaction_ticks, action_noise) — configured on the Match
mode (bot_model_path, bot_reaction_ticks, bot_action_noise in
match.tscn) or any AIShipController:
- Model: the main lever. An early checkpoint is an easy bot — promote
e.g.
easy.json/medium.json/hard.jsonfrom different stages of one training run (verify the gaps withevaluate.py). - reaction_ticks (default 8 = training cadence): higher = slower reactions, easier.
- action_noise: adds execution error, easier.
Promoted bots (Game/bots/promoted/)
Game/bots/*.json is a flat, ever-growing dump of every experiment/curriculum
export — useful for evaluate.py and for A/B-ing arbitrary past checkpoints
against each other in the in-game Spectate dropdown (main_menu.gd lists
Game/bots/ non-recursively, so anything one directory deeper is invisible
to it), but none of those filenames (run07.json, curric-s3-no_draws.json,
...) are meant to be the shipped bot — they get superseded constantly and
the automated curriculum pipeline (run_training.sh) only ever writes new
flat files there, never touching subdirectories.
Game/bots/promoted/<tier>.json is the small, curated, hand-maintained set
actually referenced by the shipped game — currently just easy.json
(promoted 2026-07-24 from curric-s6-unmask, the strongest checkpoint at the
time). match.tscn/spectate.tscn point their bot_model_path exports here
directly, so a promoted file is never touched by training scripts, never
overwritten by a same-named future export, and never disturbed by pruning old
experiment files from the flat dump.
To promote a new bot into a tier: copy the chosen Game/bots/<experiment>.json
to Game/bots/promoted/<tier>.json (overwriting the old one), and note the
source experiment + date in this section. Do this for medium.json/
hard.json as later curriculum stages clear the bar against easy.json in
evaluate.py.
Curriculum training
Training from scratch with self-play alone hands the network every skill
at once — finishing, defending, positioning, not stalling to a draw — off a
sparse ±40 goal reward. train.py has a curriculum flag group that stages
this the way you'd coach a human: score first, then also defend, then learn
that a draw is still a failure, and only then spend compute polishing general
movement. Each stage is a normal chained run — a new --experiment resumed
via --resume checkpoints/<previous>/final.zip, same as any other run —
just with different curriculum flags.
curriculum.py has run through two generations so far. Generation 1 (below)
ran stages 1-6 to completion/block and is archived; generation 2 (the one
curriculum.py actually runs today) starts a fresh stage 1 seeded from
generation 1's last clean pass instead of continuing to retry a stage that
kept getting worse — see "Generation 2" below.
Generation 1 (archived — see curriculum_state_gen1.json)
| Stage | Flags | What it teaches |
|---|---|---|
| 1 — score | --opponent-mode inert --attack-goal-bias 1.0 --no-allow-vertical --no-allow-pitch-roll |
Team 1 is a do-nothing placeholder ship parked at its spawn (an effectively empty net); near-goal resets always target the goal the trainee attacks; the ship can't fly or pitch/roll, only drive and yaw. Isolated finishing practice. |
| 2 — defend too | --opponent-mode self_play --no-allow-vertical --no-allow-pitch-roll |
Reintroduces a live opponent (self-play) and the default episode-start mix — the same near-goal state is now simultaneously a finishing chance for one side and a defensive save for the other. Locomotion stays grounded. |
| 3 — no draws | --draw-penalty 5 --reset-std 0.3 |
Training episodes are golden-goal (end at the first goal), so there's no in-episode goal-margin to penalize — draw_penalty is the closest available signal: a one-time penalty when an episode times out with no goal at all, on top of the existing per-tick time_penalty. Also lifts the locomotion mask (full 3D controls) by omitting --allow-vertical/--allow-pitch-roll; pair that with --reset-std since the policy never got a reward gradient on those axes before now, so expect a brief re-exploration wobble. |
| 4 — mechanics/refinement | (no curriculum flags — plain next_run.sh) |
Stock self-play, full controls, default reward/start-state mix. This is what all runs before this feature already did. |
| 5 — aggression | --opponent-mode self_play --no-allow-vertical --no-allow-pitch-roll --velocity-to-ball-weight 0.05 --ball-distance-penalty 0.006 --ball-touch-reward 0.5 |
Resumes from stage 2 (curric-s2-defend), not stage 4 — see the regression note below. Retunes ball-pursuit reward weights (up from 0.02/0.002/0.4) for much more aggressive, constantly-chasing floor play, deliberately keeping the locomotion mask on so it can't reopen the stage-3 regression. Passed 2026-07-22 (41-47 vs grounded stage 2 — close, not yet a clear win). |
| 6 — unmask | --opponent-mode self_play --velocity-to-ball-weight 0.05 --ball-distance-penalty 0.006 --ball-touch-reward 0.5 --airborne-penalty 0.003 |
Re-opens full 3D controls on top of the aggression retune — this is the same grounded-checkpoint-to-full-3D transition that regressed stage 3, but this time paired with airborne_penalty (dense, scaled by height above the floor — see ship_ai_controller.gd) so the policy learns to prefer staying grounded through incentives instead of a hard mask, and can still pick up genuinely useful aerial/wall plays instead of never touching those axes. Failed 3 attempts in a row (25% → 20% → 15% win rate vs curric-s5-aggression) and blocked — see "Generation 2" below for what replaced it. |
Stages 3-4 regressed; stage 6 deliberately reopens the same transition with a mitigation. The locomotion-mask inference bugfix (
8c15c46) revealed that stage 3's evals up to that point had been running with an unfairly unmasked grounded reference. Re-evaluated fairly,curric-s2-defend(grounded) beats bothcurric-s3-no_draws(26-60) andcurric-s4-mechanics(24-57) — lifting the locomotion mask to full 3D in stage 3 was a clear regression in floor play that self-play never earned back in 20M steps. Stage 5 sidesteps this by resuming and evaluating against stage 2 directly (curriculum.py'sresume_from_experiment/reference_experimentstage-dict overrides) instead of chaining through stages 3-4. Stage 6 is where full 3D flight comes back — not masked away this time, but discouraged viaairborne_penaltyand given ~12x the training time to settle. Seecurriculum_state_gen1.json's log for the full eval numbers.
Stage 6 (unmask, retry1, retry2) all used identical flags —
curriculum.py always reuses STAGES[stage_index]["flags"] on retry, only
the resume checkpoint changes — so continued training just drifted the same
policy further rather than converging differently (25% → 20% → 15% win rate
vs curric-s5-aggression). After 3 failed attempts the script blocked for
human review; rather than pile up retry4, retry5, ... on a lineage that
kept getting worse, generation 2 (below) replaces it with a fresh stage 1.
Generation 2 (current)
curriculum.py's live STAGES list now contains a single stage, unmask
(displays as stage 1 — curric-s1-unmask), which picks up exactly where
generation 1's regression analysis left off. It resumes directly from
FOUNDATION_EXPERIMENT (curric-s5-aggression's own checkpoint — the last
stage that passed cleanly) via resume_from_experiment/reference_experiment
overrides, rather than re-running stages 1-5 or continuing generation 1's
drifted retry2:
--opponent-mode self_play --velocity-to-ball-weight 0.08 --ball-distance-penalty 0.01 --ball-touch-reward 0.7 --airborne-penalty 0.003 --ball-velocity-to-goal-weight 0.06 --goal-reward 80 --draw-penalty 5
Compared to generation 1's stage 6:
velocity_to_ball_weight(0.05→0.08) andball_distance_penalty(0.006→0.01) — the actual ball-chasing terms, unchanged since stage 5 despite three failed attempts — plusball_touch_reward(0.5→0.7).- Two scoring-specific terms newly exposed via
train.py(they already existed asship_ai_controller.gd/training_mode.gd@exports, just not as CLI flags):ball_velocity_to_goal_weight(0.004 default → 0.06) rewards the ball actually moving toward the goal, not just being chased/touched;goal_reward(40 default → 80) is the terminal reward for scoring itself. draw_penalty 5(proven effective in stage 3 against passivity), which generation 1's stage 6 had never set — previously all carrot for scoring, no stick for never scoring.reset_retry_checkpoint: Trueon the stage dict, so if this stage itself fails and retries,resume_checkpoint()resets toFOUNDATION_EXPERIMENTagain instead of drifting a failed attempt further — the specific bug that made generation 1's 3 retries monotonically worse instead of converging.
Deliberately not added: a cooldown/cap on ball_velocity_to_goal_weight to
guard against a bot farming near-misses (bouncing the ball toward goal
repeatedly without finishing) instead of actually scoring. Unlike the
touch-farming bug (see ship_ai_controller.gd's ball_touch_reward
comments) this term is already direction-scaled by construction (it's a
velocity-toward-goal quantity, not an undirected contact count), so the risk
is theoretical rather than demonstrated. If this stage's eval shows high
ball_velocity_to_goal_weight accrual without a matching rise in actual
goals scored, that's the signal to add one.
Every experiment name curriculum.py generates is now timestamped
(YYYYMMDD-HHMM-<name>, e.g. 20260727-0930-curric-s1-unmask), applied once
in run_stage_attempt — this keeps generation 2's names from colliding with
generation 1's plain ones (both checkpoint directories and TensorBoard run
names come straight from --experiment) and makes run order obvious in
TensorBoard without cross-referencing curriculum_state.json.
All curriculum flags default to leaving Godot's own @export defaults
alone (train.py only forwards a flag when you pass it), so ordinary runs
are unaffected. Full flag list: --opponent-mode {self_play,inert,frozen},
--opponent-model <path> (for frozen), --draw-penalty,
--attack-goal-bias, --kickoff-chance, --near-goal-chance,
--allow-vertical/--no-allow-vertical, --allow-pitch-roll/--no-allow-pitch-roll,
--velocity-to-ball-weight, --ball-distance-penalty, --ball-touch-reward,
--airborne-penalty, --ball-velocity-to-goal-weight, --goal-reward.
Running it automatically
training/curriculum.py (started via curriculum.sh, same detached-tmux
pattern as start_training.sh) drives all stages end to end: for each
stage it runs run_training.sh (pull, train, export, commit+push) with that
stage's flags, then evaluates the resulting checkpoint against a reference
bot over 100 episodes — the fixed rookie.json baseline for a from-scratch
stage 1 (no resume_from_experiment/reference_experiment override on
STAGES[0]), the previous stage's promoted checkpoint by default for
stages 2+, or an explicit override in that stage's dict when it deliberately
skips a since-regressed branch (generation 1's stage 5) or seeds from a
fixed foundation checkpoint (generation 2's stage 1 — see above).
cd training
./curriculum.sh # start/resume the curriculum
./curriculum.sh --seed-checkpoint checkpoints/run11/final.zip # override stage 1's resume source for this run
The gate is deliberately lenient: it blocks a stage only on a clear
regression (the reference beating the candidate by 15+ points of win
rate), not "must show improvement." A 40-episode eval already misled us once
in this project — run11 was the first model to deliberately score a goal,
but its head-to-head eval read as a loss on sample noise alone. A strict
gate would have retried that stage forever for the wrong reason; a loose one
still catches a genuinely broken stage. Progress and every attempt's eval
result are logged to curriculum_state.json (committed alongside
eval_history.json after each attempt).
A stage gets up to 2 retries (3 attempts total) before the script stops and
asks for a human look — it will not retry indefinitely or advance past a
stage that keeps failing on its own. By default a retry resumes from that
stage's own previous attempt with a fresh --reset-std; a stage can instead
set reset_retry_checkpoint: True (generation 2's stage 1 does) to always
reset to its normal resume source instead — see the generation 1 → 2
postmortem above for why blind same-checkpoint retries can make things
monotonically worse. Once you've looked at why a block happened (more
timesteps? a flag needs adjusting? the eval itself was misleading?), re-run
with --force-retry to try again or --skip-to-next-stage if you judge the
result good enough despite the gate.
Running a stage by hand (e.g. to experiment with flags before trusting the
orchestrator) still works exactly as the table above describes — just call
next_run.sh/run_training.sh directly with that stage's flags.
Self-play notes
By default both ships share the live policy (mirrored, team-relative
observations — see ship_observations.gd), so training is against the
current self. --opponent-mode inert/frozen (see Curriculum training
above) replace that with a placeholder or a fixed exported policy for one
side of a run; frozen is a single-fixed-model slice of full league play.
Fixed-opponent training against a pool of past checkpoints sampled per
episode (to avoid strategy collapse on long self-play runs) is still
deferred — see TODO.md.