AIShipController (eval + real gameplay) ran the raw policy output unmasked regardless of allow_vertical/allow_pitch_roll, while ShipAIController (training) correctly discarded those axes for grounded curriculum stages. A grounded-trained model's untrained vertical/pitch-roll output reached the ship as noise during eval, understating it against models that were never handicapped this way.
11 KiB
Training the AI bot
Cosmic Clash bots are trained with reinforcement learning (self-play PPO): two ships in a headless arena share one policy that learns by playing against itself. Training runs in Python (Godot RL Agents bridge + Stable-Baselines3); the trained policy is exported to a small JSON file and runs inside the game in pure GDScript — shipped bots need no Python, no .NET, no network.
How it fits together
Game/scenes/training.tscn+scripts/training_mode.gd— headless self-play environment: two RL ships, randomized episode starts, goal rewards. Contains the vendored godot_rl_agentsSyncnode that talks TCP to the trainer.scripts/ship_ai_controller.gd— training-side bridge (observations, rewards, action mapping).scripts/ship_observations.gdis the shared observation builder — training and in-game inference must stay identical, so never fork it.training/train.py— PPO trainer; launches N parallel headless Godot instances (2 agents each).training/export_policy.py— SB3 checkpoint → JSON policy for the game.scripts/ai_ship_controller.gd+scripts/policy_network.gd— in-game inference (GDScript MLP forward pass).training/evaluate.py— pits two exported policies against each other and appends totraining/eval_history.json.
Hardware
The environment is our own headless Godot sim — fully cross-platform:
- Any machine (e.g. the M4 Mac mini): fine for pipeline development, smoke runs, and short experiments. Env stepping is CPU-bound; the policy is a small MLP, so even CPU-only PPO updates are cheap.
- Linux + NVIDIA GPU (e.g. the RTX 3090 box): recommended for real
multi-hour/overnight runs. PyTorch CUDA works out of the box; more CPU
cores also mean more parallel Godot instances (
--n-parallel).
There is no hard GPU requirement (unlike Rocket League tooling) — a GPU mainly speeds up learning updates on long runs.
For the Linux/3090 remote-training workflow (setup, throughput tuning, auto-copying results back to the Mac, dashboard over the network), see TRAINING_LINUX.md.
Setup
Needs Python 3.10+ and a Godot 4.7 binary.
cd training
python3.12 -m venv .venv # macOS: brew install python@3.12
.venv/bin/pip install -r requirements.txt
On Linux, download the Godot 4.7 Linux binary and point at it:
export GODOT_BIN=~/godot/Godot_v4.7.1-stable_linux.x86_64
(macOS default is /Applications/Godot.app/Contents/MacOS/Godot; override
with GODOT_BIN or --godot_bin if yours lives elsewhere.)
Run a training session
cd training
.venv/bin/python train.py --experiment run01 --timesteps 20000000 --n-parallel 6 --speedup 16
- Checkpoints land in
training/checkpoints/run01/every--checkpoint-everysteps (default 100k), plusfinal.zipon exit (also written on Ctrl-C). - Resume with
--resume checkpoints/run01/final.zip. --n-parallel= Godot instances (2 agents each). Scale with CPU cores.--speedup= in-engine physics speedup. Raise until CPU saturates.--wandbmirrors logs to Weights & Biases (pip install wandbfirst).
Expect the smoke-run scale (~100k steps) to only learn crude ball-chasing; real behaviour needs tens of millions of steps (hours on the 3090 box).
Watch progress
.venv/bin/tensorboard --logdir training/logs
Key curves: rollout/ep_rew_mean (should trend up), rollout/ep_len_mean
(should trend down from 225 as goals end episodes early — 225 action steps
= the 30s episode timeout).
Reward/observation tuning
Reward weights are exported vars on ShipAIController (goal reward on
TrainingMode) — tune in training.tscn/scripts without touching the
trainer. If you change the observation layout (ship_observations.gd),
old checkpoints/exports become incompatible: retrain, and bump a note in
your experiment name.
Export a checkpoint into the game
cd training
.venv/bin/python export_policy.py checkpoints/run01/final.zip ../Game/bots/hard.json
The exporter runs a parity check (JSON forward pass vs SB3 prediction) before
writing. Models live in Game/bots/.
Evaluate progress between checkpoints
TensorBoard shows learning, but "is the new checkpoint actually better?" needs head-to-head play:
.venv/bin/python export_policy.py checkpoints/run01/ppo_5000000_steps.zip /tmp/candidate.json
.venv/bin/python evaluate.py ../Game/bots/hard.json /tmp/candidate.json --episodes 40
Golden-goal episodes (first goal wins, timeout = draw), sides swapped halfway
for fairness, using the exact inference path that ships in-game. Every run
appends to training/eval_history.json — the long-term progress record.
Evaluate each new candidate against the previous promoted bot and a fixed
early reference to see absolute progress over time.
If a model was trained with the locomotion mask on (curriculum stages 1-2 —
see below), pass --grounded-a/--grounded-b for whichever side it's on.
The eval otherwise runs AIShipController fully unmasked regardless of how a
model was trained, so a grounded model's untrained vertical/pitch-roll output
reaches the ship as noise it never had to contend with during training —
this understates it, not a neutral comparison.
Difficulty tiers
A bot is (model, reaction_ticks, action_noise) — configured on the Match
mode (bot_model_path, bot_reaction_ticks, bot_action_noise in
match.tscn) or any AIShipController:
- Model: the main lever. An early checkpoint is an easy bot — promote
e.g.
easy.json/medium.json/hard.jsonfrom different stages of one training run (verify the gaps withevaluate.py). - reaction_ticks (default 8 = training cadence): higher = slower reactions, easier.
- action_noise: adds execution error, easier.
Curriculum training
Training from scratch with self-play alone hands the network every skill
at once — finishing, defending, positioning, not stalling to a draw — off a
sparse ±40 goal reward. train.py has a curriculum flag group that stages
this the way you'd coach a human: score first, then also defend, then learn
that a draw is still a failure, and only then spend compute polishing general
movement. Each stage is a normal chained run — a new --experiment resumed
via --resume checkpoints/<previous>/final.zip, same as any other run —
just with different curriculum flags.
| Stage | Flags | What it teaches |
|---|---|---|
| 1 — score | --opponent-mode inert --attack-goal-bias 1.0 --no-allow-vertical --no-allow-pitch-roll |
Team 1 is a do-nothing placeholder ship parked at its spawn (an effectively empty net); near-goal resets always target the goal the trainee attacks; the ship can't fly or pitch/roll, only drive and yaw. Isolated finishing practice. |
| 2 — defend too | --opponent-mode self_play --no-allow-vertical --no-allow-pitch-roll |
Reintroduces a live opponent (self-play) and the default episode-start mix — the same near-goal state is now simultaneously a finishing chance for one side and a defensive save for the other. Locomotion stays grounded. |
| 3 — no draws | --draw-penalty 5 --reset-std 0.3 |
Training episodes are golden-goal (end at the first goal), so there's no in-episode goal-margin to penalize — draw_penalty is the closest available signal: a one-time penalty when an episode times out with no goal at all, on top of the existing per-tick time_penalty. Also lifts the locomotion mask (full 3D controls) by omitting --allow-vertical/--allow-pitch-roll; pair that with --reset-std since the policy never got a reward gradient on those axes before now, so expect a brief re-exploration wobble. |
| 4 — mechanics/refinement | (no curriculum flags — plain next_run.sh) |
Stock self-play, full controls, default reward/start-state mix. This is what all runs before this feature already did. |
All curriculum flags default to leaving Godot's own @export defaults
alone (train.py only forwards a flag when you pass it), so ordinary runs
are unaffected. Full flag list: --opponent-mode {self_play,inert,frozen},
--opponent-model <path> (for frozen), --draw-penalty,
--attack-goal-bias, --kickoff-chance, --near-goal-chance,
--allow-vertical/--no-allow-vertical, --allow-pitch-roll/--no-allow-pitch-roll.
Running it automatically
training/curriculum.py (started via curriculum.sh, same detached-tmux
pattern as start_training.sh) drives all four stages end to end: for each
stage it runs run_training.sh (pull, train, export, commit+push) with that
stage's flags, then evaluates the resulting checkpoint against a reference
bot — the fixed rookie.json baseline for stage 1, or the previous stage's
promoted checkpoint for stages 2-4 — over 100 episodes.
cd training
./curriculum.sh # start/resume the curriculum
./curriculum.sh --seed-checkpoint checkpoints/run11/final.zip # seed stage 1 instead of a fresh policy
The gate is deliberately lenient: it blocks a stage only on a clear
regression (the reference beating the candidate by 15+ points of win
rate), not "must show improvement." A 40-episode eval already misled us once
in this project — run11 was the first model to deliberately score a goal,
but its head-to-head eval read as a loss on sample noise alone. A strict
gate would have retried that stage forever for the wrong reason; a loose one
still catches a genuinely broken stage. Progress and every attempt's eval
result are logged to curriculum_state.json (committed alongside
eval_history.json after each attempt).
A stage gets up to 2 retries (3 attempts total, each resuming from that
stage's own previous attempt with a fresh --reset-std) before the script
stops and asks for a human look — it will not retry indefinitely or advance
past a stage that keeps failing on its own. Once you've looked at why (more
timesteps? a flag needs adjusting? the eval itself was misleading?), re-run
with --force-retry to try again or --skip-to-next-stage if you judge the
result good enough despite the gate.
Running a stage by hand (e.g. to experiment with flags before trusting the
orchestrator) still works exactly as the table above describes — just call
next_run.sh/run_training.sh directly with that stage's flags.
Self-play notes
By default both ships share the live policy (mirrored, team-relative
observations — see ship_observations.gd), so training is against the
current self. --opponent-mode inert/frozen (see Curriculum training
above) replace that with a placeholder or a fixed exported policy for one
side of a run; frozen is a single-fixed-model slice of full league play.
Fixed-opponent training against a pool of past checkpoints sampled per
episode (to avoid strategy collapse on long self-play runs) is still
deferred — see TODO.md.