10 KiB
Training the AI bot
Cosmic Clash bots are trained with reinforcement learning (self-play PPO): two ships in a headless arena share one policy that learns by playing against itself. Training runs in Python (Godot RL Agents bridge + Stable-Baselines3); the trained policy is exported to a small JSON file and runs inside the game in pure GDScript — shipped bots need no Python, no .NET, no network.
How it fits together
Game/scenes/training.tscn+scripts/training_mode.gd— headless self-play environment: two RL ships, randomized episode starts, goal rewards. Contains the vendored godot_rl_agentsSyncnode that talks TCP to the trainer.scripts/ship_ai_controller.gd— training-side bridge (observations, rewards, action mapping).scripts/ship_observations.gdis the shared observation builder — training and in-game inference must stay identical, so never fork it.training/train.py— PPO trainer; launches N parallel headless Godot instances (2 agents each).training/export_policy.py— SB3 checkpoint → JSON policy for the game.scripts/ai_ship_controller.gd+scripts/policy_network.gd— in-game inference (GDScript MLP forward pass).training/evaluate.py— pits two exported policies against each other and appends totraining/eval_history.json.
Hardware
The environment is our own headless Godot sim — fully cross-platform:
- Any machine (e.g. the M4 Mac mini): fine for pipeline development, smoke runs, and short experiments. Env stepping is CPU-bound; the policy is a small MLP, so even CPU-only PPO updates are cheap.
- Linux + NVIDIA GPU (e.g. the RTX 3090 box): recommended for real
multi-hour/overnight runs. PyTorch CUDA works out of the box; more CPU
cores also mean more parallel Godot instances (
--n-parallel).
There is no hard GPU requirement (unlike Rocket League tooling) — a GPU mainly speeds up learning updates on long runs.
For the Linux/3090 remote-training workflow (setup, throughput tuning, auto-copying results back to the Mac, dashboard over the network), see TRAINING_LINUX.md.
Setup
Needs Python 3.10+ and a Godot 4.7 binary.
cd training
python3.12 -m venv .venv # macOS: brew install python@3.12
.venv/bin/pip install -r requirements.txt
On Linux, download the Godot 4.7 Linux binary and point at it:
export GODOT_BIN=~/godot/Godot_v4.7.1-stable_linux.x86_64
(macOS default is /Applications/Godot.app/Contents/MacOS/Godot; override
with GODOT_BIN or --godot_bin if yours lives elsewhere.)
Run a training session
cd training
.venv/bin/python train.py --experiment run01 --timesteps 20000000 --n-parallel 6 --speedup 16
- Checkpoints land in
training/checkpoints/run01/every--checkpoint-everysteps (default 100k), plusfinal.zipon exit (also written on Ctrl-C). - Resume with
--resume checkpoints/run01/final.zip. --n-parallel= Godot instances (2 agents each). Scale with CPU cores.--speedup= in-engine physics speedup. Raise until CPU saturates.--wandbmirrors logs to Weights & Biases (pip install wandbfirst).
Expect the smoke-run scale (~100k steps) to only learn crude ball-chasing; real behaviour needs tens of millions of steps (hours on the 3090 box).
Watch progress
.venv/bin/tensorboard --logdir training/logs
Key curves: rollout/ep_rew_mean (should trend up), rollout/ep_len_mean
(should trend down from 225 as goals end episodes early — 225 action steps
= the 30s episode timeout).
Reward/observation tuning
Reward weights are exported vars on ShipAIController (goal reward on
TrainingMode) — tune in training.tscn/scripts without touching the
trainer. If you change the observation layout (ship_observations.gd),
old checkpoints/exports become incompatible: retrain, and bump a note in
your experiment name.
Export a checkpoint into the game
cd training
.venv/bin/python export_policy.py checkpoints/run01/final.zip ../Game/bots/hard.json
The exporter runs a parity check (JSON forward pass vs SB3 prediction) before
writing. Models live in Game/bots/.
Evaluate progress between checkpoints
TensorBoard shows learning, but "is the new checkpoint actually better?" needs head-to-head play:
.venv/bin/python export_policy.py checkpoints/run01/ppo_5000000_steps.zip /tmp/candidate.json
.venv/bin/python evaluate.py ../Game/bots/hard.json /tmp/candidate.json --episodes 40
Golden-goal episodes (first goal wins, timeout = draw), sides swapped halfway
for fairness, using the exact inference path that ships in-game. Every run
appends to training/eval_history.json — the long-term progress record.
Evaluate each new candidate against the previous promoted bot and a fixed
early reference to see absolute progress over time.
Difficulty tiers
A bot is (model, reaction_ticks, action_noise) — configured on the Match
mode (bot_model_path, bot_reaction_ticks, bot_action_noise in
match.tscn) or any AIShipController:
- Model: the main lever. An early checkpoint is an easy bot — promote
e.g.
easy.json/medium.json/hard.jsonfrom different stages of one training run (verify the gaps withevaluate.py). - reaction_ticks (default 8 = training cadence): higher = slower reactions, easier.
- action_noise: adds execution error, easier.
Curriculum training
Training from scratch with self-play alone hands the network every skill
at once — finishing, defending, positioning, not stalling to a draw — off a
sparse ±40 goal reward. train.py has a curriculum flag group that stages
this the way you'd coach a human: score first, then also defend, then learn
that a draw is still a failure, and only then spend compute polishing general
movement. Each stage is a normal chained run — a new --experiment resumed
via --resume checkpoints/<previous>/final.zip, same as any other run —
just with different curriculum flags.
| Stage | Flags | What it teaches |
|---|---|---|
| 1 — score | --opponent-mode inert --attack-goal-bias 1.0 --no-allow-vertical --no-allow-pitch-roll |
Team 1 is a do-nothing placeholder ship parked at its spawn (an effectively empty net); near-goal resets always target the goal the trainee attacks; the ship can't fly or pitch/roll, only drive and yaw. Isolated finishing practice. |
| 2 — defend too | --opponent-mode self_play --no-allow-vertical --no-allow-pitch-roll |
Reintroduces a live opponent (self-play) and the default episode-start mix — the same near-goal state is now simultaneously a finishing chance for one side and a defensive save for the other. Locomotion stays grounded. |
| 3 — no draws | --draw-penalty 5 --reset-std 0.3 |
Training episodes are golden-goal (end at the first goal), so there's no in-episode goal-margin to penalize — draw_penalty is the closest available signal: a one-time penalty when an episode times out with no goal at all, on top of the existing per-tick time_penalty. Also lifts the locomotion mask (full 3D controls) by omitting --allow-vertical/--allow-pitch-roll; pair that with --reset-std since the policy never got a reward gradient on those axes before now, so expect a brief re-exploration wobble. |
| 4 — mechanics/refinement | (no curriculum flags — plain next_run.sh) |
Stock self-play, full controls, default reward/start-state mix. This is what all runs before this feature already did. |
All curriculum flags default to leaving Godot's own @export defaults
alone (train.py only forwards a flag when you pass it), so ordinary runs
are unaffected. Full flag list: --opponent-mode {self_play,inert,frozen},
--opponent-model <path> (for frozen), --draw-penalty,
--attack-goal-bias, --kickoff-chance, --near-goal-chance,
--allow-vertical/--no-allow-vertical, --allow-pitch-roll/--no-allow-pitch-roll.
Running it automatically
training/curriculum.py (started via curriculum.sh, same detached-tmux
pattern as start_training.sh) drives all four stages end to end: for each
stage it runs run_training.sh (pull, train, export, commit+push) with that
stage's flags, then evaluates the resulting checkpoint against a reference
bot — the fixed rookie.json baseline for stage 1, or the previous stage's
promoted checkpoint for stages 2-4 — over 100 episodes.
cd training
./curriculum.sh # start/resume the curriculum
./curriculum.sh --seed-checkpoint checkpoints/run11/final.zip # seed stage 1 instead of a fresh policy
The gate is deliberately lenient: it blocks a stage only on a clear
regression (the reference beating the candidate by 15+ points of win
rate), not "must show improvement." A 40-episode eval already misled us once
in this project — run11 was the first model to deliberately score a goal,
but its head-to-head eval read as a loss on sample noise alone. A strict
gate would have retried that stage forever for the wrong reason; a loose one
still catches a genuinely broken stage. Progress and every attempt's eval
result are logged to curriculum_state.json (committed alongside
eval_history.json after each attempt).
A stage gets up to 2 retries (3 attempts total, each resuming from that
stage's own previous attempt with a fresh --reset-std) before the script
stops and asks for a human look — it will not retry indefinitely or advance
past a stage that keeps failing on its own. Once you've looked at why (more
timesteps? a flag needs adjusting? the eval itself was misleading?), re-run
with --force-retry to try again or --skip-to-next-stage if you judge the
result good enough despite the gate.
Running a stage by hand (e.g. to experiment with flags before trusting the
orchestrator) still works exactly as the table above describes — just call
next_run.sh/run_training.sh directly with that stage's flags.
Self-play notes
By default both ships share the live policy (mirrored, team-relative
observations — see ship_observations.gd), so training is against the
current self. --opponent-mode inert/frozen (see Curriculum training
above) replace that with a placeholder or a fixed exported policy for one
side of a run; frozen is a single-fixed-model slice of full league play.
Fixed-opponent training against a pool of past checkpoints sampled per
episode (to avoid strategy collapse on long self-play runs) is still
deferred — see TODO.md.