mirror of
https://github.com/jcreek/CosmicClash.git
synced 2026-09-10 16:04:04 +00:00
334 lines
19 KiB
Markdown
334 lines
19 KiB
Markdown
# Training the AI bot
|
|
|
|
Cosmic Clash bots are trained with reinforcement learning (self-play PPO): two
|
|
ships in a headless arena share one policy that learns by playing against
|
|
itself. Training runs in Python ([Godot RL Agents](https://github.com/edbeeching/godot_rl_agents)
|
|
bridge + Stable-Baselines3); the trained policy is exported to a small JSON
|
|
file and runs **inside the game** in pure GDScript — shipped bots need no
|
|
Python, no .NET, no network.
|
|
|
|
## How it fits together
|
|
|
|
- `Game/scenes/training.tscn` + `scripts/training_mode.gd` — headless self-play
|
|
environment: two RL ships, randomized episode starts, goal rewards. Contains
|
|
the vendored godot_rl_agents `Sync` node that talks TCP to the trainer.
|
|
- `scripts/ship_ai_controller.gd` — training-side bridge (observations,
|
|
rewards, action mapping). `scripts/ship_observations.gd` is the *shared*
|
|
observation builder — training and in-game inference must stay identical,
|
|
so never fork it.
|
|
- `training/train.py` — PPO trainer; launches N parallel headless Godot
|
|
instances (2 agents each) — from source by default, or from a pre-built
|
|
binary via `--exported-binary` (see TRAINING_LINUX.md's "Exported-binary
|
|
training" section; `training/export_linux.sh` builds it from
|
|
`Game/export_presets.cfg`'s "Linux Training" preset).
|
|
- `training/export_policy.py` — SB3 checkpoint → JSON policy for the game.
|
|
- `scripts/ai_ship_controller.gd` + `scripts/policy_network.gd` — in-game
|
|
inference (GDScript MLP forward pass).
|
|
- `training/evaluate.py` — pits two exported policies against each other and
|
|
appends to `training/eval_history.json`.
|
|
|
|
## Hardware
|
|
|
|
The environment is our own headless Godot sim — fully cross-platform:
|
|
|
|
- **Any machine (e.g. the M4 Mac mini)**: fine for pipeline development,
|
|
smoke runs, and short experiments. Env stepping is CPU-bound; the policy is
|
|
a small MLP, so even CPU-only PPO updates are cheap.
|
|
- **Linux + NVIDIA GPU (e.g. the RTX 3090 box)**: recommended for real
|
|
multi-hour/overnight runs. PyTorch CUDA works out of the box; more CPU
|
|
cores also mean more parallel Godot instances (`--n-parallel`).
|
|
|
|
There is no hard GPU requirement (unlike Rocket League tooling) — a GPU
|
|
mainly speeds up learning updates on long runs.
|
|
|
|
For the Linux/3090 remote-training workflow (setup, throughput tuning,
|
|
auto-copying results back to the Mac, dashboard over the network), see
|
|
[TRAINING_LINUX.md](TRAINING_LINUX.md).
|
|
|
|
## Setup
|
|
|
|
Needs Python 3.10+ and a Godot 4.7 binary.
|
|
|
|
```bash
|
|
cd training
|
|
python3.12 -m venv .venv # macOS: brew install python@3.12
|
|
.venv/bin/pip install -r requirements.txt
|
|
```
|
|
|
|
On Linux, download the Godot 4.7 Linux binary and point at it:
|
|
|
|
```bash
|
|
export GODOT_BIN=~/godot/Godot_v4.7.1-stable_linux.x86_64
|
|
```
|
|
|
|
(macOS default is `/Applications/Godot.app/Contents/MacOS/Godot`; override
|
|
with `GODOT_BIN` or `--godot_bin` if yours lives elsewhere.)
|
|
|
|
## Run a training session
|
|
|
|
```bash
|
|
cd training
|
|
.venv/bin/python train.py --experiment run01 --timesteps 20000000 --n-parallel 6 --speedup 16
|
|
```
|
|
|
|
- Checkpoints land in `training/checkpoints/run01/` every `--checkpoint-every`
|
|
steps (default 100k), plus `final.zip` on exit (also written on Ctrl-C).
|
|
- Resume with `--resume checkpoints/run01/final.zip`.
|
|
- `--n-parallel` = Godot instances (2 agents each). Scale with CPU cores.
|
|
- `--speedup` = in-engine physics speedup. Raise until CPU saturates.
|
|
- `--wandb` mirrors logs to Weights & Biases (`pip install wandb` first).
|
|
|
|
Expect the smoke-run scale (~100k steps) to only learn crude ball-chasing;
|
|
real behaviour needs tens of millions of steps (hours on the 3090 box).
|
|
|
|
### Watch progress
|
|
|
|
```bash
|
|
.venv/bin/tensorboard --logdir training/logs
|
|
```
|
|
|
|
Key curves: `rollout/ep_rew_mean` (should trend up), `rollout/ep_len_mean`
|
|
(should trend *down* from 225 as goals end episodes early — 225 action steps
|
|
= the 30s episode timeout).
|
|
|
|
### Reward/observation tuning
|
|
|
|
Reward weights are exported vars on `ShipAIController` (goal reward on
|
|
`TrainingMode`) — tune in `training.tscn`/scripts without touching the
|
|
trainer. If you change the *observation* layout (`ship_observations.gd`),
|
|
old checkpoints/exports become incompatible: retrain, and bump a note in
|
|
your experiment name.
|
|
|
|
## Export a checkpoint into the game
|
|
|
|
```bash
|
|
cd training
|
|
.venv/bin/python export_policy.py checkpoints/run01/final.zip ../Game/bots/hard.json
|
|
```
|
|
|
|
The exporter runs a parity check (JSON forward pass vs SB3 prediction) before
|
|
writing. Models live in `Game/bots/`.
|
|
|
|
## Evaluate progress between checkpoints
|
|
|
|
TensorBoard shows learning, but "is the new checkpoint actually *better*?"
|
|
needs head-to-head play:
|
|
|
|
```bash
|
|
.venv/bin/python export_policy.py checkpoints/run01/ppo_5000000_steps.zip /tmp/candidate.json
|
|
.venv/bin/python evaluate.py ../Game/bots/hard.json /tmp/candidate.json --episodes 40
|
|
```
|
|
|
|
Golden-goal episodes (first goal wins, timeout = draw), sides swapped halfway
|
|
for fairness, using the exact inference path that ships in-game. Every run
|
|
appends to `training/eval_history.json` — the long-term progress record.
|
|
Evaluate each new candidate against the previous promoted bot and a fixed
|
|
early reference to see absolute progress over time.
|
|
|
|
If a model was trained with the locomotion mask on (curriculum stages 1, 2,
|
|
and 5 — see below), pass `--grounded-a`/`--grounded-b` for whichever side it's on.
|
|
The eval otherwise runs `AIShipController` fully unmasked regardless of how a
|
|
model was trained, so a grounded model's untrained vertical/pitch-roll output
|
|
reaches the ship as noise it never had to contend with during training —
|
|
this understates it, not a neutral comparison.
|
|
|
|
## Difficulty tiers
|
|
|
|
A bot is `(model, reaction_ticks, action_noise)` — configured on the Match
|
|
mode (`bot_model_path`, `bot_reaction_ticks`, `bot_action_noise` in
|
|
`match.tscn`) or any `AIShipController`:
|
|
|
|
- **Model**: the main lever. An early checkpoint *is* an easy bot — promote
|
|
e.g. `easy.json` / `medium.json` / `hard.json` from different stages of one
|
|
training run (verify the gaps with `evaluate.py`).
|
|
- **reaction_ticks** (default 8 = training cadence): higher = slower
|
|
reactions, easier.
|
|
- **action_noise**: adds execution error, easier.
|
|
|
|
### Promoted bots (`Game/bots/promoted/`)
|
|
|
|
`Game/bots/*.json` is a flat, ever-growing dump of every experiment/curriculum
|
|
export — useful for `evaluate.py` and for A/B-ing arbitrary past checkpoints
|
|
against each other in the in-game Spectate dropdown (`main_menu.gd` lists
|
|
`Game/bots/` non-recursively, so anything one directory deeper is invisible
|
|
to it), but none of those filenames (`run07.json`, `curric-s3-no_draws.json`,
|
|
...) are meant to be *the* shipped bot — they get superseded constantly and
|
|
the automated curriculum pipeline (`run_training.sh`) only ever writes new
|
|
flat files there, never touching subdirectories.
|
|
|
|
`Game/bots/promoted/<tier>.json` is the small, curated, hand-maintained set
|
|
actually referenced by the shipped game — currently just `easy.json`
|
|
(promoted 2026-07-24 from `curric-s6-unmask`, the strongest checkpoint at the
|
|
time). `match.tscn`/`spectate.tscn` point their `bot_model_path` exports here
|
|
directly, so a promoted file is never touched by training scripts, never
|
|
overwritten by a same-named future export, and never disturbed by pruning old
|
|
experiment files from the flat dump.
|
|
|
|
To promote a new bot into a tier: copy the chosen `Game/bots/<experiment>.json`
|
|
to `Game/bots/promoted/<tier>.json` (overwriting the old one), and note the
|
|
source experiment + date in this section. Do this for `medium.json`/
|
|
`hard.json` as later curriculum stages clear the bar against `easy.json` in
|
|
`evaluate.py`.
|
|
|
|
## Curriculum training
|
|
|
|
Training from scratch with self-play alone hands the network every skill
|
|
at once — finishing, defending, positioning, not stalling to a draw — off a
|
|
sparse ±40 goal reward. `train.py` has a `curriculum` flag group that stages
|
|
this the way you'd coach a human: score first, then also defend, then learn
|
|
that a draw is still a failure, and only then spend compute polishing general
|
|
movement. Each stage is a normal chained run — a new `--experiment` resumed
|
|
via `--resume checkpoints/<previous>/final.zip`, same as any other run —
|
|
just with different curriculum flags.
|
|
|
|
`curriculum.py` has run through two generations so far. Generation 1 (below)
|
|
ran stages 1-6 to completion/block and is archived; generation 2 (the one
|
|
`curriculum.py` actually runs today) starts a fresh stage 1 seeded from
|
|
generation 1's last clean pass instead of continuing to retry a stage that
|
|
kept getting worse — see "Generation 2" below.
|
|
|
|
### Generation 1 (archived — see `curriculum_state_gen1.json`)
|
|
|
|
| Stage | Flags | What it teaches |
|
|
|---|---|---|
|
|
| 1 — score | `--opponent-mode inert --attack-goal-bias 1.0 --no-allow-vertical --no-allow-pitch-roll` | Team 1 is a do-nothing placeholder ship parked at its spawn (an effectively empty net); near-goal resets always target the goal the trainee attacks; the ship can't fly or pitch/roll, only drive and yaw. Isolated finishing practice. |
|
|
| 2 — defend too | `--opponent-mode self_play --no-allow-vertical --no-allow-pitch-roll` | Reintroduces a live opponent (self-play) and the default episode-start mix — the same near-goal state is now simultaneously a finishing chance for one side and a defensive save for the other. Locomotion stays grounded. |
|
|
| 3 — no draws | `--draw-penalty 5 --reset-std 0.3` | Training episodes are golden-goal (end at the *first* goal), so there's no in-episode goal-margin to penalize — `draw_penalty` is the closest available signal: a one-time penalty when an episode times out with no goal at all, on top of the existing per-tick `time_penalty`. Also lifts the locomotion mask (full 3D controls) by omitting `--allow-vertical`/`--allow-pitch-roll`; pair that with `--reset-std` since the policy never got a reward gradient on those axes before now, so expect a brief re-exploration wobble. |
|
|
| 4 — mechanics/refinement | *(no curriculum flags — plain `next_run.sh`)* | Stock self-play, full controls, default reward/start-state mix. This is what all runs before this feature already did. |
|
|
| 5 — aggression | `--opponent-mode self_play --no-allow-vertical --no-allow-pitch-roll --velocity-to-ball-weight 0.05 --ball-distance-penalty 0.006 --ball-touch-reward 0.5` | **Resumes from stage 2 (`curric-s2-defend`), not stage 4** — see the regression note below. Retunes ball-pursuit reward weights (up from 0.02/0.002/0.4) for much more aggressive, constantly-chasing floor play, deliberately keeping the locomotion mask on so it can't reopen the stage-3 regression. Passed 2026-07-22 (41-47 vs grounded stage 2 — close, not yet a clear win). |
|
|
| 6 — unmask | `--opponent-mode self_play --velocity-to-ball-weight 0.05 --ball-distance-penalty 0.006 --ball-touch-reward 0.5 --airborne-penalty 0.003` | Re-opens full 3D controls on top of the aggression retune — this is the same grounded-checkpoint-to-full-3D transition that regressed stage 3, but this time paired with `airborne_penalty` (dense, scaled by height above the floor — see `ship_ai_controller.gd`) so the policy learns to *prefer* staying grounded through incentives instead of a hard mask, and can still pick up genuinely useful aerial/wall plays instead of never touching those axes. **Failed 3 attempts in a row** (25% → 20% → 15% win rate vs `curric-s5-aggression`) and blocked — see "Generation 2" below for what replaced it. |
|
|
|
|
> **Stages 3-4 regressed; stage 6 deliberately reopens the same transition
|
|
> with a mitigation.** The locomotion-mask inference bugfix (`8c15c46`)
|
|
> revealed that stage 3's evals up to that point had been running with an
|
|
> unfairly unmasked grounded reference. Re-evaluated fairly, `curric-s2-defend`
|
|
> (grounded) beats both `curric-s3-no_draws` (26-60) and `curric-s4-mechanics`
|
|
> (24-57) — lifting the locomotion mask to full 3D in stage 3 was a clear
|
|
> regression in floor play that self-play never earned back in 20M steps.
|
|
> Stage 5 sidesteps this by resuming and evaluating against stage 2 directly
|
|
> (`curriculum.py`'s `resume_from_experiment`/`reference_experiment` stage-dict
|
|
> overrides) instead of chaining through stages 3-4. Stage 6 is where full 3D
|
|
> flight comes back — not masked away this time, but discouraged via
|
|
> `airborne_penalty` and given ~12x the training time to settle. See
|
|
> `curriculum_state_gen1.json`'s log for the full eval numbers.
|
|
|
|
Stage 6 (`unmask`, `retry1`, `retry2`) all used identical flags —
|
|
`curriculum.py` always reuses `STAGES[stage_index]["flags"]` on retry, only
|
|
the resume checkpoint changes — so continued training just drifted the same
|
|
policy further rather than converging differently (25% → 20% → 15% win rate
|
|
vs `curric-s5-aggression`). After 3 failed attempts the script blocked for
|
|
human review; rather than pile up `retry4`, `retry5`, ... on a lineage that
|
|
kept getting worse, generation 2 (below) replaces it with a fresh stage 1.
|
|
|
|
### Generation 2 (current)
|
|
|
|
`curriculum.py`'s live `STAGES` list now contains a single stage, `unmask`
|
|
(displays as stage 1 — `curric-s1-unmask`), which picks up exactly where
|
|
generation 1's regression analysis left off. It resumes directly from
|
|
`FOUNDATION_EXPERIMENT` (`curric-s5-aggression`'s own checkpoint — the last
|
|
stage that passed cleanly) via `resume_from_experiment`/`reference_experiment`
|
|
overrides, rather than re-running stages 1-5 or continuing generation 1's
|
|
drifted `retry2`:
|
|
|
|
`--opponent-mode self_play --velocity-to-ball-weight 0.08 --ball-distance-penalty 0.01 --ball-touch-reward 0.7 --airborne-penalty 0.003 --ball-velocity-to-goal-weight 0.06 --goal-reward 80 --draw-penalty 5`
|
|
|
|
Compared to generation 1's stage 6:
|
|
- `velocity_to_ball_weight` (0.05→0.08) and `ball_distance_penalty`
|
|
(0.006→0.01) — the actual ball-chasing terms, unchanged since stage 5
|
|
despite three failed attempts — plus `ball_touch_reward` (0.5→0.7).
|
|
- Two scoring-specific terms newly exposed via `train.py` (they already
|
|
existed as `ship_ai_controller.gd`/`training_mode.gd` `@export`s, just not
|
|
as CLI flags): `ball_velocity_to_goal_weight` (0.004 default → 0.06)
|
|
rewards the ball actually moving toward the goal, not just being
|
|
chased/touched; `goal_reward` (40 default → 80) is the terminal reward for
|
|
scoring itself.
|
|
- `draw_penalty 5` (proven effective in stage 3 against passivity), which
|
|
generation 1's stage 6 had never set — previously all carrot for scoring,
|
|
no stick for never scoring.
|
|
- `reset_retry_checkpoint: True` on the stage dict, so if this stage itself
|
|
fails and retries, `resume_checkpoint()` resets to `FOUNDATION_EXPERIMENT`
|
|
again instead of drifting a failed attempt further — the specific bug that
|
|
made generation 1's 3 retries monotonically worse instead of converging.
|
|
|
|
Deliberately not added: a cooldown/cap on `ball_velocity_to_goal_weight` to
|
|
guard against a bot farming near-misses (bouncing the ball toward goal
|
|
repeatedly without finishing) instead of actually scoring. Unlike the
|
|
touch-farming bug (see `ship_ai_controller.gd`'s `ball_touch_reward`
|
|
comments) this term is already direction-scaled by construction (it's a
|
|
velocity-toward-goal quantity, not an undirected contact count), so the risk
|
|
is theoretical rather than demonstrated. If this stage's eval shows high
|
|
`ball_velocity_to_goal_weight` accrual without a matching rise in actual
|
|
goals scored, that's the signal to add one.
|
|
|
|
Every experiment name `curriculum.py` generates is now timestamped
|
|
(`YYYYMMDD-HHMM-<name>`, e.g. `20260727-0930-curric-s1-unmask`), applied once
|
|
in `run_stage_attempt` — this keeps generation 2's names from colliding with
|
|
generation 1's plain ones (both checkpoint directories and TensorBoard run
|
|
names come straight from `--experiment`) and makes run order obvious in
|
|
TensorBoard without cross-referencing `curriculum_state.json`.
|
|
|
|
All curriculum flags default to leaving Godot's own `@export` defaults
|
|
alone (`train.py` only forwards a flag when you pass it), so ordinary runs
|
|
are unaffected. Full flag list: `--opponent-mode {self_play,inert,frozen}`,
|
|
`--opponent-model <path>` (for `frozen`), `--draw-penalty`,
|
|
`--attack-goal-bias`, `--kickoff-chance`, `--near-goal-chance`,
|
|
`--allow-vertical`/`--no-allow-vertical`, `--allow-pitch-roll`/`--no-allow-pitch-roll`,
|
|
`--velocity-to-ball-weight`, `--ball-distance-penalty`, `--ball-touch-reward`,
|
|
`--airborne-penalty`, `--ball-velocity-to-goal-weight`, `--goal-reward`.
|
|
|
|
### Running it automatically
|
|
|
|
`training/curriculum.py` (started via `curriculum.sh`, same detached-tmux
|
|
pattern as `start_training.sh`) drives all stages end to end: for each
|
|
stage it runs `run_training.sh` (pull, train, export, commit+push) with that
|
|
stage's flags, then evaluates the resulting checkpoint against a reference
|
|
bot over 100 episodes — the fixed `rookie.json` baseline for a from-scratch
|
|
stage 1 (no `resume_from_experiment`/`reference_experiment` override on
|
|
`STAGES[0]`), the previous stage's promoted checkpoint by default for
|
|
stages 2+, or an explicit override in that stage's dict when it deliberately
|
|
skips a since-regressed branch (generation 1's stage 5) or seeds from a
|
|
fixed foundation checkpoint (generation 2's stage 1 — see above).
|
|
|
|
```bash
|
|
cd training
|
|
./curriculum.sh # start/resume the curriculum
|
|
./curriculum.sh --seed-checkpoint checkpoints/run11/final.zip # override stage 1's resume source for this run
|
|
```
|
|
|
|
The gate is deliberately lenient: it blocks a stage only on a **clear
|
|
regression** (the reference beating the candidate by 15+ points of win
|
|
rate), not "must show improvement." A 40-episode eval already misled us once
|
|
in this project — run11 was the first model to deliberately score a goal,
|
|
but its head-to-head eval read as a loss on sample noise alone. A strict
|
|
gate would have retried that stage forever for the wrong reason; a loose one
|
|
still catches a genuinely broken stage. Progress and every attempt's eval
|
|
result are logged to `curriculum_state.json` (committed alongside
|
|
`eval_history.json` after each attempt).
|
|
|
|
A stage gets up to 2 retries (3 attempts total) before the script stops and
|
|
asks for a human look — it will not retry indefinitely or advance past a
|
|
stage that keeps failing on its own. By default a retry resumes from that
|
|
stage's own previous attempt with a fresh `--reset-std`; a stage can instead
|
|
set `reset_retry_checkpoint: True` (generation 2's stage 1 does) to always
|
|
reset to its normal resume source instead — see the generation 1 → 2
|
|
postmortem above for why blind same-checkpoint retries can make things
|
|
monotonically worse. Once you've looked at why a block happened (more
|
|
timesteps? a flag needs adjusting? the eval itself was misleading?), re-run
|
|
with `--force-retry` to try again or `--skip-to-next-stage` if you judge the
|
|
result good enough despite the gate.
|
|
|
|
Running a stage by hand (e.g. to experiment with flags before trusting the
|
|
orchestrator) still works exactly as the table above describes — just call
|
|
`next_run.sh`/`run_training.sh` directly with that stage's flags.
|
|
|
|
## Self-play notes
|
|
|
|
By default both ships share the live policy (mirrored, team-relative
|
|
observations — see `ship_observations.gd`), so training is against the
|
|
current self. `--opponent-mode inert`/`frozen` (see Curriculum training
|
|
above) replace that with a placeholder or a fixed exported policy for one
|
|
side of a run; `frozen` is a single-fixed-model slice of full league play.
|
|
Fixed-opponent training against a *pool* of past checkpoints sampled per
|
|
episode (to avoid strategy collapse on long self-play runs) is still
|
|
deferred — see TODO.md.
|