8.0 KiB
Training on the Linux / RTX 3090 box
Remote-training workflow: run long training sessions on the Linux machine and
watch the dashboard from any machine on the network. All training artifacts
(checkpoints, TensorBoard logs, exported bots) are committed to git — no
result ever depends on a single machine, and moving models between the box
and the Mac is just git pull. General training concepts and the
export/evaluate workflow live in TRAINING.md — this doc is
only what differs on the Linux box.
Both workflows below are wrapped in idempotent scripts in training/ —
re-running either is always safe.
One-time setup
GitHub auth first (git-over-HTTPS no longer accepts account passwords, so
clone over SSH — this key also lets run_training.sh push results):
ssh-keygen -t ed25519 # accept the defaults
cat ~/.ssh/id_ed25519.pub # add at github.com/settings/keys → "New SSH key"
cd ~/ai-training
git clone git@github.com:jcreek/CosmicClash.git
# (submodules are editor tooling only — training doesn't need them)
~/ai-training/CosmicClash/training/setup_linux.sh
setup_linux.sh is safe to re-run any time (after a Godot upgrade, a broken
venv, a fresh clone — it checks each step before acting). It:
- downloads the Godot 4.7.1 Linux binary to
~/ai-training/godot/if missing (override the location by exportingGODOT_BIN); - creates
training/.venvif missing and installs requirements; - verifies CUDA torch, reinstalling from the CUDA wheel index if the box got a CPU-only build;
- runs the Godot import pass (fresh clones have no
.godot/cache, soclass_namescripts aren't registered until the project imports once); - finishes with the headless smoke test — the game must boot without rendering.
Maximising throughput
Env stepping is CPU-bound (each --n-parallel instance is one headless Godot
process simulating 2 agents); the 3090 only accelerates the PPO updates. So
the two levers are instance count and in-engine speedup:
--n-parallel: start atnprocminus 2 (leave headroom for the trainer process itself). The Mac mini sustains 6; a many-core box should take considerably more. Each instance opens its own TCP port upward from--port(default 11008).--speedup: in-engine physics time-scale. 16 is proven; try 24–32 and keep raising whiletime/fpsin the console/TensorBoard still scales up. Back off if fps stops improving (CPU saturated) or physics glitches appear (ball tunnelling, ships escaping the arena — watch for respawn warnings in the Godot output).
Tune by watching time/fps: run a 2-minute smoke run per setting and keep
the best. Reference: 6 instances × speedup 16 ≈ 1,385 steps/s on an M4 Mac
mini — a 20M-step run in ~4 h. Doubling fps halves that.
Run training
One command does everything (needs tmux: sudo apt install tmux):
cd ~/ai-training/CosmicClash/training
./start_training.sh run03 --timesteps 20000000 --n-parallel 14 --speedup 24
start_training.sh launches a detached tmux session with two windows —
training survives SSH disconnects — and prints the dashboard URL:
trainrunsrun_training.sh:git pull→ train → export the policy JSON → commit and push checkpoints, logs, and the exported bot;dashboardserves TensorBoard on0.0.0.0:6006for the whole network (reused if one is already running).
Re-running start_training.sh while a session exists just attaches you to
it (detach again with Ctrl-B then D) — it will never start a second
trainer. To stop training, attach and Ctrl-C: the trainer writes final.zip
on the way out and the script still exports, commits, and pushes what it has.
Resuming a previous policy (continues its timestep counter; --timesteps is
additional steps). Lessons from run01/run02 hard-coded into flags:
./start_training.sh run03 --timesteps 20000000 --n-parallel 14 --speedup 24 \
--resume checkpoints/run02/final.zip --ent-coef 0.001 --reset-std 0.3
--ent-coef— entropy bonus.0.0001collapsed the policy std to 0.075 by 20M steps (no exploration left);0.005blew it up to 3.0 (random play).0.001is the current middle. Healthytrain/stddrifts between ~0.2 and ~1.0 — check it 30–45 min in before committing to a long run.--reset-std— on resume, restores exploration a collapsed checkpoint lost.
Results travel via git
run_training.sh commits and pushes everything a run produces:
training/checkpoints/<exp>/— periodic checkpoints +final.zip, for future--resume, evaluation, and difficulty tiers (an early checkpoint is an easy bot);training/logs/— TensorBoard history;Game/bots/<exp>.json— the exported policy, immediately playable (point Match or Spectate mode atres://bots/<exp>.json);training/eval_history.json— if evaluations ran.
On the Mac (or anywhere), collecting the results is just git pull. A
20M-step run adds roughly 40 MB of checkpoints — acceptable growth for the
guarantee that training is never lost with a machine.
Exported-binary training (faster parallel startup)
By default (and in every example above) train.py runs the project from
source via --godot_bin — each of the --n-parallel instances re-parses
project settings and re-imports scripts/resources on launch. An exported
build skips that: resources are pre-imported and packed once at export time,
so each instance just loads a binary. Worth it once --n-parallel is large
enough that per-instance startup overhead adds up (i.e. this box, not the Mac
mini's --n-parallel 6).
Opt in once — run_training.sh (and so start_training.sh/next_run.sh/
curriculum.sh, which all funnel through it) takes it from there automatically:
./export_linux.sh # one-time opt-in: builds training/build/CosmicClash.x86_64
./next_run.sh # from here on, every standing/curriculum run uses it
You don't need to (and shouldn't) pass --exported-binary yourself through
those entry points — run_training.sh adds it whenever training/build/
exists, after re-exporting against whatever git pull just fetched. Calling
train.py directly still takes it explicitly, same as any other flag:
.venv/bin/python train.py --experiment run04 --exported-binary build/CosmicClash.x86_64 \
--timesteps 20000000 --n-parallel 14 --speedup 24
To go back to a source run permanently, delete training/build/ — with it
gone, run_training.sh stops re-exporting and stops adding the flag, so
next_run.sh/curriculum.sh revert to plain source runs with no code changes.
export_linux.sh builds from the "Linux Training" preset in
Game/export_presets.cfg, which is training-only — its
custom_features="training" activates project.godot's
run/main_scene.training override, so the resulting binary boots straight
into training.tscn on its own. This indirection is required, not
incidental: official Godot export templates have path/scene overrides
compiled out, so passing --scene at launch time (the way the source run
does) hard-aborts an exported binary with "compiled without support for path
overrides" — there's no way to redirect an exported build to a different
scene at runtime. Because the main scene is baked in at export time, this
preset can't later double as a normal "ship the game" Linux build (which
would need main_menu.tscn and no training feature tag) — a real game export
would need its own separate preset.
setup_linux.sh installs the export templates this needs alongside the
Godot binary. If you never opt in (no training/build/ directory), this
costs nothing — training stays a plain source run.
Dashboard over the network
start_training.sh already serves TensorBoard on all interfaces — browse to
the URL it prints (http://10.0.0.20:6006) from the Mac or anything on
the LAN.
- If
ufwis active on the box:sudo ufw allow 6006/tcp. - If you'd rather not open a port, tunnel instead:
ssh -L 6006:localhost:6006 10.0.0.20from the Mac, then browsehttp://localhost:6006.