5.1 KiB
Training on the Linux / RTX 3090 box
Remote-training workflow: run long training sessions on the Linux machine and
watch the dashboard from any machine on the network. All training artifacts
(checkpoints, TensorBoard logs, exported bots) are committed to git — no
result ever depends on a single machine, and moving models between the box
and the Mac is just git pull. General training concepts and the
export/evaluate workflow live in TRAINING.md — this doc is
only what differs on the Linux box.
Both workflows below are wrapped in idempotent scripts in training/ —
re-running either is always safe.
One-time setup
GitHub auth first (git-over-HTTPS no longer accepts account passwords, so
clone over SSH — this key also lets run_training.sh push results):
ssh-keygen -t ed25519 # accept the defaults
cat ~/.ssh/id_ed25519.pub # add at github.com/settings/keys → "New SSH key"
cd ~/ai-training
git clone git@github.com:jcreek/CosmicClash.git
# (submodules are editor tooling only — training doesn't need them)
~/ai-training/CosmicClash/training/setup_linux.sh
setup_linux.sh is safe to re-run any time (after a Godot upgrade, a broken
venv, a fresh clone — it checks each step before acting). It:
- downloads the Godot 4.7.1 Linux binary to
~/ai-training/godot/if missing (override the location by exportingGODOT_BIN); - creates
training/.venvif missing and installs requirements; - verifies CUDA torch, reinstalling from the CUDA wheel index if the box got a CPU-only build;
- runs the Godot import pass (fresh clones have no
.godot/cache, soclass_namescripts aren't registered until the project imports once); - finishes with the headless smoke test — the game must boot without rendering.
Maximising throughput
Env stepping is CPU-bound (each --n-parallel instance is one headless Godot
process simulating 2 agents); the 3090 only accelerates the PPO updates. So
the two levers are instance count and in-engine speedup:
--n-parallel: start atnprocminus 2 (leave headroom for the trainer process itself). The Mac mini sustains 6; a many-core box should take considerably more. Each instance opens its own TCP port upward from--port(default 11008).--speedup: in-engine physics time-scale. 16 is proven; try 24–32 and keep raising whiletime/fpsin the console/TensorBoard still scales up. Back off if fps stops improving (CPU saturated) or physics glitches appear (ball tunnelling, ships escaping the arena — watch for respawn warnings in the Godot output).
Tune by watching time/fps: run a 2-minute smoke run per setting and keep
the best. Reference: 6 instances × speedup 16 ≈ 1,385 steps/s on an M4 Mac
mini — a 20M-step run in ~4 h. Doubling fps halves that.
Run training
run_training.sh <experiment> [train.py args...] wraps the whole cycle:
git pull → train → export the policy JSON → commit and push checkpoints,
logs, and the exported bot. Run it inside tmux/screen so an SSH
disconnect doesn't kill training. Ctrl-C is safe: the trainer writes
final.zip on the way out, and the script still exports, commits, and
pushes what it has.
Fresh run:
cd ~/ai-training/CosmicClash/training
./run_training.sh run03 --timesteps 20000000 --n-parallel 14 --speedup 24
Resuming a previous policy (continues its timestep counter; --timesteps is
additional steps). Lessons from run01/run02 hard-coded into flags:
./run_training.sh run03 --timesteps 20000000 --n-parallel 14 --speedup 24 \
--resume checkpoints/run02/final.zip --ent-coef 0.001 --reset-std 0.3
--ent-coef— entropy bonus.0.0001collapsed the policy std to 0.075 by 20M steps (no exploration left);0.005blew it up to 3.0 (random play).0.001is the current middle. Healthytrain/stddrifts between ~0.2 and ~1.0 — check it 30–45 min in before committing to a long run.--reset-std— on resume, restores exploration a collapsed checkpoint lost.
Results travel via git
run_training.sh commits and pushes everything a run produces:
training/checkpoints/<exp>/— periodic checkpoints +final.zip, for future--resume, evaluation, and difficulty tiers (an early checkpoint is an easy bot);training/logs/— TensorBoard history;Game/bots/<exp>.json— the exported policy, immediately playable (point Match or Spectate mode atres://bots/<exp>.json);training/eval_history.json— if evaluations ran.
On the Mac (or anywhere), collecting the results is just git pull. A
20M-step run adds roughly 40 MB of checkpoints — acceptable growth for the
guarantee that training is never lost with a machine.
Dashboard over the network
On the Linux box, bind TensorBoard to all interfaces instead of localhost:
cd ~/ai-training/CosmicClash/training
.venv/bin/tensorboard --logdir logs --host 0.0.0.0 --port 6006
Then from the Mac (or anything on the LAN): http://<linux-box-hostname>:6006.
- If
ufwis active on the box:sudo ufw allow 6006/tcp. - If you'd rather not open a port, tunnel instead:
ssh -L 6006:localhost:6006 <linux-box>from the Mac, then browsehttp://localhost:6006.