5.4 KiB
Training on the Linux / RTX 3090 box
Remote-training workflow: run long training sessions on the Linux machine and
watch the dashboard from any machine on the network. All training artifacts
(checkpoints, TensorBoard logs, exported bots) are committed to git — no
result ever depends on a single machine, and moving models between the box
and the Mac is just git pull. General training concepts and the
export/evaluate workflow live in TRAINING.md — this doc is
only what differs on the Linux box.
Both workflows below are wrapped in idempotent scripts in training/ —
re-running either is always safe.
One-time setup
GitHub auth first (git-over-HTTPS no longer accepts account passwords, so
clone over SSH — this key also lets run_training.sh push results):
ssh-keygen -t ed25519 # accept the defaults
cat ~/.ssh/id_ed25519.pub # add at github.com/settings/keys → "New SSH key"
cd ~/ai-training
git clone git@github.com:jcreek/CosmicClash.git
# (submodules are editor tooling only — training doesn't need them)
~/ai-training/CosmicClash/training/setup_linux.sh
setup_linux.sh is safe to re-run any time (after a Godot upgrade, a broken
venv, a fresh clone — it checks each step before acting). It:
- downloads the Godot 4.7.1 Linux binary to
~/ai-training/godot/if missing (override the location by exportingGODOT_BIN); - creates
training/.venvif missing and installs requirements; - verifies CUDA torch, reinstalling from the CUDA wheel index if the box got a CPU-only build;
- runs the Godot import pass (fresh clones have no
.godot/cache, soclass_namescripts aren't registered until the project imports once); - finishes with the headless smoke test — the game must boot without rendering.
Maximising throughput
Env stepping is CPU-bound (each --n-parallel instance is one headless Godot
process simulating 2 agents); the 3090 only accelerates the PPO updates. So
the two levers are instance count and in-engine speedup:
--n-parallel: start atnprocminus 2 (leave headroom for the trainer process itself). The Mac mini sustains 6; a many-core box should take considerably more. Each instance opens its own TCP port upward from--port(default 11008).--speedup: in-engine physics time-scale. 16 is proven; try 24–32 and keep raising whiletime/fpsin the console/TensorBoard still scales up. Back off if fps stops improving (CPU saturated) or physics glitches appear (ball tunnelling, ships escaping the arena — watch for respawn warnings in the Godot output).
Tune by watching time/fps: run a 2-minute smoke run per setting and keep
the best. Reference: 6 instances × speedup 16 ≈ 1,385 steps/s on an M4 Mac
mini — a 20M-step run in ~4 h. Doubling fps halves that.
Run training
One command does everything (needs tmux: sudo apt install tmux):
cd ~/ai-training/CosmicClash/training
./start_training.sh run03 --timesteps 20000000 --n-parallel 14 --speedup 24
start_training.sh launches a detached tmux session with two windows —
training survives SSH disconnects — and prints the dashboard URL:
trainrunsrun_training.sh:git pull→ train → export the policy JSON → commit and push checkpoints, logs, and the exported bot;dashboardserves TensorBoard on0.0.0.0:6006for the whole network (reused if one is already running).
Re-running start_training.sh while a session exists just attaches you to
it (detach again with Ctrl-B then D) — it will never start a second
trainer. To stop training, attach and Ctrl-C: the trainer writes final.zip
on the way out and the script still exports, commits, and pushes what it has.
Resuming a previous policy (continues its timestep counter; --timesteps is
additional steps). Lessons from run01/run02 hard-coded into flags:
./start_training.sh run03 --timesteps 20000000 --n-parallel 14 --speedup 24 \
--resume checkpoints/run02/final.zip --ent-coef 0.001 --reset-std 0.3
--ent-coef— entropy bonus.0.0001collapsed the policy std to 0.075 by 20M steps (no exploration left);0.005blew it up to 3.0 (random play).0.001is the current middle. Healthytrain/stddrifts between ~0.2 and ~1.0 — check it 30–45 min in before committing to a long run.--reset-std— on resume, restores exploration a collapsed checkpoint lost.
Results travel via git
run_training.sh commits and pushes everything a run produces:
training/checkpoints/<exp>/— periodic checkpoints +final.zip, for future--resume, evaluation, and difficulty tiers (an early checkpoint is an easy bot);training/logs/— TensorBoard history;Game/bots/<exp>.json— the exported policy, immediately playable (point Match or Spectate mode atres://bots/<exp>.json);training/eval_history.json— if evaluations ran.
On the Mac (or anywhere), collecting the results is just git pull. A
20M-step run adds roughly 40 MB of checkpoints — acceptable growth for the
guarantee that training is never lost with a machine.
Dashboard over the network
start_training.sh already serves TensorBoard on all interfaces — browse to
the URL it prints (http://10.0.0.20:6006) from the Mac or anything on
the LAN.
- If
ufwis active on the box:sudo ufw allow 6006/tcp. - If you'd rather not open a port, tunnel instead:
ssh -L 6006:localhost:6006 10.0.0.20from the Mac, then browsehttp://localhost:6006.