Files
CosmicClash/TRAINING_LINUX.md
T

5.4 KiB
Raw Blame History

Training on the Linux / RTX 3090 box

Remote-training workflow: run long training sessions on the Linux machine and watch the dashboard from any machine on the network. All training artifacts (checkpoints, TensorBoard logs, exported bots) are committed to git — no result ever depends on a single machine, and moving models between the box and the Mac is just git pull. General training concepts and the export/evaluate workflow live in TRAINING.md — this doc is only what differs on the Linux box.

Both workflows below are wrapped in idempotent scripts in training/ — re-running either is always safe.

One-time setup

GitHub auth first (git-over-HTTPS no longer accepts account passwords, so clone over SSH — this key also lets run_training.sh push results):

ssh-keygen -t ed25519        # accept the defaults
cat ~/.ssh/id_ed25519.pub    # add at github.com/settings/keys → "New SSH key"

cd ~/ai-training
git clone git@github.com:jcreek/CosmicClash.git
# (submodules are editor tooling only — training doesn't need them)

~/ai-training/CosmicClash/training/setup_linux.sh

setup_linux.sh is safe to re-run any time (after a Godot upgrade, a broken venv, a fresh clone — it checks each step before acting). It:

  • downloads the Godot 4.7.1 Linux binary to ~/ai-training/godot/ if missing (override the location by exporting GODOT_BIN);
  • creates training/.venv if missing and installs requirements;
  • verifies CUDA torch, reinstalling from the CUDA wheel index if the box got a CPU-only build;
  • runs the Godot import pass (fresh clones have no .godot/ cache, so class_name scripts aren't registered until the project imports once);
  • finishes with the headless smoke test — the game must boot without rendering.

Maximising throughput

Env stepping is CPU-bound (each --n-parallel instance is one headless Godot process simulating 2 agents); the 3090 only accelerates the PPO updates. So the two levers are instance count and in-engine speedup:

  • --n-parallel: start at nproc minus 2 (leave headroom for the trainer process itself). The Mac mini sustains 6; a many-core box should take considerably more. Each instance opens its own TCP port upward from --port (default 11008).
  • --speedup: in-engine physics time-scale. 16 is proven; try 2432 and keep raising while time/fps in the console/TensorBoard still scales up. Back off if fps stops improving (CPU saturated) or physics glitches appear (ball tunnelling, ships escaping the arena — watch for respawn warnings in the Godot output).

Tune by watching time/fps: run a 2-minute smoke run per setting and keep the best. Reference: 6 instances × speedup 16 ≈ 1,385 steps/s on an M4 Mac mini — a 20M-step run in ~4 h. Doubling fps halves that.

Run training

One command does everything (needs tmux: sudo apt install tmux):

cd ~/ai-training/CosmicClash/training
./start_training.sh run03 --timesteps 20000000 --n-parallel 14 --speedup 24

start_training.sh launches a detached tmux session with two windows — training survives SSH disconnects — and prints the dashboard URL:

  • train runs run_training.sh: git pull → train → export the policy JSON → commit and push checkpoints, logs, and the exported bot;
  • dashboard serves TensorBoard on 0.0.0.0:6006 for the whole network (reused if one is already running).

Re-running start_training.sh while a session exists just attaches you to it (detach again with Ctrl-B then D) — it will never start a second trainer. To stop training, attach and Ctrl-C: the trainer writes final.zip on the way out and the script still exports, commits, and pushes what it has.

Resuming a previous policy (continues its timestep counter; --timesteps is additional steps). Lessons from run01/run02 hard-coded into flags:

./start_training.sh run03 --timesteps 20000000 --n-parallel 14 --speedup 24 \
    --resume checkpoints/run02/final.zip --ent-coef 0.001 --reset-std 0.3
  • --ent-coef — entropy bonus. 0.0001 collapsed the policy std to 0.075 by 20M steps (no exploration left); 0.005 blew it up to 3.0 (random play). 0.001 is the current middle. Healthy train/std drifts between ~0.2 and ~1.0 — check it 3045 min in before committing to a long run.
  • --reset-std — on resume, restores exploration a collapsed checkpoint lost.

Results travel via git

run_training.sh commits and pushes everything a run produces:

  • training/checkpoints/<exp>/ — periodic checkpoints + final.zip, for future --resume, evaluation, and difficulty tiers (an early checkpoint is an easy bot);
  • training/logs/ — TensorBoard history;
  • Game/bots/<exp>.json — the exported policy, immediately playable (point Match or Spectate mode at res://bots/<exp>.json);
  • training/eval_history.json — if evaluations ran.

On the Mac (or anywhere), collecting the results is just git pull. A 20M-step run adds roughly 40 MB of checkpoints — acceptable growth for the guarantee that training is never lost with a machine.

Dashboard over the network

start_training.sh already serves TensorBoard on all interfaces — browse to the URL it prints (http://10.0.0.20:6006) from the Mac or anything on the LAN.

  • If ufw is active on the box: sudo ufw allow 6006/tcp.
  • If you'd rather not open a port, tunnel instead: ssh -L 6006:localhost:6006 10.0.0.20 from the Mac, then browse http://localhost:6006.