# Training on the Linux / RTX 3090 box Remote-training workflow: run long training sessions on the Linux machine and watch the dashboard from any machine on the network. **All training artifacts (checkpoints, TensorBoard logs, exported bots) are committed to git** — no result ever depends on a single machine, and moving models between the box and the Mac is just `git pull`. General training concepts and the export/evaluate workflow live in [TRAINING.md](TRAINING.md) — this doc is only what differs on the Linux box. Both workflows below are wrapped in idempotent scripts in `training/` — re-running either is always safe. ## One-time setup GitHub auth first (git-over-HTTPS no longer accepts account passwords, so clone over SSH — this key also lets `run_training.sh` push results): ```bash ssh-keygen -t ed25519 # accept the defaults cat ~/.ssh/id_ed25519.pub # add at github.com/settings/keys → "New SSH key" cd ~/ai-training git clone git@github.com:jcreek/CosmicClash.git # (submodules are editor tooling only — training doesn't need them) ~/ai-training/CosmicClash/training/setup_linux.sh ``` `setup_linux.sh` is safe to re-run any time (after a Godot upgrade, a broken venv, a fresh clone — it checks each step before acting). It: - downloads the Godot 4.7.1 Linux binary to `~/ai-training/godot/` if missing (override the location by exporting `GODOT_BIN`); - creates `training/.venv` if missing and installs requirements; - verifies CUDA torch, reinstalling from the CUDA wheel index if the box got a CPU-only build; - runs the Godot import pass (fresh clones have no `.godot/` cache, so `class_name` scripts aren't registered until the project imports once); - finishes with the headless smoke test — the game must boot without rendering. ## Maximising throughput Env stepping is CPU-bound (each `--n-parallel` instance is one headless Godot process simulating 2 agents); the 3090 only accelerates the PPO updates. So the two levers are instance count and in-engine speedup: - **`--n-parallel`**: start at `nproc` minus 2 (leave headroom for the trainer process itself). The Mac mini sustains 6; a many-core box should take considerably more. Each instance opens its own TCP port upward from `--port` (default 11008). - **`--speedup`**: in-engine physics time-scale. 16 is proven; try 24–32 and keep raising while `time/fps` in the console/TensorBoard still scales up. Back off if fps stops improving (CPU saturated) or physics glitches appear (ball tunnelling, ships escaping the arena — watch for respawn warnings in the Godot output). Tune by watching `time/fps`: run a 2-minute smoke run per setting and keep the best. Reference: 6 instances × speedup 16 ≈ 1,385 steps/s on an M4 Mac mini — a 20M-step run in ~4 h. Doubling fps halves that. ## Run training `run_training.sh [train.py args...]` wraps the whole cycle: `git pull` → train → export the policy JSON → commit and push checkpoints, logs, and the exported bot. Run it inside `tmux`/`screen` so an SSH disconnect doesn't kill training. Ctrl-C is safe: the trainer writes `final.zip` on the way out, and the script still exports, commits, and pushes what it has. Fresh run: ```bash cd ~/ai-training/CosmicClash/training ./run_training.sh run03 --timesteps 20000000 --n-parallel 14 --speedup 24 ``` Resuming a previous policy (continues its timestep counter; `--timesteps` is *additional* steps). Lessons from run01/run02 hard-coded into flags: ```bash ./run_training.sh run03 --timesteps 20000000 --n-parallel 14 --speedup 24 \ --resume checkpoints/run02/final.zip --ent-coef 0.001 --reset-std 0.3 ``` - `--ent-coef` — entropy bonus. `0.0001` collapsed the policy std to 0.075 by 20M steps (no exploration left); `0.005` blew it up to 3.0 (random play). `0.001` is the current middle. Healthy `train/std` drifts between ~0.2 and ~1.0 — check it 30–45 min in before committing to a long run. - `--reset-std` — on resume, restores exploration a collapsed checkpoint lost. ## Results travel via git `run_training.sh` commits and pushes everything a run produces: - `training/checkpoints//` — periodic checkpoints + `final.zip`, for future `--resume`, evaluation, and difficulty tiers (an early checkpoint *is* an easy bot); - `training/logs/` — TensorBoard history; - `Game/bots/.json` — the exported policy, immediately playable (point Match or Spectate mode at `res://bots/.json`); - `training/eval_history.json` — if evaluations ran. On the Mac (or anywhere), collecting the results is just `git pull`. A 20M-step run adds roughly 40 MB of checkpoints — acceptable growth for the guarantee that training is never lost with a machine. ## Dashboard over the network On the Linux box, bind TensorBoard to all interfaces instead of localhost: ```bash cd ~/ai-training/CosmicClash/training .venv/bin/tensorboard --logdir logs --host 0.0.0.0 --port 6006 ``` Then from the Mac (or anything on the LAN): `http://:6006`. - If `ufw` is active on the box: `sudo ufw allow 6006/tcp`. - If you'd rather not open a port, tunnel instead: `ssh -L 6006:localhost:6006 ` from the Mac, then browse `http://localhost:6006`.