Files
CosmicClash/TRAINING_LINUX.md
T

124 lines
5.1 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Training on the Linux / RTX 3090 box
Remote-training workflow: run long training sessions on the Linux machine and
watch the dashboard from any machine on the network. **All training artifacts
(checkpoints, TensorBoard logs, exported bots) are committed to git** — no
result ever depends on a single machine, and moving models between the box
and the Mac is just `git pull`. General training concepts and the
export/evaluate workflow live in [TRAINING.md](TRAINING.md) — this doc is
only what differs on the Linux box.
Both workflows below are wrapped in idempotent scripts in `training/`
re-running either is always safe.
## One-time setup
GitHub auth first (git-over-HTTPS no longer accepts account passwords, so
clone over SSH — this key also lets `run_training.sh` push results):
```bash
ssh-keygen -t ed25519 # accept the defaults
cat ~/.ssh/id_ed25519.pub # add at github.com/settings/keys → "New SSH key"
cd ~/ai-training
git clone git@github.com:jcreek/CosmicClash.git
# (submodules are editor tooling only — training doesn't need them)
~/ai-training/CosmicClash/training/setup_linux.sh
```
`setup_linux.sh` is safe to re-run any time (after a Godot upgrade, a broken
venv, a fresh clone — it checks each step before acting). It:
- downloads the Godot 4.7.1 Linux binary to `~/ai-training/godot/` if missing
(override the location by exporting `GODOT_BIN`);
- creates `training/.venv` if missing and installs requirements;
- verifies CUDA torch, reinstalling from the CUDA wheel index if the box got
a CPU-only build;
- runs the Godot import pass (fresh clones have no `.godot/` cache, so
`class_name` scripts aren't registered until the project imports once);
- finishes with the headless smoke test — the game must boot without
rendering.
## Maximising throughput
Env stepping is CPU-bound (each `--n-parallel` instance is one headless Godot
process simulating 2 agents); the 3090 only accelerates the PPO updates. So
the two levers are instance count and in-engine speedup:
- **`--n-parallel`**: start at `nproc` minus 2 (leave headroom for the
trainer process itself). The Mac mini sustains 6; a many-core box should
take considerably more. Each instance opens its own TCP port upward from
`--port` (default 11008).
- **`--speedup`**: in-engine physics time-scale. 16 is proven; try 2432 and
keep raising while `time/fps` in the console/TensorBoard still scales up.
Back off if fps stops improving (CPU saturated) or physics glitches appear
(ball tunnelling, ships escaping the arena — watch for respawn warnings in
the Godot output).
Tune by watching `time/fps`: run a 2-minute smoke run per setting and keep
the best. Reference: 6 instances × speedup 16 ≈ 1,385 steps/s on an M4 Mac
mini — a 20M-step run in ~4 h. Doubling fps halves that.
## Run training
`run_training.sh <experiment> [train.py args...]` wraps the whole cycle:
`git pull` → train → export the policy JSON → commit and push checkpoints,
logs, and the exported bot. Run it inside `tmux`/`screen` so an SSH
disconnect doesn't kill training. Ctrl-C is safe: the trainer writes
`final.zip` on the way out, and the script still exports, commits, and
pushes what it has.
Fresh run:
```bash
cd ~/ai-training/CosmicClash/training
./run_training.sh run03 --timesteps 20000000 --n-parallel 14 --speedup 24
```
Resuming a previous policy (continues its timestep counter; `--timesteps` is
*additional* steps). Lessons from run01/run02 hard-coded into flags:
```bash
./run_training.sh run03 --timesteps 20000000 --n-parallel 14 --speedup 24 \
--resume checkpoints/run02/final.zip --ent-coef 0.001 --reset-std 0.3
```
- `--ent-coef` — entropy bonus. `0.0001` collapsed the policy std to 0.075 by
20M steps (no exploration left); `0.005` blew it up to 3.0 (random play).
`0.001` is the current middle. Healthy `train/std` drifts between ~0.2 and
~1.0 — check it 3045 min in before committing to a long run.
- `--reset-std` — on resume, restores exploration a collapsed checkpoint lost.
## Results travel via git
`run_training.sh` commits and pushes everything a run produces:
- `training/checkpoints/<exp>/` — periodic checkpoints + `final.zip`, for
future `--resume`, evaluation, and difficulty tiers (an early checkpoint
*is* an easy bot);
- `training/logs/` — TensorBoard history;
- `Game/bots/<exp>.json` — the exported policy, immediately playable (point
Match or Spectate mode at `res://bots/<exp>.json`);
- `training/eval_history.json` — if evaluations ran.
On the Mac (or anywhere), collecting the results is just `git pull`. A
20M-step run adds roughly 40 MB of checkpoints — acceptable growth for the
guarantee that training is never lost with a machine.
## Dashboard over the network
On the Linux box, bind TensorBoard to all interfaces instead of localhost:
```bash
cd ~/ai-training/CosmicClash/training
.venv/bin/tensorboard --logdir logs --host 0.0.0.0 --port 6006
```
Then from the Mac (or anything on the LAN): `http://<linux-box-hostname>:6006`.
- If `ufw` is active on the box: `sudo ufw allow 6006/tcp`.
- If you'd rather not open a port, tunnel instead:
`ssh -L 6006:localhost:6006 <linux-box>` from the Mac, then browse
`http://localhost:6006`.