mirror of
https://github.com/jcreek/CosmicClash.git
synced 2026-09-10 16:04:04 +00:00
124 lines
5.1 KiB
Markdown
124 lines
5.1 KiB
Markdown
# Training on the Linux / RTX 3090 box
|
||
|
||
Remote-training workflow: run long training sessions on the Linux machine and
|
||
watch the dashboard from any machine on the network. **All training artifacts
|
||
(checkpoints, TensorBoard logs, exported bots) are committed to git** — no
|
||
result ever depends on a single machine, and moving models between the box
|
||
and the Mac is just `git pull`. General training concepts and the
|
||
export/evaluate workflow live in [TRAINING.md](TRAINING.md) — this doc is
|
||
only what differs on the Linux box.
|
||
|
||
Both workflows below are wrapped in idempotent scripts in `training/` —
|
||
re-running either is always safe.
|
||
|
||
## One-time setup
|
||
|
||
GitHub auth first (git-over-HTTPS no longer accepts account passwords, so
|
||
clone over SSH — this key also lets `run_training.sh` push results):
|
||
|
||
```bash
|
||
ssh-keygen -t ed25519 # accept the defaults
|
||
cat ~/.ssh/id_ed25519.pub # add at github.com/settings/keys → "New SSH key"
|
||
|
||
cd ~/ai-training
|
||
git clone git@github.com:jcreek/CosmicClash.git
|
||
# (submodules are editor tooling only — training doesn't need them)
|
||
|
||
~/ai-training/CosmicClash/training/setup_linux.sh
|
||
```
|
||
|
||
`setup_linux.sh` is safe to re-run any time (after a Godot upgrade, a broken
|
||
venv, a fresh clone — it checks each step before acting). It:
|
||
|
||
- downloads the Godot 4.7.1 Linux binary to `~/ai-training/godot/` if missing
|
||
(override the location by exporting `GODOT_BIN`);
|
||
- creates `training/.venv` if missing and installs requirements;
|
||
- verifies CUDA torch, reinstalling from the CUDA wheel index if the box got
|
||
a CPU-only build;
|
||
- runs the Godot import pass (fresh clones have no `.godot/` cache, so
|
||
`class_name` scripts aren't registered until the project imports once);
|
||
- finishes with the headless smoke test — the game must boot without
|
||
rendering.
|
||
|
||
## Maximising throughput
|
||
|
||
Env stepping is CPU-bound (each `--n-parallel` instance is one headless Godot
|
||
process simulating 2 agents); the 3090 only accelerates the PPO updates. So
|
||
the two levers are instance count and in-engine speedup:
|
||
|
||
- **`--n-parallel`**: start at `nproc` minus 2 (leave headroom for the
|
||
trainer process itself). The Mac mini sustains 6; a many-core box should
|
||
take considerably more. Each instance opens its own TCP port upward from
|
||
`--port` (default 11008).
|
||
- **`--speedup`**: in-engine physics time-scale. 16 is proven; try 24–32 and
|
||
keep raising while `time/fps` in the console/TensorBoard still scales up.
|
||
Back off if fps stops improving (CPU saturated) or physics glitches appear
|
||
(ball tunnelling, ships escaping the arena — watch for respawn warnings in
|
||
the Godot output).
|
||
|
||
Tune by watching `time/fps`: run a 2-minute smoke run per setting and keep
|
||
the best. Reference: 6 instances × speedup 16 ≈ 1,385 steps/s on an M4 Mac
|
||
mini — a 20M-step run in ~4 h. Doubling fps halves that.
|
||
|
||
## Run training
|
||
|
||
`run_training.sh <experiment> [train.py args...]` wraps the whole cycle:
|
||
`git pull` → train → export the policy JSON → commit and push checkpoints,
|
||
logs, and the exported bot. Run it inside `tmux`/`screen` so an SSH
|
||
disconnect doesn't kill training. Ctrl-C is safe: the trainer writes
|
||
`final.zip` on the way out, and the script still exports, commits, and
|
||
pushes what it has.
|
||
|
||
Fresh run:
|
||
|
||
```bash
|
||
cd ~/ai-training/CosmicClash/training
|
||
./run_training.sh run03 --timesteps 20000000 --n-parallel 14 --speedup 24
|
||
```
|
||
|
||
Resuming a previous policy (continues its timestep counter; `--timesteps` is
|
||
*additional* steps). Lessons from run01/run02 hard-coded into flags:
|
||
|
||
```bash
|
||
./run_training.sh run03 --timesteps 20000000 --n-parallel 14 --speedup 24 \
|
||
--resume checkpoints/run02/final.zip --ent-coef 0.001 --reset-std 0.3
|
||
```
|
||
|
||
- `--ent-coef` — entropy bonus. `0.0001` collapsed the policy std to 0.075 by
|
||
20M steps (no exploration left); `0.005` blew it up to 3.0 (random play).
|
||
`0.001` is the current middle. Healthy `train/std` drifts between ~0.2 and
|
||
~1.0 — check it 30–45 min in before committing to a long run.
|
||
- `--reset-std` — on resume, restores exploration a collapsed checkpoint lost.
|
||
|
||
## Results travel via git
|
||
|
||
`run_training.sh` commits and pushes everything a run produces:
|
||
|
||
- `training/checkpoints/<exp>/` — periodic checkpoints + `final.zip`, for
|
||
future `--resume`, evaluation, and difficulty tiers (an early checkpoint
|
||
*is* an easy bot);
|
||
- `training/logs/` — TensorBoard history;
|
||
- `Game/bots/<exp>.json` — the exported policy, immediately playable (point
|
||
Match or Spectate mode at `res://bots/<exp>.json`);
|
||
- `training/eval_history.json` — if evaluations ran.
|
||
|
||
On the Mac (or anywhere), collecting the results is just `git pull`. A
|
||
20M-step run adds roughly 40 MB of checkpoints — acceptable growth for the
|
||
guarantee that training is never lost with a machine.
|
||
|
||
## Dashboard over the network
|
||
|
||
On the Linux box, bind TensorBoard to all interfaces instead of localhost:
|
||
|
||
```bash
|
||
cd ~/ai-training/CosmicClash/training
|
||
.venv/bin/tensorboard --logdir logs --host 0.0.0.0 --port 6006
|
||
```
|
||
|
||
Then from the Mac (or anything on the LAN): `http://<linux-box-hostname>:6006`.
|
||
|
||
- If `ufw` is active on the box: `sudo ufw allow 6006/tcp`.
|
||
- If you'd rather not open a port, tunnel instead:
|
||
`ssh -L 6006:localhost:6006 <linux-box>` from the Mac, then browse
|
||
`http://localhost:6006`.
|