mirror of
https://github.com/jcreek/CosmicClash.git
synced 2026-09-10 16:04:04 +00:00
chore(training): Track training artifacts in git, add idempotent Linux setup/run scripts, and commit run01 results
This commit is contained in:
+49
-55
@@ -1,44 +1,44 @@
|
||||
# Training on the Linux / RTX 3090 box
|
||||
|
||||
Remote-training workflow: run long training sessions on the Linux machine,
|
||||
watch the dashboard from any machine on the network, and ship the trained
|
||||
model back to the Mac mini automatically when the run finishes. General
|
||||
training concepts and the export/evaluate workflow live in
|
||||
[TRAINING.md](TRAINING.md) — this doc is only what differs on the Linux box.
|
||||
Remote-training workflow: run long training sessions on the Linux machine and
|
||||
watch the dashboard from any machine on the network. **All training artifacts
|
||||
(checkpoints, TensorBoard logs, exported bots) are committed to git** — no
|
||||
result ever depends on a single machine, and moving models between the box
|
||||
and the Mac is just `git pull`. General training concepts and the
|
||||
export/evaluate workflow live in [TRAINING.md](TRAINING.md) — this doc is
|
||||
only what differs on the Linux box.
|
||||
|
||||
Both workflows below are wrapped in idempotent scripts in `training/` —
|
||||
re-running either is always safe.
|
||||
|
||||
## One-time setup
|
||||
|
||||
GitHub auth first (git-over-HTTPS no longer accepts account passwords, so
|
||||
clone over SSH — this key also lets `run_training.sh` push results):
|
||||
|
||||
```bash
|
||||
# GitHub auth (one-time): git-over-HTTPS no longer accepts account passwords,
|
||||
# so clone over SSH. Generate a key, then add the printed public key at
|
||||
# github.com/settings/keys → "New SSH key".
|
||||
ssh-keygen -t ed25519 # accept the defaults
|
||||
cat ~/.ssh/id_ed25519.pub
|
||||
cat ~/.ssh/id_ed25519.pub # add at github.com/settings/keys → "New SSH key"
|
||||
|
||||
cd ~/ai-training
|
||||
git clone git@github.com:jcreek/CosmicClash.git
|
||||
# (submodules are editor tooling only — training doesn't need them)
|
||||
|
||||
# Godot 4.7.1 Linux binary
|
||||
mkdir -p ~/ai-training/godot && cd ~/ai-training/godot
|
||||
wget https://github.com/godotengine/godot/releases/download/4.7.1-stable/Godot_v4.7.1-stable_linux.x86_64.zip
|
||||
unzip Godot_v4.7.1-stable_linux.x86_64.zip
|
||||
echo 'export GODOT_BIN=~/ai-training/godot/Godot_v4.7.1-stable_linux.x86_64' >> ~/.bashrc && source ~/.bashrc
|
||||
|
||||
# Python env
|
||||
cd ~/ai-training/CosmicClash/training
|
||||
python3 -m venv .venv
|
||||
.venv/bin/pip install -r requirements.txt
|
||||
|
||||
# CUDA sanity check — should print True
|
||||
.venv/bin/python -c "import torch; print(torch.cuda.is_available())"
|
||||
|
||||
# Headless smoke test — the game must boot without rendering
|
||||
$GODOT_BIN --headless --path ../Game res://scenes/free_play.tscn --quit-after 300
|
||||
~/ai-training/CosmicClash/training/setup_linux.sh
|
||||
```
|
||||
|
||||
If the CUDA check prints `False`, reinstall torch from the CUDA index
|
||||
(`pip install torch --index-url https://download.pytorch.org/whl/cu121`).
|
||||
`setup_linux.sh` is safe to re-run any time (after a Godot upgrade, a broken
|
||||
venv, a fresh clone — it checks each step before acting). It:
|
||||
|
||||
- downloads the Godot 4.7.1 Linux binary to `~/ai-training/godot/` if missing
|
||||
(override the location by exporting `GODOT_BIN`);
|
||||
- creates `training/.venv` if missing and installs requirements;
|
||||
- verifies CUDA torch, reinstalling from the CUDA wheel index if the box got
|
||||
a CPU-only build;
|
||||
- runs the Godot import pass (fresh clones have no `.godot/` cache, so
|
||||
`class_name` scripts aren't registered until the project imports once);
|
||||
- finishes with the headless smoke test — the game must boot without
|
||||
rendering.
|
||||
|
||||
## Maximising throughput
|
||||
|
||||
@@ -60,22 +60,27 @@ Tune by watching `time/fps`: run a 2-minute smoke run per setting and keep
|
||||
the best. Reference: 6 instances × speedup 16 ≈ 1,385 steps/s on an M4 Mac
|
||||
mini — a 20M-step run in ~4 h. Doubling fps halves that.
|
||||
|
||||
## Start a run
|
||||
## Run training
|
||||
|
||||
`run_training.sh <experiment> [train.py args...]` wraps the whole cycle:
|
||||
`git pull` → train → export the policy JSON → commit and push checkpoints,
|
||||
logs, and the exported bot. Run it inside `tmux`/`screen` so an SSH
|
||||
disconnect doesn't kill training. Ctrl-C is safe: the trainer writes
|
||||
`final.zip` on the way out, and the script still exports, commits, and
|
||||
pushes what it has.
|
||||
|
||||
Fresh run:
|
||||
|
||||
```bash
|
||||
cd ~/ai-training/CosmicClash/training
|
||||
.venv/bin/python train.py --experiment run03 --timesteps 20000000 \
|
||||
--n-parallel 14 --speedup 24
|
||||
./run_training.sh run03 --timesteps 20000000 --n-parallel 14 --speedup 24
|
||||
```
|
||||
|
||||
Resuming a previous policy (continues its timestep counter; `--timesteps` is
|
||||
*additional* steps). Lessons from run01/run02 hard-coded into flags:
|
||||
|
||||
```bash
|
||||
.venv/bin/python train.py --experiment run03 --timesteps 20000000 \
|
||||
--n-parallel 14 --speedup 24 \
|
||||
./run_training.sh run03 --timesteps 20000000 --n-parallel 14 --speedup 24 \
|
||||
--resume checkpoints/run02/final.zip --ent-coef 0.001 --reset-std 0.3
|
||||
```
|
||||
|
||||
@@ -85,32 +90,21 @@ Resuming a previous policy (continues its timestep counter; `--timesteps` is
|
||||
~1.0 — check it 30–45 min in before committing to a long run.
|
||||
- `--reset-std` — on resume, restores exploration a collapsed checkpoint lost.
|
||||
|
||||
Run inside `tmux`/`screen` so an SSH disconnect doesn't kill training.
|
||||
Ctrl-C is safe: `final.zip` is written on the way out.
|
||||
## Results travel via git
|
||||
|
||||
## Auto-copy the result to the Mac mini when training finishes
|
||||
`run_training.sh` commits and pushes everything a run produces:
|
||||
|
||||
One-time: enable **System Settings → General → Sharing → Remote Login** on
|
||||
the Mac mini, and `ssh-copy-id jcreek@Joshs-Mac-mini.local` from the Linux
|
||||
box so rsync runs unattended.
|
||||
- `training/checkpoints/<exp>/` — periodic checkpoints + `final.zip`, for
|
||||
future `--resume`, evaluation, and difficulty tiers (an early checkpoint
|
||||
*is* an easy bot);
|
||||
- `training/logs/` — TensorBoard history;
|
||||
- `Game/bots/<exp>.json` — the exported policy, immediately playable (point
|
||||
Match or Spectate mode at `res://bots/<exp>.json`);
|
||||
- `training/eval_history.json` — if evaluations ran.
|
||||
|
||||
Chain export + copy onto the training command (`;` not `&&`, so the copy
|
||||
still happens after a Ctrl-C — `final.zip` exists either way):
|
||||
|
||||
```bash
|
||||
EXP=run03
|
||||
.venv/bin/python train.py --experiment $EXP --timesteps 20000000 --n-parallel 14 --speedup 24 ; \
|
||||
.venv/bin/python export_policy.py checkpoints/$EXP/final.zip ../Game/bots/$EXP.json && \
|
||||
rsync -av checkpoints/$EXP/final.zip \
|
||||
jcreek@Joshs-Mac-mini.local:~/Documents/repos/GitHub/CosmicClash/training/checkpoints/$EXP/ && \
|
||||
rsync -av ../Game/bots/$EXP.json \
|
||||
jcreek@Joshs-Mac-mini.local:~/Documents/repos/GitHub/CosmicClash/Game/bots/
|
||||
```
|
||||
|
||||
That lands both the raw checkpoint (for future `--resume` / evaluation on the
|
||||
Mac) and the exported JSON policy (immediately playable — point Match or
|
||||
Spectate mode at `res://bots/<exp>.json`). Add a third rsync of `logs/` if
|
||||
you also want the TensorBoard history archived on the Mac.
|
||||
On the Mac (or anywhere), collecting the results is just `git pull`. A
|
||||
20M-step run adds roughly 40 MB of checkpoints — acceptable growth for the
|
||||
guarantee that training is never lost with a machine.
|
||||
|
||||
## Dashboard over the network
|
||||
|
||||
|
||||
Reference in New Issue
Block a user