Commit Graph

104 Commits

Author SHA1 Message Date
CosmicClash Training Bot 8e5eea4f46 chore(training): Add 20260812-0424-gen5-s4-handling-retry2 checkpoints, logs, and exported policy 2026-08-12 10:30:03 +01:00
CosmicClash Training Bot 052b47c04f chore(training): generation 5 progress after 20260811-2216-gen5-s4-handling-retry1 2026-08-12 04:24:21 +01:00
CosmicClash Training Bot 739c4b09d5 chore(training): Add 20260811-2216-gen5-s4-handling-retry1 checkpoints, logs, and exported policy 2026-08-12 04:22:59 +01:00
CosmicClash Training Bot 6a34eefceb chore(training): generation 5 progress after 20260811-1606-gen5-s4-handling 2026-08-11 22:16:10 +01:00
CosmicClash Training Bot 722ad7cc85 chore(training): Add 20260811-1606-gen5-s4-handling checkpoints, logs, and exported policy 2026-08-11 22:14:46 +01:00
Josh Creek b8e2a7b57a fix(training): cut grounded_upright_reward, restart stage-4 from Stage-3 foundation
grounded_upright_reward at 0.015 overshot: four force-retries pushed
upright_fraction from 0.265 to a plateauing 0.331, then the fifth jumped it
to 0.696 (55% over the 0.45 floor) while goal_rate collapsed 0.542->0.366
and forward_motion_fraction fell 0.244->0.184 (vertical_thrust_mean went
negative) - the policy learned to sit pinned upright and farm the bonus
instead of chasing the ball. It was sized "comparable to
time_penalty/ball_distance_penalty" but at 0.015/tick it was actually above
ball_distance_penalty's 0.01/tick worst case, so idling near the ball beat
playing. Cut to 0.004/tick (episode ceiling ~7.2, below
ball_distance_penalty's ~18 worst case). Delete the five blocked attempts
and reset generation5_state.json so the next run starts fresh from the
Stage-3 foundation rather than continuing from the farming checkpoint.
2026-08-11 15:57:27 +01:00
CosmicClash Training Bot f5b0a79cea chore(training): generation 5 progress after 20260811-0858-gen5-s4-handling-retry4 2026-08-11 15:12:24 +01:00
CosmicClash Training Bot 25560a65b5 chore(training): Add 20260811-0858-gen5-s4-handling-retry4 checkpoints, logs, and exported policy 2026-08-11 15:10:43 +01:00
CosmicClash Training Bot be1429e99e chore(training): generation 5 progress after 20260810-1338-gen5-s4-handling-retry3 2026-08-10 19:52:02 +01:00
CosmicClash Training Bot 651c048115 chore(training): Add 20260810-1338-gen5-s4-handling-retry3 checkpoints, logs, and exported policy 2026-08-10 19:50:40 +01:00
CosmicClash Training Bot d1ccf4dbc7 chore(training): generation 5 progress after 20260810-0211-gen5-s4-handling-retry2 2026-08-10 08:22:52 +01:00
CosmicClash Training Bot 677fff2c82 chore(training): Add 20260810-0211-gen5-s4-handling-retry2 checkpoints, logs, and exported policy 2026-08-10 08:21:17 +01:00
CosmicClash Training Bot dd1165c805 chore(training): generation 5 progress after 20260809-1955-gen5-s4-handling-retry1 2026-08-10 02:11:41 +01:00
CosmicClash Training Bot aa8d895b14 chore(training): Add 20260809-1955-gen5-s4-handling-retry1 checkpoints, logs, and exported policy 2026-08-10 02:10:10 +01:00
CosmicClash Training Bot 7174ff7cf9 chore(training): generation 5 progress after 20260809-1340-gen5-s4-handling 2026-08-09 19:55:13 +01:00
CosmicClash Training Bot dcf618dda1 chore(training): Add 20260809-1340-gen5-s4-handling checkpoints, logs, and exported policy 2026-08-09 19:53:46 +01:00
Josh Creek 6f7536f03c fix(training): correct non-forward penalty math and add a grounding incentive
Adversarial review of the previous stage-4 retune found two problems:
non_forward_speed used planar_speed - forward_component, which under-charges
diagonal motion relative to true lateral speed (e.g. ~29% penalty at 45
degrees off the nose instead of the correct ~71%); fixed to the Pythagorean
magnitude for forward-facing angles, full speed for backward-facing ones.

Also, ground_tilt_penalty and non_forward_penalty only ever cost reward near
the floor with nothing offsetting them above it, which could teach a policy
that's still bad at ground handling to just avoid the floor rather than get
better at it. Added grounded_upright_reward (ship_ai_controller.gd) plus a
new ShipObservations.is_floor_contact helper for genuine belly-on-floor
contact detection, so grounding well while upright is the locally profitable
choice, not just the least-punished one.
2026-08-09 13:23:00 +01:00
Josh Creek c56f5ed1a3 chore(training): retune stage-4 handling penalties and restart from Stage-3 foundation
Stage 4's upright/forward-motion telemetry plateaued flat across all three
blocked attempts because ground_tilt_penalty (0.003) was too weak to matter
and nothing penalized sideways/reverse motion at all. Raise
ground_tilt_penalty to 0.05 and add a new non_forward_penalty term
(ship_ai_controller.gd) that directly costs non-forward planar velocity near
the floor, independent of the ball. Delete the three blocked attempts'
checkpoints/logs/exports and reset generation5_state.json so the next run
starts fresh from the Stage-3 foundation checkpoint instead of continuing
from the drifted retry2 weights.
2026-08-09 13:09:12 +01:00
CosmicClash Training Bot 005cd0c66e chore(training): generation 5 progress after 20260809-0328-gen5-s4-handling-retry2 2026-08-09 09:36:44 +01:00
CosmicClash Training Bot c6f0f2084f chore(training): Add 20260809-0328-gen5-s4-handling-retry2 checkpoints, logs, and exported policy 2026-08-09 09:35:20 +01:00
CosmicClash Training Bot 837deedc08 chore(training): generation 5 progress after 20260808-2120-gen5-s4-handling-retry1 2026-08-09 03:28:48 +01:00
CosmicClash Training Bot 686b115b88 chore(training): Add 20260808-2120-gen5-s4-handling-retry1 checkpoints, logs, and exported policy 2026-08-09 03:27:31 +01:00
CosmicClash Training Bot d617c8032c chore(training): generation 5 progress after 20260808-1508-gen5-s4-handling 2026-08-08 21:20:09 +01:00
CosmicClash Training Bot 11b41ab562 chore(training): Add 20260808-1508-gen5-s4-handling checkpoints, logs, and exported policy 2026-08-08 21:18:39 +01:00
Josh Creek 341a67f6da feat(training): add generation 5 curriculum 2026-08-08 14:56:17 +01:00
Josh Creek 33952b3cd0 feat(training): pair evaluations across physical sides 2026-08-08 14:55:12 +01:00
Josh Creek 57a298dc06 fix(ai): map team-relative rotation actions 2026-08-08 14:54:27 +01:00
CosmicClash Training Bot b7f646ba48 chore(training): Add 20260806-1939-curric-s3-gauntlet checkpoints, logs, and exported policy 2026-08-08 08:24:30 +01:00
CosmicClash Training Bot 9fd1764522 chore(training): curriculum progress after 20260805-1926-curric-s2-selfplay 2026-08-06 19:39:02 +01:00
CosmicClash Training Bot a370b15020 chore(training): Add 20260805-1926-curric-s2-selfplay checkpoints, logs, and exported policy 2026-08-06 19:37:58 +01:00
CosmicClash Training Bot f8641836df chore(training): curriculum progress after 20260805-0953-curric-s1-bootstrap 2026-08-05 19:26:22 +01:00
CosmicClash Training Bot 9f78d37c79 chore(training): Add 20260805-0953-curric-s1-bootstrap checkpoints, logs, and exported policy 2026-08-05 19:26:20 +01:00
Josh Creek 0e046aa9f1 chore(training): scrap generation 1-3 training data for generation 4
All checkpoints/logs/exported policies here are a continuous-Gaussian,
31-input action/observation shape that generation 4's MultiDiscrete
redesign is structurally incompatible with -- nothing to resume from (see
the prior commit and TRAINING.md's "Generation 4" section). Game/bots/promoted/
(easy.json, and the new reference-grounded.json copied from
curric-s5-aggression before this) is untouched -- both remain valid,
playable evaluation opponents forever via PolicyNetwork's format-versioned
JSON despite their own checkpoints/generation being gone.

- training/checkpoints/*, training/logs/* removed (~3.5GB of working tree,
  all generation 1-3 experiment runs).
- Game/bots/*.json flat dump removed (superseded exports; main_menu.gd's
  Spectate dropdown will just be empty until the first generation-4 export).
- curriculum_state.json -> curriculum_state_gen3.json, archived alongside
  the existing _gen1/_gen2 logs (all three are referenced as postmortem
  evidence in TRAINING.md/curriculum.py). A fresh curriculum_state.json
  will be created on the next curriculum.py run (load_state() already
  handles a missing file).

training/eval_history.json is deliberately NOT reset -- it's the one
continuous cross-generation progress record.

NOT YET PUSHED: this needs the remote Linux training box quiesced first
(kill any active tmux session, confirm it's synced to origin) so its own
run_training.sh doesn't race a still-running job's final commit against
this deletion.
2026-08-04 23:28:57 +01:00
Josh Creek 1811e9333e feat(training): curriculum generation 4 — MultiDiscrete action space redesign
Three curriculum generations (2026-07-21 through 2026-08-04) all tried
gating *when* the policy could use vertical thrust/pitch-roll on top of a
continuous Gaussian action space, and all three failed the same way: PPO's
action-distribution std collapsed within ~10% of steps and never recovered,
landing at a 15-32% win rate vs the grounded reference regardless of
mechanism (hard mask, then a gradual ramp). Generation 3's final attempt
just landed at 24% — the worst of the three.

Root cause, verified against this project's own physics: hovering this ship
requires *holding* thrust.y ~= 0.408 continuously (mass 5.0, vertical_thrust
120, gravity 9.8). A collapsed near-zero-mean Gaussian can brush that value
but never sustain it long enough to earn the reward gradient that would
move the mean — no amount of gating *when* the axis acts fixes a problem in
*how* the policy represents a decision on it. This also independently found
and fixes a real bug: godot_rl never marks an episode timeout as a
truncation, so PPO was bootstrapping V(s)=0 on every 30s draw in every
generation to date.

- Game/scripts/ship_action_codec.gd (new): single source of truth for a
  per-axis MultiDiscrete action space (7 heads, nvec [5,5,5,5,5,5,2]) shared
  by training and in-game inference, replacing the continuous Gaussian.
  thrust_y's bins are deliberately asymmetric so a random policy drifts
  through the volume instead of floor-pinning. Legacy continuous decode
  (ai_ship_controller.gd's old logic) preserved verbatim so every
  pre-generation-4 export (e.g. Game/bots/promoted/easy.json) keeps working
  unchanged via an optional "action_space" JSON field.
- ship_observations.gd: append own contact state (SIZE 31 -> 35, append-only)
  so the value function can see what wall_contact_penalty fires on.
- ship_ai_controller.gd: action space/decode via the codec; drop the
  vertical_ramp/pitch_roll_ramp mechanism entirely; tilt_penalty default
  lowered 4x (aerial approaches require pitching); flight telemetry
  (airborne_fraction, mean_altitude, air_touch_fraction, vertical_thrust_mean)
  and truncation-snapshot fields on get_info().
- training_mode.gd: new air_drill_chance state-setter branch (ball spawned
  high, ships low, kept clear of walls) so aerial practice is forced by the
  environment instead of relying on reward-driven exploration alone; snapshot
  terminal observations before a timeout reset for the truncation fix.
- cosmic_env.py: remap ShipAIController's truncated/terminal_obs info into
  SB3's TimeLimit.truncated/terminal_observation keys.
- train.py: --reset-logits (+ --reset-logits-heads) replaces the
  now-meaningless --reset-std; new EntropyFloorCallback (a persistent
  per-rollout ent_coef controller replacing the one-shot std-reset shock)
  and per-head entropy logging; FlightTelemetryCallback; --air-drill-chance/
  --tilt-penalty flags; optional AbortIfCallback kill-criterion.
- export_policy.py: writes the action_space block for MultiDiscrete models;
  index-level parity check (argmax per head) instead of comparing floats.
- curriculum.py: full rewrite — 3 stages (bootstrap/selfplay/gauntlet), no
  grounded stage, full action space live from step 1; deletes generation
  1-3's checkpoint-lineage machinery (nothing to resume from); final report
  evaluates against both promoted/easy.json and the new
  promoted/reference-grounded.json (a copy of curric-s5-aggression, the
  strongest grounded-era artifact, kept as a fixed yardstick).
- run_training.sh/.gitignore: commit only final.zip, not the ~2400
  intermediate checkpoint files a single stage was writing (~500MB ->
  ~0.2MB per run); requirements.txt pinned (behaviour here now depends on
  specific library internals, not just public APIs).
- test_action_space.py (new): offline rung-0 check catching a head-order
  mismatch before it silently corrupts 24h of training.

Validated: GDScript compiles clean (Godot --headless --import + script
validation), free_play.tscn and training.tscn both boot headless without
errors, offline action-space assertions pass. Not yet run: the actual
smoke-training/A-B validation ladder steps in TRAINING.md's "Generation 4"
section, before committing to the full ~32h curriculum.

See TRAINING.md's "Generation 4" section for the full design writeup.
2026-08-04 23:27:57 +01:00
CosmicClash Training Bot 6198dc67fc chore(training): Add 20260803-1829-curric-s4-unmask-retry2 checkpoints, logs, and exported policy 2026-08-04 22:06:04 +01:00
CosmicClash Training Bot 20f6be7e28 chore(training): curriculum progress after 20260802-1458-curric-s4-unmask-retry1 2026-08-03 18:29:13 +01:00
CosmicClash Training Bot 5fe53b2406 chore(training): Add 20260802-1458-curric-s4-unmask-retry1 checkpoints, logs, and exported policy 2026-08-03 18:27:21 +01:00
CosmicClash Training Bot 62dc0a2981 chore(training): curriculum progress after 20260801-1131-curric-s4-unmask 2026-08-02 14:58:25 +01:00
CosmicClash Training Bot f62ddde369 chore(training): Add 20260801-1131-curric-s4-unmask checkpoints, logs, and exported policy 2026-08-02 14:56:35 +01:00
CosmicClash Training Bot fa53d72f63 chore(training): curriculum progress after 20260801-0658-curric-s3-unmask-ramp75 2026-08-01 11:31:57 +01:00
CosmicClash Training Bot e5a0df63c5 chore(training): Add 20260801-0658-curric-s3-unmask-ramp75 checkpoints, logs, and exported policy 2026-08-01 11:31:48 +01:00
CosmicClash Training Bot be1bf37e0b chore(training): curriculum progress after 20260801-0223-curric-s2-unmask-ramp50 2026-08-01 06:58:20 +01:00
CosmicClash Training Bot 1615791ee0 chore(training): Add 20260801-0223-curric-s2-unmask-ramp50 checkpoints, logs, and exported policy 2026-08-01 06:58:11 +01:00
CosmicClash Training Bot e22b4814f5 chore(training): curriculum progress after 20260731-2149-curric-s1-unmask-ramp25 2026-08-01 02:23:37 +01:00
CosmicClash Training Bot 23b2cd19df chore(training): Add 20260731-2149-curric-s1-unmask-ramp25 checkpoints, logs, and exported policy 2026-08-01 02:23:27 +01:00
Josh Creek 3fd1c00895 feat(training): Replace all-or-nothing unmask with a gradual ramp
Generation 2's single "unmask" stage (flip vertical/pitch-roll locomotion
from grounded-only to full 3D in one step) failed 3 independent 240M-step
attempts, landing at a stable 32% / 28% / 31% win rate vs curric-s5-aggression
each time -- not noise, and not fixable by more training time (attempts 2-3
each continued the same checkpoint lineage for another full 240M steps with
zero improvement). Every attempt shows train/std collapsing from ~0.30 to
~0.13-0.15 within the first ~10% of steps and never recovering: the policy
locks the newly-opened axes back down before ever meaningfully exploring
them.

Replaces the boolean allow_vertical/allow_pitch_roll mask on ShipAIController
with float vertical_ramp/pitch_roll_ramp multipliers (0.0-1.0), scaling axis
effect in set_action() instead of gating it outright -- the action space
never changes shape, so checkpoints stay resumable across ramp values. The
single unmask stage in curriculum.py becomes 4: three ungated warmup stages
(25%/50%/75% authority, airborne_penalty ramping in step) that train,
checkpoint, and always advance with no eval gate, then the measured stage at
full authority -- same reference, opponent mode, and 240M budget as the 3
failed attempts, for a direct comparison. Adds a "gated" flag/branch to
main()'s loop for the ungated stages.

This is generation 3 of the curriculum; generation 2's state is archived to
curriculum_state_gen2.json (mirroring the earlier gen1 -> gen2 archival) and
curriculum_state.json resets fresh, since its stage 0 no longer means what it
used to. See TRAINING.md's "Generation 3" section for the full postmortem,
stage table, and the open question about whether scaling action effect in
Godot (which PPO's own entropy/exploration math never sees) actually
addresses the collapse.
2026-07-31 21:47:53 +01:00
CosmicClash Training Bot 0759e1514b chore(training): curriculum progress after 20260730-1224-curric-s1-unmask-retry2 2026-07-31 16:07:16 +01:00
CosmicClash Training Bot f03037a612 chore(training): Add 20260730-1224-curric-s1-unmask-retry2 checkpoints, logs, and exported policy 2026-07-31 16:04:58 +01:00
CosmicClash Training Bot 6084991f1c chore(training): curriculum progress after 20260729-0837-curric-s1-unmask-retry1 2026-07-30 12:24:45 +01:00
CosmicClash Training Bot f41702e7ea chore(training): Add 20260729-0837-curric-s1-unmask-retry1 checkpoints, logs, and exported policy 2026-07-30 12:22:54 +01:00