Commit Graph

183 Commits

Author SHA1 Message Date
Josh Creek 00b900d864 Merge pull request #30 from jcreek/feat/multiplayer
Feat/multiplayer
2026-09-06 10:58:50 +01:00
CosmicClash Training Bot 6983ddd7df chore(training): generation 5 progress after 20260903-1146-gen5-s6-league-retry3 2026-09-05 01:01:25 +01:00
CosmicClash Training Bot 7d247bc516 chore(training): Add 20260903-1146-gen5-s6-league-retry3 checkpoints, logs, and exported policy 2026-09-05 00:55:28 +01:00
CosmicClash Training Bot 4519f2db82 chore(training): generation 5 progress after 20260901-2301-gen5-s6-league-retry2 2026-09-03 11:46:05 +01:00
CosmicClash Training Bot 0f1a7403e4 chore(training): Add 20260901-2301-gen5-s6-league-retry2 checkpoints, logs, and exported policy 2026-09-03 11:40:08 +01:00
CosmicClash Training Bot 8a55c33666 chore(training): generation 5 progress after 20260831-0724-gen5-s6-league-retry1 2026-09-01 23:01:35 +01:00
CosmicClash Training Bot ad6f9cc148 chore(training): Add 20260831-0724-gen5-s6-league-retry1 checkpoints, logs, and exported policy 2026-09-01 22:55:32 +01:00
Josh Creek e56850a236 fix(training): make policy evaluation portable 2026-09-01 18:56:31 +01:00
Josh Creek e376675fa6 feat(training): add wall and rebound curriculum states 2026-09-01 17:37:52 +01:00
Josh Creek 4c1ed87344 test(training): verify 2v2 evaluator command 2026-09-01 17:32:10 +01:00
Josh Creek 7b2f9c26f4 feat(training): add opt-in teamplay evaluation 2026-09-01 17:30:40 +01:00
Josh Creek 9004800326 test(training): require multi-seed curriculum evaluation 2026-09-01 17:26:44 +01:00
CosmicClash Training Bot c51d5ee369 chore(training): generation 5 progress after 20260829-1649-gen5-s6-league 2026-08-31 07:24:53 +01:00
CosmicClash Training Bot 6d837b2bf3 chore(training): Add 20260829-1649-gen5-s6-league checkpoints, logs, and exported policy 2026-08-31 07:19:18 +01:00
Josh Creek 4f13b4eca9 chore(training): close Stage 5 by human override, re-derive its air-touch gate
productive_air_touch_episode_fraction's 0.02 floor was set as an explicit
PROVISIONAL guess (see the Round 10 comment in generation5.py) with
instructions to re-derive it from attempt 1's measured tail. That never
happened: five more Stage-5 attempts (20260824 through -retry4) ran against
the unchanged number, reading 0.00004/0.00006/0.00002/0.00018/0.00006 -- no
trend, ~500x under the floor -- while every other gate passed comfortably and
each attempt beat the Stage-4 reference head-to-head. Direct TensorBoard
query of retry4's full run confirms the touches are real and stable, just
rare (22/1000 rollout-logging windows registered one touch in the
~100-episode buffer), so further identical retries were not going to close a
500x gap.

Lowered the floor to 0.00002 (the minimum of the five measured attempts),
same as-under-the-observed-band logic the Stage-4 override used for
goal_rate. Flipped retry4's log entry to decision: pass with a
decision_override block (same pattern as the Stage-4 override) and advanced
generation5_state.json to Stage 6 attempt 0. Documented in TRAINING.md and
flagged Stage 6's own 0.015 floor for the same metric as equally unvalidated.
2026-08-29 16:45:09 +01:00
CosmicClash Training Bot b946f78d1f chore(training): generation 5 progress after 20260828-0214-gen5-s5-intercepts-retry4 2026-08-29 00:16:31 +01:00
CosmicClash Training Bot ca17265bf1 chore(training): Add 20260828-0214-gen5-s5-intercepts-retry4 checkpoints, logs, and exported policy 2026-08-29 00:14:04 +01:00
CosmicClash Training Bot d235342889 chore(training): generation 5 progress after 20260827-0426-gen5-s5-intercepts-retry3 2026-08-28 02:14:11 +01:00
CosmicClash Training Bot 706ac9bd2c chore(training): Add 20260827-0426-gen5-s5-intercepts-retry3 checkpoints, logs, and exported policy 2026-08-28 02:11:26 +01:00
CosmicClash Training Bot 5de72b0636 chore(training): generation 5 progress after 20260826-0631-gen5-s5-intercepts-retry2 2026-08-27 04:26:36 +01:00
CosmicClash Training Bot 295be3d26a chore(training): Add 20260826-0631-gen5-s5-intercepts-retry2 checkpoints, logs, and exported policy 2026-08-27 04:24:16 +01:00
CosmicClash Training Bot 3af7ed077c chore(training): generation 5 progress after 20260825-0835-gen5-s5-intercepts-retry1 2026-08-26 06:31:14 +01:00
CosmicClash Training Bot e364a7dd06 chore(training): Add 20260825-0835-gen5-s5-intercepts-retry1 checkpoints, logs, and exported policy 2026-08-26 06:28:51 +01:00
CosmicClash Training Bot d06a67ade6 chore(training): generation 5 progress after 20260824-1052-gen5-s5-intercepts 2026-08-25 08:35:28 +01:00
CosmicClash Training Bot 6d94366693 chore(training): Add 20260824-1052-gen5-s5-intercepts checkpoints, logs, and exported policy 2026-08-25 08:32:51 +01:00
Josh Creek cb06300685 feat(training): reopen stage 5 with a gate that can see the behaviour
Stage 5 blocked after nine attempts and ~540M steps, every one on
productive_air_touch_fraction. Instrumenting the environment rather than
retuning the reward again found three separate causes, none of which was the
policy's competence.

The gate could not register the behaviour. productive_air_touch_fraction
divides by TOTAL touches in the episode, so a strong ground game dilutes it for
identical aerial play. Stage 4's entire purpose is improving that ground game
(it took forward_motion_fraction 0.24 -> 0.48), so Stage 4's success drove
Stage 5's gate toward zero and the two stages were working against each other.
It also explains why every non-zero reading in the whole lineage came from
degenerate episodes whose single touch happened to be aerial: per-episode 1.0,
which is exactly 0.0100 once meaned over SB3's 100-episode buffer, and 0.0100
was every run's observed maximum. Replaced with
productive_air_touch_episode_fraction, which asks whether the episode contained
a productive aerial at all and cannot be diluted by ground play.

The bar was never derived from anything. AIR_TOUCH_HEIGHT was 5.0 and four
rounds of aerial mechanisms were built on top of it without anyone measuring
where the ball goes. New ball-altitude telemetry over normal match play: the
ball averages ~1.6m, the average episode's peak is ~2.4m, and it clears 5m for
~5% of ticks. Lowered to 3.0, this project's existing airborne threshold, with
_place_air_intercept's band retuned 8-14m -> 6-10m. Simulated against real
physics the pair strictly dominates the old one: 67.8% reach (was 53.2%), 57.3%
above-bar touches (was 41.2%), 5.2m of climb instead of 8.2m. The band could
not be lowered alone -- at a 5m bar, 8-14m was optimal and 5-8m collapses
above-bar touches to 4.3%. This reverses Round 9's explicit "AIR_TOUCH_HEIGHT
stays 5.0"; that objection was about comparability, and a metric that read 0.0
for nine attempts has no history to protect. Pre-2026-08-24 air-touch figures
are not comparable with later ones.

Note AIR_TOUCH_HEIGHT also gates air_touch_bonus_weight's payout, so unlike
Round 9 this DOES change the reward function and the usual "don't resume a
policy shaped by a different reward balance" rule is engaged rather than exempt.
Resuming retry2 anyway is justified on narrower grounds: the changed term has
never once fired (productive_air_touch_fraction exactly 0.0 across nine
attempts, air_touch_fraction at ~0.0003 noise), so no learned value estimate is
attached to it, while the ground handling and scoring retry2 does know are
untouched. The flip side is that at a 3m bar a fully-aligned aerial touch now
pays 0.7 + 0.5 = 1.2 against a ground touch's 0.7, which is the intended
incentive but is a live reward change -- if attempts show touch farming near 3m
rather than genuine intercepts, cut air_touch_bonus_weight rather than raising
the threshold back.

The policy could not climb, and the entropy controller could not see it. Its
target is a sum over heads, which read 21% of h_max -- on target -- while
thrust_y alone sat at 14% of its own ceiling. The measured consequence was a
policy commanding ~0.03 mean vertical thrust when hovering needs 0.408
(120/5 = 24 m/s^2 against 9.8 gravity), leaving it in free fall ~84% of every
episode. Added --min-head-entropy-frac so one starved head raises ent_coef
regardless of the aggregate, and --ent-coef-max because a probe pinned the old
0.05 ceiling for its entire duration with the head still starved.

A 200k-step probe from retry2 with all three in place moved air_touch_fraction
from 0/74 rollouts non-zero to 5/98, ent_coef 0.0102 -> 0.0416 and
vertical_thrust_mean 0.031 -> 0.089, with goal_rate, upright_fraction and
forward_motion_fraction all holding. The gate metric was still 0.0 at that
scale, so its 0.02 floor is marked provisional in generation5.py and should be
re-derived from attempt 1's tail rather than trusted.

Stage 5 expands to 90M timesteps and MAX_RETRIES 4, its goal_rate floor drops
0.75 -> 0.72 (every attempt landed 0.7217-0.7369 and was failed by ~2-4% while
winning its paired evaluations 54-25, 63-23 and 47-32), and state resumes from
20260823-1734-gen5-s5-intercepts-retry2 via resume_override.

Verified: generation5.py --dry-run resolves the resume to retry2 with the new
flags, 123 unit tests pass, probe artifacts removed.
2026-08-24 10:49:05 +01:00
Josh Creek e1f512c94e feat(bots): promote gen5 stage-5 policy to the Hard tier
Hard has been a label-only duplicate of medium.json since medium was promoted
on 2026-08-17. Promote 20260823-1734-gen5-s5-intercepts-retry2 into
hard.json so the tier is a genuinely distinct policy, and so the strongest bot
the curriculum has produced survives the next round's checkpoint pruning —
promoted files are never touched by training scripts.

Stage 5 blocked after three attempts, so like medium.json this comes from a run
recorded as decision: "fail". Both failing floors are covered in TRAINING.md:
goal_rate 0.7369 vs 0.75 is marginal, and productive_air_touch_fraction 0.0001
vs 0.005 is a bar no policy in the lineage has approached, against a metric
quantised at 0.01 per ~100-episode window. On every other axis it is the best
yet: upright_fraction 0.757 against a 0.40 floor that the pre-Round-6 lineage
never pushed past 0.331, and forward_motion_fraction 0.479 against 0.20.

Chosen over attempt 2 (retry1) on a tiebreak, not a margin. retry1 posts a much
wider indirect result against medium.json (63-23-14 vs 47-32-21), but a direct
100-episode head-to-head between the two finished 36-39 with 25 draws, so that
gap does not reflect a real strength difference. Attempt 3 is the later
checkpoint (it resumed from attempt 2) and edges every telemetry metric.

Verified: hard.json is byte-identical to its source export, matches easy/medium
on input_size 83, 3 layers and action space, and beats medium.json 19-7-4 in a
fresh 30-episode paired run. Tiers stay monotonic: hard > medium > easy.

That head-to-head also showed a 17% physical side imbalance (physical teams
0-1 = 29-46), reproduced at 13% in the 30-episode check. Inside the 20% bar
used elsewhere and equal across both models, but noted in TRAINING.md as worth
investigating rather than assuming variance.
2026-08-24 08:46:06 +01:00
CosmicClash Training Bot dffc2812e1 chore(training): generation 5 progress after 20260823-1734-gen5-s5-intercepts-retry2 2026-08-24 08:07:11 +01:00
CosmicClash Training Bot 614ec9cda9 chore(training): Add 20260823-1734-gen5-s5-intercepts-retry2 checkpoints, logs, and exported policy 2026-08-24 08:04:54 +01:00
CosmicClash Training Bot 17f588b95b chore(training): generation 5 progress after 20260823-0258-gen5-s5-intercepts-retry1 2026-08-23 17:34:58 +01:00
CosmicClash Training Bot 7c949d8679 chore(training): Add 20260823-0258-gen5-s5-intercepts-retry1 checkpoints, logs, and exported policy 2026-08-23 17:32:48 +01:00
CosmicClash Training Bot ba1887fe93 chore(training): generation 5 progress after 20260822-1242-gen5-s5-intercepts 2026-08-23 02:58:22 +01:00
CosmicClash Training Bot b83e030a44 chore(training): Add 20260822-1242-gen5-s5-intercepts checkpoints, logs, and exported policy 2026-08-23 02:55:57 +01:00
CosmicClash Training Bot 5fcdb256b3 chore(training): restore resume_override for stage-5 retry2 after crashed push 2026-08-22 12:42:13 +01:00
CosmicClash Training Bot b69291a7d3 chore(training): Add 20260821-1516-gen5-s5-intercepts checkpoints, logs, and exported policy 2026-08-22 11:50:40 +01:00
Josh Creek 818f8e89cd fix(training): make the stage-5 air-intercept drill physically solvable
productive_air_touch_fraction sat at exactly 0.0 across nine Stage-5
attempts and 540M timesteps. Two rounds of reward shaping were aimed at
it (air_approach_weight, then air_touch_bonus_weight); both worked --
airborne_fraction 0.223->0.258, mean_altitude 2.59->3.25,
vertical_thrust_mean 0.004->0.063 -- and the ship now visibly plays the
ball in the air. The metric could not see it because it counts only
touches with the ball above AIR_TOUCH_HEIGHT (5m), and
_place_air_intercept never produced a reachable one.

Simulating the spawn distribution against the ship's flight envelope
(vertical_thrust 120 / mass 5 = 24 m/s^2 less gravity, drag capping
climb near 12 m/s): a ball spawned 6-12m up at 6-11 m/s is above 5m for
a median of 0.80s, while the ship spawned 7-13m behind, 3-10m below, and
at a dead stop. An ideal interceptor -- point mass, instant attitude, no
righting torque, zero reaction delay -- makes that touch in 0.00% of
episodes and reaches the ball at all in 0.5%.

Retune the drill instead of the reward: ball higher (8-14m) and slower
(4-8 m/s), ship closer (4-9m behind), narrower lateral spread, and a
6-14 m/s planar run-up rather than a standing start -- the dead stop was
the largest single factor. Ideal interceptor now reaches the ball in
~98% of episodes and above 5m in ~37%, so the 0.005 floor has headroom.
AIR_TOUCH_HEIGHT stays 5.0 so the metric remains comparable with earlier
generations.

Resume from retry2 rather than restarting from Stage 4: that rule guards
against a changed reward function invalidating the value function, and
the reward function is untouched here -- only the state distribution
moved, so the policy that already learned to fly is what should be
pointed at a reachable target. Adds a one-shot resume_override to
generation5_state.json, consumed on first use.
2026-08-21 15:14:36 +01:00
CosmicClash Training Bot 23e3dd18f9 chore(training): generation 5 progress after 20260821-0056-gen5-s5-intercepts-retry2 2026-08-21 13:49:28 +01:00
CosmicClash Training Bot 0699b14d4e chore(training): Add 20260821-0056-gen5-s5-intercepts-retry2 checkpoints, logs, and exported policy 2026-08-21 13:48:12 +01:00
CosmicClash Training Bot 9708bfafa3 chore(training): generation 5 progress after 20260820-1157-gen5-s5-intercepts-retry1 2026-08-21 00:56:33 +01:00
CosmicClash Training Bot 12bc4d7e8a chore(training): Add 20260820-1157-gen5-s5-intercepts-retry1 checkpoints, logs, and exported policy 2026-08-21 00:55:09 +01:00
CosmicClash Training Bot aa049aaff4 chore(training): generation 5 progress after 20260819-2307-gen5-s5-intercepts 2026-08-20 11:57:16 +01:00
CosmicClash Training Bot 7782d63660 chore(training): Add 20260819-2307-gen5-s5-intercepts checkpoints, logs, and exported policy 2026-08-20 11:55:54 +01:00
Josh Creek 602fa297d0 chore(training): add air_touch_bonus_weight and restart stage-5 intercepts
air_approach_weight alone didn't move productive_air_touch_fraction after a
further 180M steps (360M cumulative across all six Stage-5 attempts): an
unredirected air-intercept ball falls short of the goal from gravity and
just lands on the floor, so the already-solved ground game collects the
same episode reward whether or not anything touched the ball in the air.
air_touch_bonus_weight adds a conjunctive event bonus on top of
ball_touch_reward for a touch that's both genuinely aerial and
goal-directed, targeting the actual measured behaviour instead of only the
approach to it.
2026-08-19 22:46:04 +01:00
CosmicClash Training Bot 03f49e59c8 chore(training): generation 5 progress after 20260819-1321-gen5-s5-intercepts-retry2 2026-08-19 22:34:06 +01:00
CosmicClash Training Bot fde098c6c9 chore(training): Add 20260819-1321-gen5-s5-intercepts-retry2 checkpoints, logs, and exported policy 2026-08-19 22:32:53 +01:00
CosmicClash Training Bot 2f6e7b1201 chore(training): generation 5 progress after 20260819-0412-gen5-s5-intercepts-retry1 2026-08-19 13:21:42 +01:00
CosmicClash Training Bot 0db21b20a6 chore(training): Add 20260819-0412-gen5-s5-intercepts-retry1 checkpoints, logs, and exported policy 2026-08-19 13:20:33 +01:00
CosmicClash Training Bot 6fc2efacc6 chore(training): generation 5 progress after 20260818-1903-gen5-s5-intercepts 2026-08-19 04:12:42 +01:00
CosmicClash Training Bot 24b99ff8f9 chore(training): Add 20260818-1903-gen5-s5-intercepts checkpoints, logs, and exported policy 2026-08-19 04:11:29 +01:00
Josh Creek 88591e031f chore(training): add air_approach_weight and restart stage-5 intercepts
Stage 5 blocked all three attempts on productive_air_touch_fraction
stuck exactly at 0.0 across a continuous 180M-step lineage, while
goal_rate/upright_fraction/forward_motion_fraction kept improving on
the same budget. forward_velocity_to_ball_weight (the term that solved
Stage 4's ground pursuit) is hard-gated below GROUND_HANDLING_HEIGHT
and does nothing in the air, so Stage 5's air_intercept_chance had no
matching aerial incentive to learn from. air_approach_weight adds the
airborne mirror (nose-first 3D closing speed, no uprightness
multiplier) and folds into HANDLING_REWARD_FLAGS so Stage 6 inherits
it too. Deleted the three blocked attempts and reset state to resume
Stage 5 from the Stage-4 checkpoint with the new term.
2026-08-18 16:03:56 +01:00