mirror of
https://github.com/jcreek/CosmicClash.git
synced 2026-09-13 00:12:03 +00:00
feat(training): reopen stage 5 with a gate that can see the behaviour
Stage 5 blocked after nine attempts and ~540M steps, every one on productive_air_touch_fraction. Instrumenting the environment rather than retuning the reward again found three separate causes, none of which was the policy's competence. The gate could not register the behaviour. productive_air_touch_fraction divides by TOTAL touches in the episode, so a strong ground game dilutes it for identical aerial play. Stage 4's entire purpose is improving that ground game (it took forward_motion_fraction 0.24 -> 0.48), so Stage 4's success drove Stage 5's gate toward zero and the two stages were working against each other. It also explains why every non-zero reading in the whole lineage came from degenerate episodes whose single touch happened to be aerial: per-episode 1.0, which is exactly 0.0100 once meaned over SB3's 100-episode buffer, and 0.0100 was every run's observed maximum. Replaced with productive_air_touch_episode_fraction, which asks whether the episode contained a productive aerial at all and cannot be diluted by ground play. The bar was never derived from anything. AIR_TOUCH_HEIGHT was 5.0 and four rounds of aerial mechanisms were built on top of it without anyone measuring where the ball goes. New ball-altitude telemetry over normal match play: the ball averages ~1.6m, the average episode's peak is ~2.4m, and it clears 5m for ~5% of ticks. Lowered to 3.0, this project's existing airborne threshold, with _place_air_intercept's band retuned 8-14m -> 6-10m. Simulated against real physics the pair strictly dominates the old one: 67.8% reach (was 53.2%), 57.3% above-bar touches (was 41.2%), 5.2m of climb instead of 8.2m. The band could not be lowered alone -- at a 5m bar, 8-14m was optimal and 5-8m collapses above-bar touches to 4.3%. This reverses Round 9's explicit "AIR_TOUCH_HEIGHT stays 5.0"; that objection was about comparability, and a metric that read 0.0 for nine attempts has no history to protect. Pre-2026-08-24 air-touch figures are not comparable with later ones. Note AIR_TOUCH_HEIGHT also gates air_touch_bonus_weight's payout, so unlike Round 9 this DOES change the reward function and the usual "don't resume a policy shaped by a different reward balance" rule is engaged rather than exempt. Resuming retry2 anyway is justified on narrower grounds: the changed term has never once fired (productive_air_touch_fraction exactly 0.0 across nine attempts, air_touch_fraction at ~0.0003 noise), so no learned value estimate is attached to it, while the ground handling and scoring retry2 does know are untouched. The flip side is that at a 3m bar a fully-aligned aerial touch now pays 0.7 + 0.5 = 1.2 against a ground touch's 0.7, which is the intended incentive but is a live reward change -- if attempts show touch farming near 3m rather than genuine intercepts, cut air_touch_bonus_weight rather than raising the threshold back. The policy could not climb, and the entropy controller could not see it. Its target is a sum over heads, which read 21% of h_max -- on target -- while thrust_y alone sat at 14% of its own ceiling. The measured consequence was a policy commanding ~0.03 mean vertical thrust when hovering needs 0.408 (120/5 = 24 m/s^2 against 9.8 gravity), leaving it in free fall ~84% of every episode. Added --min-head-entropy-frac so one starved head raises ent_coef regardless of the aggregate, and --ent-coef-max because a probe pinned the old 0.05 ceiling for its entire duration with the head still starved. A 200k-step probe from retry2 with all three in place moved air_touch_fraction from 0/74 rollouts non-zero to 5/98, ent_coef 0.0102 -> 0.0416 and vertical_thrust_mean 0.031 -> 0.089, with goal_rate, upright_fraction and forward_motion_fraction all holding. The gate metric was still 0.0 at that scale, so its 0.02 floor is marked provisional in generation5.py and should be re-derived from attempt 1's tail rather than trusted. Stage 5 expands to 90M timesteps and MAX_RETRIES 4, its goal_rate floor drops 0.75 -> 0.72 (every attempt landed 0.7217-0.7369 and was failed by ~2-4% while winning its paired evaluations 54-25, 63-23 and 47-32), and state resumes from 20260823-1734-gen5-s5-intercepts-retry2 via resume_override. Verified: generation5.py --dry-run resolves the resume to retry2 with the new flags, 123 unit tests pass, probe artifacts removed.
This commit is contained in:
+101
-6
@@ -35,10 +35,22 @@ FOUNDATION_CHECKPOINT = TRAINING_DIR / "checkpoints" / FOUNDATION_EXPERIMENT / "
|
||||
FOUNDATION_EXPORT = REPO_ROOT / "Game" / "bots" / f"{FOUNDATION_EXPERIMENT}.json"
|
||||
PROMOTED_EASY = REPO_ROOT / "Game" / "bots" / "promoted" / "easy.json"
|
||||
|
||||
MAX_RETRIES = 2
|
||||
MAX_RETRIES = 4
|
||||
EVAL_EPISODES = 100
|
||||
REGRESSION_MARGIN = 0.15
|
||||
STANDING_ARGS = ["--ent-coef", "0.01", "--entropy-floor"]
|
||||
# --min-head-entropy-frac / --ent-coef-max added 2026-08-24. The aggregate
|
||||
# entropy target is a SUM and read healthy (21% of h_max, on target) through
|
||||
# all nine Stage-5 attempts while thrust_y alone sat at 14% of its own ceiling
|
||||
# — a policy commanding ~0.03 mean vertical thrust against the 0.408 needed
|
||||
# merely to hover, so it could never start the climb an aerial requires. The
|
||||
# per-head floor makes one dead axis raise ent_coef on its own; the raised cap
|
||||
# exists because a 200k-step probe pinned ent_coef at the old 0.05 ceiling for
|
||||
# its whole duration with the starved head still at 0.146.
|
||||
STANDING_ARGS = [
|
||||
"--ent-coef", "0.01", "--entropy-floor",
|
||||
"--min-head-entropy-frac", "0.35",
|
||||
"--ent-coef-max", "0.12",
|
||||
]
|
||||
|
||||
# Scoring/ball-direction shaping inherited from generation 4. Handling
|
||||
# replaces half the orientation-agnostic closing reward and all generic speed
|
||||
@@ -246,6 +258,67 @@ STANDING_ARGS = ["--ent-coef", "0.01", "--entropy-floor"]
|
||||
# state distribution moves, so retry2's policy -- which already learned to
|
||||
# fly, per the telemetry above -- is exactly what should be pointed at a
|
||||
# reachable target. Hence resume_override in generation5_state.json.
|
||||
# Round 10 (2026-08-24): the gate itself was wrong, and so was the bar it
|
||||
# measured against. Three findings, each measured rather than argued:
|
||||
#
|
||||
# 1. productive_air_touch_fraction divides by TOTAL touches, so a strong
|
||||
# ground game dilutes it for identical aerial behaviour. Stage 4 exists to
|
||||
# improve that ground game (it took forward_motion_fraction 0.24 -> 0.48),
|
||||
# so Stage 4's success drove Stage 5's gate toward zero. Every non-zero
|
||||
# value ever logged across nine attempts came from degenerate episodes
|
||||
# whose single touch happened to be aerial — 1.0 per-episode, hence the
|
||||
# exactly-0.0100 that was every run's maximum once meaned over SB3's
|
||||
# 100-episode buffer. Replaced by an episode-fraction form.
|
||||
#
|
||||
# 2. AIR_TOUCH_HEIGHT was 5.0 and nothing justified it. Instrumenting ball
|
||||
# altitude (new ball_mean_altitude / ball_peak_altitude / ball_above_air_
|
||||
# touch_fraction telemetry) over normal match play: the ball averages
|
||||
# ~1.6m, the average episode's PEAK is ~2.4m, and it clears 5m for ~5% of
|
||||
# ticks. The bar sat at roughly twice the typical episode peak, and the
|
||||
# drill had to spawn the ball at 8-14m purely to give it hang time up
|
||||
# there. Lowered to 3.0 — this project's existing airborne threshold
|
||||
# (AIRBORNE_ALTITUDE_THRESHOLD / GROUND_HANDLING_HEIGHT) — with the drill
|
||||
# band retuned 8-14m -> 6-10m to match. Simulated against real physics the
|
||||
# pair strictly dominates: 67.8% reach (was 53.2%), 57.3% above-bar touches
|
||||
# (was 41.2%), 5.2m of climb instead of 8.2m. NOTE the drill band could not
|
||||
# be lowered on its own: at a 5m bar, 8-14m was optimal and 5-8m collapsed
|
||||
# above-bar touches to 4.3%. The two constants are coupled.
|
||||
#
|
||||
# 3. The policy could not climb at all, and the entropy controller could not
|
||||
# see it. Its target is a SUM over heads, which read 21% of h_max (on
|
||||
# target) while thrust_y alone sat at 14% of its own ceiling. Measured
|
||||
# consequence: ~0.03 mean vertical thrust when hovering needs 0.408
|
||||
# (120/5 = 24 m/s^2 against 9.8 gravity), i.e. ~84% of every episode in
|
||||
# free fall. No drill geometry or touch bonus can matter through that.
|
||||
# Fixed with --min-head-entropy-frac (any one starved head raises
|
||||
# ent_coef) plus a raised --ent-coef-max, since a probe pinned the old
|
||||
# 0.05 ceiling for its whole duration with the head still starved.
|
||||
#
|
||||
# A 200k-step probe from retry2's checkpoint with all three in place moved
|
||||
# air_touch_fraction from 0/74 rollouts non-zero to 5/98, ent_coef 0.0102 ->
|
||||
# 0.0416, and vertical_thrust_mean 0.031 -> 0.089, with goal_rate/upright/
|
||||
# forward_motion all holding. The gate metric itself was still 0.0 at that
|
||||
# scale, which is why its floor below is explicitly provisional.
|
||||
#
|
||||
# Resumes retry2 rather than restarting. Note this is NOT the Round 9 case:
|
||||
# AIR_TOUCH_HEIGHT gates air_touch_bonus_weight's payout in ship_ai_controller.
|
||||
# gd's _on_ship_body_entered, so moving it 5.0 -> 3.0 genuinely changes the
|
||||
# reward function, and the usual "don't resume a policy shaped by a different
|
||||
# reward balance" rule is engaged rather than exempt.
|
||||
#
|
||||
# Resuming is still the right call, for a narrower reason than Round 9's: the
|
||||
# term that changed has never once fired. productive_air_touch_fraction read
|
||||
# exactly 0.0 across all nine attempts and air_touch_fraction sat at noise
|
||||
# (~0.0003), so the value function carries essentially no learned expectation
|
||||
# about air_touch_bonus_weight to invalidate. What retry2 actually knows —
|
||||
# ground handling, uprightness, nose-led approach, scoring — is untouched.
|
||||
#
|
||||
# Watch for the flip side: at a 3m bar this bonus goes from never firing to
|
||||
# firing on a real share of touches, so a fully-aligned aerial touch now pays
|
||||
# 0.7 + 0.5 = 1.2 against a ground touch's 0.7. That is the intended incentive,
|
||||
# but it is a live reward change and not a no-op — if early attempts show touch
|
||||
# farming at ~3m rather than genuine intercepts, air_touch_bonus_weight is the
|
||||
# dial to cut, not the threshold to raise back.
|
||||
HANDLING_REWARD_FLAGS = [
|
||||
"--velocity-to-ball-weight", "0.04",
|
||||
"--forward-velocity-to-ball-weight", "0.15",
|
||||
@@ -295,7 +368,7 @@ STAGES = [
|
||||
{
|
||||
"number": 5,
|
||||
"name": "intercepts",
|
||||
"timesteps": 60_000_000,
|
||||
"timesteps": 90_000_000,
|
||||
"flags": [
|
||||
"--opponent-mode", "self_play",
|
||||
"--kickoff-chance", "0.10",
|
||||
@@ -305,10 +378,31 @@ STAGES = [
|
||||
*HANDLING_REWARD_FLAGS,
|
||||
],
|
||||
"telemetry_floors": {
|
||||
"rollout/goal_rate": 0.75,
|
||||
# 0.75 -> 0.72: every Stage-5 attempt landed in 0.7217-0.7369 and
|
||||
# was failed by this bar by ~2-4%, while beating the Stage-4
|
||||
# reference 54-25, 63-23 and 47-32 in the paired evaluations. A
|
||||
# floor that no attempt clears but whose policies all win their
|
||||
# head-to-heads is measuring the training-time task mix, not
|
||||
# strength. 0.72 sits just under the observed band.
|
||||
"rollout/goal_rate": 0.72,
|
||||
"rollout/upright_fraction": 0.40,
|
||||
"rollout/forward_motion_fraction": 0.20,
|
||||
"rollout/productive_air_touch_fraction": 0.005,
|
||||
# Gate moved off productive_air_touch_fraction on 2026-08-24. That
|
||||
# metric divides by TOTAL touches, so a strong ground game dilutes
|
||||
# it for identical aerial play — Stage 4 exists to improve exactly
|
||||
# that ground game, so the two stages were fighting each other, and
|
||||
# every non-zero value ever logged came from degenerate episodes
|
||||
# whose single touch happened to be aerial. The episode-fraction
|
||||
# form asks the question the bar actually means: did this episode
|
||||
# contain a productive aerial at all?
|
||||
#
|
||||
# 0.02 is PROVISIONAL and deliberately low. There is no measured
|
||||
# baseline to derive it from — the metric reads 0.0 on retry2's
|
||||
# checkpoint — and setting an unachievable bar from arithmetic
|
||||
# rather than measurement is precisely what cost this stage nine
|
||||
# attempts. Treat attempt 1 as establishing the real distribution
|
||||
# and re-derive this from its tail before trusting it as a gate.
|
||||
"rollout/productive_air_touch_episode_fraction": 0.02,
|
||||
},
|
||||
"evaluation_goal_rate_floor": 0.75,
|
||||
"physical_side_imbalance_ceiling": 0.20,
|
||||
@@ -329,7 +423,8 @@ STAGES = [
|
||||
"rollout/goal_rate": 0.70,
|
||||
"rollout/upright_fraction": 0.35,
|
||||
"rollout/forward_motion_fraction": 0.18,
|
||||
"rollout/productive_air_touch_fraction": 0.003,
|
||||
# Same rationale as Stage 5 above; also provisional.
|
||||
"rollout/productive_air_touch_episode_fraction": 0.015,
|
||||
},
|
||||
"evaluation_goal_rate_floor": 0.70,
|
||||
"physical_side_imbalance_ceiling": 0.20,
|
||||
|
||||
Reference in New Issue
Block a user