chore(training): close Stage 5 by human override, re-derive its air-touch gate

productive_air_touch_episode_fraction's 0.02 floor was set as an explicit
PROVISIONAL guess (see the Round 10 comment in generation5.py) with
instructions to re-derive it from attempt 1's measured tail. That never
happened: five more Stage-5 attempts (20260824 through -retry4) ran against
the unchanged number, reading 0.00004/0.00006/0.00002/0.00018/0.00006 -- no
trend, ~500x under the floor -- while every other gate passed comfortably and
each attempt beat the Stage-4 reference head-to-head. Direct TensorBoard
query of retry4's full run confirms the touches are real and stable, just
rare (22/1000 rollout-logging windows registered one touch in the
~100-episode buffer), so further identical retries were not going to close a
500x gap.

Lowered the floor to 0.00002 (the minimum of the five measured attempts),
same as-under-the-observed-band logic the Stage-4 override used for
goal_rate. Flipped retry4's log entry to decision: pass with a
decision_override block (same pattern as the Stage-4 override) and advanced
generation5_state.json to Stage 6 attempt 0. Documented in TRAINING.md and
flagged Stage 6's own 0.015 floor for the same metric as equally unvalidated.
This commit is contained in:
Josh Creek
2026-08-29 16:45:09 +01:00
parent b946f78d1f
commit 4f13b4eca9
3 changed files with 85 additions and 11 deletions
+28 -7
View File
@@ -396,13 +396,34 @@ STAGES = [
# form asks the question the bar actually means: did this episode
# contain a productive aerial at all?
#
# 0.02 is PROVISIONAL and deliberately low. There is no measured
# baseline to derive it from — the metric reads 0.0 on retry2's
# checkpoint — and setting an unachievable bar from arithmetic
# rather than measurement is precisely what cost this stage nine
# attempts. Treat attempt 1 as establishing the real distribution
# and re-derive this from its tail before trusting it as a gate.
"rollout/productive_air_touch_episode_fraction": 0.02,
# Round 11 (2026-08-29): the 0.02 above was never re-derived, and
# the comment that set it said explicitly to do that after
# attempt 1. Five more attempts (20260824 through -retry4) ran
# against it unchanged: 0.00004, 0.00006, 0.00002, 0.00018,
# 0.00006 -- no trend, all within one order of magnitude of each
# other and roughly 500x under the floor. rollout/air_touch_
# fraction over retry4's full run confirms this is real signal
# rather than a broken metric (22 of 1000 rollout-logging windows
# registered exactly one aerial touch in the ~100-episode SB3
# buffer) -- just a rare event at this training-time drill mix,
# not a growing one. Every other gate cleared comfortably on all
# five attempts (retry4: goal_rate 0.796 vs 0.72, upright 0.778
# vs 0.40, forward_motion 0.493 vs 0.20) and every attempt beat
# the Stage-4 reference head-to-head (retry4: 53-26-21, sides
# 29-11 / 24-15). Lowered to 0.00002 -- the minimum of the five
# measured attempts, same "just under the observed band" logic
# Stage 4's own override used for goal_rate (see TRAINING.md) --
# so this floor now tests for regression against real behaviour
# instead of an unvalidated guess. retry4 closed Stage 5 by
# human override under the corrected floor rather than a sixth
# identical retry; see TRAINING.md and generation5_state.json's
# decision_override on that entry.
#
# Stage 6's 0.015 below carries the exact same provisional-guess
# problem and has never run a single attempt. Re-derive it from
# measured data the same way once Stage 6 actually produces a
# tail -- don't assume it transfers from this number.
"rollout/productive_air_touch_episode_fraction": 0.00002,
},
"evaluation_goal_rate_floor": 0.75,
"physical_side_imbalance_ceiling": 0.20,
+12 -4
View File
@@ -1,7 +1,7 @@
{
"stage_index": 1,
"attempt": 4,
"status": "blocked",
"stage_index": 2,
"attempt": 0,
"status": "in_progress",
"log": [
{
"stage_index": 0,
@@ -1185,7 +1185,15 @@
"win_rate_a": 0.53
}
],
"decision": "fail"
"decision": "pass",
"decision_override": {
"date": "2026-08-29",
"original_decision": "fail",
"original_failures": [
"rollout/productive_air_touch_episode_fraction=0.0001 < 0.0200"
],
"reason": "Human override. Only miss across this and four earlier attempts was rollout/productive_air_touch_episode_fraction, against a 0.02 floor the stage's own comment flagged as PROVISIONAL and said to re-derive from attempt 1's tail -- that step was never done. Attempt 1 (20260824) already measured 0.00004, ~500x under the floor, and the next three attempts (0.00006, 0.00002, 0.00018) plus this one (0.00006) show no trend toward 0.02, just noise at the same order of magnitude. rollout/air_touch_fraction over this attempt's full run confirms the touches are real, just rare (22/1000 rollout-logging windows registered exactly one aerial touch in the ~100-episode buffer). Every other gate passed comfortably: goal_rate 0.7962 vs 0.72, upright_fraction 0.7778 vs 0.40, forward_motion_fraction 0.4935 vs 0.20, and the paired evaluation beat Stage 4 53-26-21 with balanced sides (29-11 as team 0, 24-15 as team 1). Floor lowered 0.02 -> 0.00002 (the minimum of the five measured attempts) in generation5.py so it now tests real regression instead of an unvalidated guess. Stage 6 resumes from this checkpoint."
}
}
]
}