mirror of
https://github.com/jcreek/CosmicClash.git
synced 2026-09-15 06:42:09 +00:00
feat(training): curriculum generation 4 — MultiDiscrete action space redesign
Three curriculum generations (2026-07-21 through 2026-08-04) all tried gating *when* the policy could use vertical thrust/pitch-roll on top of a continuous Gaussian action space, and all three failed the same way: PPO's action-distribution std collapsed within ~10% of steps and never recovered, landing at a 15-32% win rate vs the grounded reference regardless of mechanism (hard mask, then a gradual ramp). Generation 3's final attempt just landed at 24% — the worst of the three. Root cause, verified against this project's own physics: hovering this ship requires *holding* thrust.y ~= 0.408 continuously (mass 5.0, vertical_thrust 120, gravity 9.8). A collapsed near-zero-mean Gaussian can brush that value but never sustain it long enough to earn the reward gradient that would move the mean — no amount of gating *when* the axis acts fixes a problem in *how* the policy represents a decision on it. This also independently found and fixes a real bug: godot_rl never marks an episode timeout as a truncation, so PPO was bootstrapping V(s)=0 on every 30s draw in every generation to date. - Game/scripts/ship_action_codec.gd (new): single source of truth for a per-axis MultiDiscrete action space (7 heads, nvec [5,5,5,5,5,5,2]) shared by training and in-game inference, replacing the continuous Gaussian. thrust_y's bins are deliberately asymmetric so a random policy drifts through the volume instead of floor-pinning. Legacy continuous decode (ai_ship_controller.gd's old logic) preserved verbatim so every pre-generation-4 export (e.g. Game/bots/promoted/easy.json) keeps working unchanged via an optional "action_space" JSON field. - ship_observations.gd: append own contact state (SIZE 31 -> 35, append-only) so the value function can see what wall_contact_penalty fires on. - ship_ai_controller.gd: action space/decode via the codec; drop the vertical_ramp/pitch_roll_ramp mechanism entirely; tilt_penalty default lowered 4x (aerial approaches require pitching); flight telemetry (airborne_fraction, mean_altitude, air_touch_fraction, vertical_thrust_mean) and truncation-snapshot fields on get_info(). - training_mode.gd: new air_drill_chance state-setter branch (ball spawned high, ships low, kept clear of walls) so aerial practice is forced by the environment instead of relying on reward-driven exploration alone; snapshot terminal observations before a timeout reset for the truncation fix. - cosmic_env.py: remap ShipAIController's truncated/terminal_obs info into SB3's TimeLimit.truncated/terminal_observation keys. - train.py: --reset-logits (+ --reset-logits-heads) replaces the now-meaningless --reset-std; new EntropyFloorCallback (a persistent per-rollout ent_coef controller replacing the one-shot std-reset shock) and per-head entropy logging; FlightTelemetryCallback; --air-drill-chance/ --tilt-penalty flags; optional AbortIfCallback kill-criterion. - export_policy.py: writes the action_space block for MultiDiscrete models; index-level parity check (argmax per head) instead of comparing floats. - curriculum.py: full rewrite — 3 stages (bootstrap/selfplay/gauntlet), no grounded stage, full action space live from step 1; deletes generation 1-3's checkpoint-lineage machinery (nothing to resume from); final report evaluates against both promoted/easy.json and the new promoted/reference-grounded.json (a copy of curric-s5-aggression, the strongest grounded-era artifact, kept as a fixed yardstick). - run_training.sh/.gitignore: commit only final.zip, not the ~2400 intermediate checkpoint files a single stage was writing (~500MB -> ~0.2MB per run); requirements.txt pinned (behaviour here now depends on specific library internals, not just public APIs). - test_action_space.py (new): offline rung-0 check catching a head-order mismatch before it silently corrupts 24h of training. Validated: GDScript compiles clean (Godot --headless --import + script validation), free_play.tscn and training.tscn both boot headless without errors, offline action-space assertions pass. Not yet run: the actual smoke-training/A-B validation ladder steps in TRAINING.md's "Generation 4" section, before committing to the full ~32h curriculum. See TRAINING.md's "Generation 4" section for the full design writeup.
This commit is contained in:
+32
-2
@@ -10,6 +10,7 @@ learning policy: self-play by construction.
|
||||
import pathlib
|
||||
import subprocess
|
||||
|
||||
import numpy as np
|
||||
from godot_rl.core.godot_env import GodotEnv
|
||||
from godot_rl.wrappers.stable_baselines_wrapper import StableBaselinesGodotEnv
|
||||
|
||||
@@ -72,8 +73,13 @@ class CosmicClashEnv(GodotEnv):
|
||||
class CosmicClashVecEnv(StableBaselinesGodotEnv):
|
||||
"""SB3 VecEnv over N parallel CosmicClashEnv instances.
|
||||
|
||||
convert_action_space=True flattens the env's (Box(6), Discrete(2)) action
|
||||
space into a single Box(7): thrust xyz, rotation xyz, turbo (>0 means on).
|
||||
convert_action_space=True: godot_rl's ActionSpaceProcessor reports a
|
||||
gym.spaces.MultiDiscrete when every per-axis action entry is Discrete
|
||||
(see ShipActionCodec/ShipAIController.get_action_space) — nvec
|
||||
[5,5,5,5,5,5,2] for rotation xyz, thrust xyz, turbo, in that
|
||||
gymnasium-sorted key order. No conversion logic here needs to change for
|
||||
that; this class's only functional addition is the truncation-info
|
||||
remap below.
|
||||
"""
|
||||
|
||||
def __init__(self, godot_bin: str, n_parallel: int = 1, seed: int = 0, port: int = GodotEnv.DEFAULT_PORT, **kwargs):
|
||||
@@ -90,3 +96,27 @@ class CosmicClashVecEnv(StableBaselinesGodotEnv):
|
||||
self.n_parallel = n_parallel
|
||||
self._check_valid_action_space()
|
||||
self.results = None
|
||||
|
||||
def step(self, action):
|
||||
"""Remap ShipAIController.get_info()'s "truncated"/"terminal_obs"
|
||||
into the keys SB3's on_policy_algorithm looks for
|
||||
("TimeLimit.truncated"/"terminal_observation") so PPO bootstraps
|
||||
V(s) through an episode timeout instead of treating every 30s draw
|
||||
as a true terminal state.
|
||||
|
||||
Godot_rl's own godot_env.py never sets either key (it returns the
|
||||
same `done` array for both term and trunc, "# TODO update API to
|
||||
term, trunc") and StableBaselinesGodotEnv.step() only ever returns
|
||||
that single collapsed `dones` array to SB3 — so without this, PPO
|
||||
has no way to distinguish "episode ended because a goal was scored"
|
||||
(a genuine terminal, V(s)=0 is correct) from "episode ended because
|
||||
the 30s clock ran out" (an artificial boundary that should be
|
||||
bootstrapped through), and was silently treating every draw as the
|
||||
former in every curriculum generation to date.
|
||||
"""
|
||||
obs, rewards, dones, infos = super().step(action)
|
||||
for info in infos:
|
||||
if info.pop("truncated", False):
|
||||
info["TimeLimit.truncated"] = True
|
||||
info["terminal_observation"] = {"obs": np.array(info.pop("terminal_obs"), dtype=np.float32)}
|
||||
return obs, rewards, dones, infos
|
||||
|
||||
Reference in New Issue
Block a user