mirror of
https://github.com/jcreek/CosmicClash.git
synced 2026-09-10 16:04:04 +00:00
1811e9333e
Three curriculum generations (2026-07-21 through 2026-08-04) all tried gating *when* the policy could use vertical thrust/pitch-roll on top of a continuous Gaussian action space, and all three failed the same way: PPO's action-distribution std collapsed within ~10% of steps and never recovered, landing at a 15-32% win rate vs the grounded reference regardless of mechanism (hard mask, then a gradual ramp). Generation 3's final attempt just landed at 24% — the worst of the three. Root cause, verified against this project's own physics: hovering this ship requires *holding* thrust.y ~= 0.408 continuously (mass 5.0, vertical_thrust 120, gravity 9.8). A collapsed near-zero-mean Gaussian can brush that value but never sustain it long enough to earn the reward gradient that would move the mean — no amount of gating *when* the axis acts fixes a problem in *how* the policy represents a decision on it. This also independently found and fixes a real bug: godot_rl never marks an episode timeout as a truncation, so PPO was bootstrapping V(s)=0 on every 30s draw in every generation to date. - Game/scripts/ship_action_codec.gd (new): single source of truth for a per-axis MultiDiscrete action space (7 heads, nvec [5,5,5,5,5,5,2]) shared by training and in-game inference, replacing the continuous Gaussian. thrust_y's bins are deliberately asymmetric so a random policy drifts through the volume instead of floor-pinning. Legacy continuous decode (ai_ship_controller.gd's old logic) preserved verbatim so every pre-generation-4 export (e.g. Game/bots/promoted/easy.json) keeps working unchanged via an optional "action_space" JSON field. - ship_observations.gd: append own contact state (SIZE 31 -> 35, append-only) so the value function can see what wall_contact_penalty fires on. - ship_ai_controller.gd: action space/decode via the codec; drop the vertical_ramp/pitch_roll_ramp mechanism entirely; tilt_penalty default lowered 4x (aerial approaches require pitching); flight telemetry (airborne_fraction, mean_altitude, air_touch_fraction, vertical_thrust_mean) and truncation-snapshot fields on get_info(). - training_mode.gd: new air_drill_chance state-setter branch (ball spawned high, ships low, kept clear of walls) so aerial practice is forced by the environment instead of relying on reward-driven exploration alone; snapshot terminal observations before a timeout reset for the truncation fix. - cosmic_env.py: remap ShipAIController's truncated/terminal_obs info into SB3's TimeLimit.truncated/terminal_observation keys. - train.py: --reset-logits (+ --reset-logits-heads) replaces the now-meaningless --reset-std; new EntropyFloorCallback (a persistent per-rollout ent_coef controller replacing the one-shot std-reset shock) and per-head entropy logging; FlightTelemetryCallback; --air-drill-chance/ --tilt-penalty flags; optional AbortIfCallback kill-criterion. - export_policy.py: writes the action_space block for MultiDiscrete models; index-level parity check (argmax per head) instead of comparing floats. - curriculum.py: full rewrite — 3 stages (bootstrap/selfplay/gauntlet), no grounded stage, full action space live from step 1; deletes generation 1-3's checkpoint-lineage machinery (nothing to resume from); final report evaluates against both promoted/easy.json and the new promoted/reference-grounded.json (a copy of curric-s5-aggression, the strongest grounded-era artifact, kept as a fixed yardstick). - run_training.sh/.gitignore: commit only final.zip, not the ~2400 intermediate checkpoint files a single stage was writing (~500MB -> ~0.2MB per run); requirements.txt pinned (behaviour here now depends on specific library internals, not just public APIs). - test_action_space.py (new): offline rung-0 check catching a head-order mismatch before it silently corrupts 24h of training. Validated: GDScript compiles clean (Godot --headless --import + script validation), free_play.tscn and training.tscn both boot headless without errors, offline action-space assertions pass. Not yet run: the actual smoke-training/A-B validation ladder steps in TRAINING.md's "Generation 4" section, before committing to the full ~32h curriculum. See TRAINING.md's "Generation 4" section for the full design writeup.
97 lines
3.6 KiB
GDScript
97 lines
3.6 KiB
GDScript
class_name AIShipController
|
|
extends ShipController
|
|
|
|
# Drives a ship from a trained self-play policy (see TRAINING.md). Builds the
|
|
# same canonical observation as training (ShipObservations) and runs the
|
|
# policy MLP in GDScript (PolicyNetwork) — the shipped bot has no Python,
|
|
# .NET, or network dependency.
|
|
#
|
|
# Difficulty is (model, reaction_ticks, action_noise): weaker checkpoints make
|
|
# easier bots outright, and the two knobs handicap a given model further —
|
|
# slower reactions and noisier execution. Models live in res://bots/.
|
|
|
|
@export_file("*.json") var model_path: String = ""
|
|
# Decide a new action every N physics ticks, holding the last one between
|
|
# decisions. 8 matches the training action_repeat; larger = slower reactions.
|
|
@export_range(1, 60) var reaction_ticks: int = 8
|
|
# Uniform noise magnitude added to each action axis (0 = play at full skill).
|
|
@export_range(0.0, 1.0) var action_noise: float = 0.0
|
|
|
|
# Only meaningful for a "continuous"-action_space model (see
|
|
# ShipActionCodec) — i.e. one exported before curriculum generation 4, such
|
|
# as Game/bots/promoted/reference-grounded.json. Must mirror whatever the
|
|
# model was actually trained with: a model trained grounded (mask on) never
|
|
# got a reward gradient on these axes, so its raw output there is untrained
|
|
# noise — leaving this true for such a model doesn't make it fly well, it
|
|
# just lets that noise reach the ship instead of being discarded like it was
|
|
# in training. Set false to match a grounded-trained model's actual
|
|
# behaviour. Generation-4-onward (multi_discrete) models train the full
|
|
# action space from the start, so these flags are ignored for them.
|
|
@export var allow_vertical := true
|
|
@export var allow_pitch_roll := true
|
|
|
|
var _policy: PolicyNetwork
|
|
var _action := ShipAction.new()
|
|
var _ticks_until_decision := 0
|
|
|
|
var _ship: Ship
|
|
var _opponent: Ship
|
|
var _ball: RigidBody3D
|
|
var _attack_goal_position: Vector3
|
|
var _scene_refs_ready := false
|
|
|
|
|
|
func _ready():
|
|
if not model_path.is_empty():
|
|
_policy = PolicyNetwork.load_from_file(model_path)
|
|
|
|
|
|
func get_action() -> ShipAction:
|
|
if _policy == null:
|
|
return _action # unloaded model: behaves like the inert placeholder
|
|
if not _scene_refs_ready and not _discover_scene_refs():
|
|
return _action
|
|
|
|
_ticks_until_decision -= 1
|
|
if _ticks_until_decision <= 0:
|
|
_ticks_until_decision = reaction_ticks
|
|
_decide()
|
|
return _action
|
|
|
|
|
|
func _decide() -> void:
|
|
var obs := ShipObservations.build(_ship, _opponent, _ball, _attack_goal_position)
|
|
var out := _policy.forward(obs)
|
|
# See ShipActionCodec for the decode — the single source of truth shared
|
|
# with the training side, so this must never reimplement layout/ordering
|
|
# locally (see that file's header for why).
|
|
if _policy.action_space.get("type", "continuous") == "continuous":
|
|
_action = ShipActionCodec.from_continuous(out, action_noise)
|
|
if not allow_pitch_roll:
|
|
_action.rotation.x = 0.0
|
|
_action.rotation.z = 0.0
|
|
if not allow_vertical:
|
|
_action.thrust.y = 0.0
|
|
else:
|
|
_action = ShipActionCodec.from_logits(out, action_noise)
|
|
|
|
|
|
# Find ship/ball/opponent/goal once everything is spawned. ShipAction axes
|
|
# are body-frame so only observations need team context (ShipObservations).
|
|
func _discover_scene_refs() -> bool:
|
|
_ship = get_parent() as Ship
|
|
if _ship == null or not is_inside_tree():
|
|
return false
|
|
_ball = get_tree().get_first_node_in_group("ball")
|
|
for node in get_tree().get_nodes_in_group("ship"):
|
|
if node != _ship:
|
|
_opponent = node
|
|
break
|
|
for goal in get_tree().get_nodes_in_group("goal"):
|
|
if goal.team == 1 - _ship.team:
|
|
_attack_goal_position = goal.global_position
|
|
if _ball == null:
|
|
return false
|
|
_scene_refs_ready = true
|
|
return true
|