mirror of
https://github.com/jcreek/CosmicClash.git
synced 2026-09-10 16:04:04 +00:00
1811e9333e
Three curriculum generations (2026-07-21 through 2026-08-04) all tried gating *when* the policy could use vertical thrust/pitch-roll on top of a continuous Gaussian action space, and all three failed the same way: PPO's action-distribution std collapsed within ~10% of steps and never recovered, landing at a 15-32% win rate vs the grounded reference regardless of mechanism (hard mask, then a gradual ramp). Generation 3's final attempt just landed at 24% — the worst of the three. Root cause, verified against this project's own physics: hovering this ship requires *holding* thrust.y ~= 0.408 continuously (mass 5.0, vertical_thrust 120, gravity 9.8). A collapsed near-zero-mean Gaussian can brush that value but never sustain it long enough to earn the reward gradient that would move the mean — no amount of gating *when* the axis acts fixes a problem in *how* the policy represents a decision on it. This also independently found and fixes a real bug: godot_rl never marks an episode timeout as a truncation, so PPO was bootstrapping V(s)=0 on every 30s draw in every generation to date. - Game/scripts/ship_action_codec.gd (new): single source of truth for a per-axis MultiDiscrete action space (7 heads, nvec [5,5,5,5,5,5,2]) shared by training and in-game inference, replacing the continuous Gaussian. thrust_y's bins are deliberately asymmetric so a random policy drifts through the volume instead of floor-pinning. Legacy continuous decode (ai_ship_controller.gd's old logic) preserved verbatim so every pre-generation-4 export (e.g. Game/bots/promoted/easy.json) keeps working unchanged via an optional "action_space" JSON field. - ship_observations.gd: append own contact state (SIZE 31 -> 35, append-only) so the value function can see what wall_contact_penalty fires on. - ship_ai_controller.gd: action space/decode via the codec; drop the vertical_ramp/pitch_roll_ramp mechanism entirely; tilt_penalty default lowered 4x (aerial approaches require pitching); flight telemetry (airborne_fraction, mean_altitude, air_touch_fraction, vertical_thrust_mean) and truncation-snapshot fields on get_info(). - training_mode.gd: new air_drill_chance state-setter branch (ball spawned high, ships low, kept clear of walls) so aerial practice is forced by the environment instead of relying on reward-driven exploration alone; snapshot terminal observations before a timeout reset for the truncation fix. - cosmic_env.py: remap ShipAIController's truncated/terminal_obs info into SB3's TimeLimit.truncated/terminal_observation keys. - train.py: --reset-logits (+ --reset-logits-heads) replaces the now-meaningless --reset-std; new EntropyFloorCallback (a persistent per-rollout ent_coef controller replacing the one-shot std-reset shock) and per-head entropy logging; FlightTelemetryCallback; --air-drill-chance/ --tilt-penalty flags; optional AbortIfCallback kill-criterion. - export_policy.py: writes the action_space block for MultiDiscrete models; index-level parity check (argmax per head) instead of comparing floats. - curriculum.py: full rewrite — 3 stages (bootstrap/selfplay/gauntlet), no grounded stage, full action space live from step 1; deletes generation 1-3's checkpoint-lineage machinery (nothing to resume from); final report evaluates against both promoted/easy.json and the new promoted/reference-grounded.json (a copy of curric-s5-aggression, the strongest grounded-era artifact, kept as a fixed yardstick). - run_training.sh/.gitignore: commit only final.zip, not the ~2400 intermediate checkpoint files a single stage was writing (~500MB -> ~0.2MB per run); requirements.txt pinned (behaviour here now depends on specific library internals, not just public APIs). - test_action_space.py (new): offline rung-0 check catching a head-order mismatch before it silently corrupts 24h of training. Validated: GDScript compiles clean (Godot --headless --import + script validation), free_play.tscn and training.tscn both boot headless without errors, offline action-space assertions pass. Not yet run: the actual smoke-training/A-B validation ladder steps in TRAINING.md's "Generation 4" section, before committing to the full ~32h curriculum. See TRAINING.md's "Generation 4" section for the full design writeup.
91 lines
3.3 KiB
GDScript
91 lines
3.3 KiB
GDScript
class_name PolicyNetwork
|
|
extends RefCounted
|
|
|
|
# Minimal MLP forward pass for running trained policies in pure GDScript —
|
|
# no .NET build or ONNX runtime needed. Weights come from a JSON file written
|
|
# by training/export_policy.py (see TRAINING.md). The policy net is tiny
|
|
# (31 → 64 → 64 → 7 by default), and the bot only thinks every few physics
|
|
# ticks, so GDScript is plenty fast.
|
|
#
|
|
# JSON shape:
|
|
# {
|
|
# "input_size": 31,
|
|
# "layers": [
|
|
# {"weights": [[out x in floats]], "biases": [out floats], "activation": "tanh" | "linear"},
|
|
# ...
|
|
# ]
|
|
# "action_space": {"type": "multi_discrete", "heads": [{"name","bins"}, ...]} // optional
|
|
# }
|
|
#
|
|
# "action_space" is absent from every model exported before curriculum
|
|
# generation 4 (e.g. Game/bots/promoted/easy.json) — absence means
|
|
# {"type": "continuous"}, decoded via ShipActionCodec.from_continuous, the
|
|
# same flattened-Box(7)-mean-output path this class has always produced.
|
|
# This class itself never changes behaviour based on it; only the caller
|
|
# (AIShipController._decide) branches on action_space["type"].
|
|
|
|
var input_size: int = 0
|
|
var action_space: Dictionary = {"type": "continuous"}
|
|
var _layers: Array = []
|
|
|
|
|
|
static func load_from_file(path: String) -> PolicyNetwork:
|
|
if not FileAccess.file_exists(path):
|
|
push_error("PolicyNetwork: model file not found: %s" % path)
|
|
return null
|
|
var text := FileAccess.get_file_as_string(path)
|
|
var data: Variant = JSON.parse_string(text)
|
|
if data == null or not (data is Dictionary) or not data.has("layers"):
|
|
push_error("PolicyNetwork: invalid model file: %s" % path)
|
|
return null
|
|
|
|
var net := PolicyNetwork.new()
|
|
net.input_size = int(data.get("input_size", 0))
|
|
net.action_space = data.get("action_space", {"type": "continuous"})
|
|
for layer in data["layers"]:
|
|
# Flatten each layer's weights into a PackedFloat64Array for speed
|
|
var out_size: int = layer["biases"].size()
|
|
var in_size: int = layer["weights"][0].size()
|
|
var flat := PackedFloat64Array()
|
|
flat.resize(out_size * in_size)
|
|
var i := 0
|
|
for row in layer["weights"]:
|
|
for value in row:
|
|
flat[i] = value
|
|
i += 1
|
|
var biases := PackedFloat64Array(layer["biases"])
|
|
net._layers.append({
|
|
"weights": flat,
|
|
"biases": biases,
|
|
"in_size": in_size,
|
|
"out_size": out_size,
|
|
"tanh": layer.get("activation", "linear") == "tanh",
|
|
})
|
|
return net
|
|
|
|
|
|
func forward(observation: Array) -> Array:
|
|
if observation.size() < input_size:
|
|
push_error("PolicyNetwork: observation has %d values, model expects %d" % [observation.size(), input_size])
|
|
# Slice rather than trust the caller: ShipObservations.SIZE only ever
|
|
# grows (append-only), so an older/smaller model must still decode
|
|
# correctly against a newer, longer observation vector — the extra
|
|
# trailing values it never trained on are simply dropped here rather
|
|
# than corrupting the first layer's dot product by accident.
|
|
var x := PackedFloat64Array(observation.slice(0, input_size))
|
|
for layer in _layers:
|
|
var in_size: int = layer["in_size"]
|
|
var out_size: int = layer["out_size"]
|
|
var weights: PackedFloat64Array = layer["weights"]
|
|
var biases: PackedFloat64Array = layer["biases"]
|
|
var y := PackedFloat64Array()
|
|
y.resize(out_size)
|
|
for row in out_size:
|
|
var sum := biases[row]
|
|
var offset := row * in_size
|
|
for col in in_size:
|
|
sum += weights[offset + col] * x[col]
|
|
y[row] = tanh(sum) if layer["tanh"] else sum
|
|
x = y
|
|
return Array(x)
|