mirror of
https://github.com/jcreek/CosmicClash.git
synced 2026-09-10 16:04:04 +00:00
cf73074e27
A second adversarial review of the previous fix commit found two of its nine fixes silently defeated each other: the seq-range guard (fix for a MEDIUM epoch-mismatch finding) capped the exact variable the ring-overflow resync (fix for the original CRITICAL finding) depends on, making the resync unreachable in production and recreating permanent input death at a lower failure threshold, reachable via ordinary server tick loss alone. - CRITICAL: rebind the seq-range guard to InputJitterBuffer's own highest_ingested_seq (now public) instead of the consumer-side last_applied_seq, so it tracks the client's send epoch rather than a value that can lag arbitrarily far behind during a stall. - HIGH: InputLeadController's release logic still ANDed the old `lead > LEAD_MIN` gate onto the new depth-driven condition, so a backlog the controller never caused still couldn't drain. Split into two independent decisions: the seq-duplicate action follows real depth alone; lead's own bookkeeping separately never drops below its floor. - MEDIUM: widen the CI driver's movement/stalled sampling margin (run_seconds - 2.0, was - 0.5) and assert the peer is still in multiplayer.get_peers() at sample time, since the old margin let the check pass on residual starvation grace after a bot had already disconnected. - LOW: measure horizontal-only displacement in the human smoke test's movement check — the old 3D-distance bar was beatable by pure gravity settling with fully dead input. - LOW: fix a real "clean stderr" violation (match_net.gd broadcasting a departure notice to a peer whose ENet channels are already torn down, including a second peer disconnecting in the same poll batch) by deferring the notification to the next idle frame. - Wire the server's per-slot stalled bit into the client debug overlay for real — a prior commit message claimed this already reached the overlay when only the CI gate actually read it. Re-verified end-to-end against the real production RPC path (not just unit tests in isolation, which is how the composition bug got past the first round): a 2-bot CI match with a 1.5s host SIGSTOP freeze injected mid-run, well past the 0.6s threshold the review reproduced the bug at, now recovers cleanly on repeated runs with zero stderr noise.
109 lines
5.1 KiB
GDScript
109 lines
5.1 KiB
GDScript
class_name InputLeadController
|
|
extends RefCounted
|
|
|
|
# Client-owned input_lead control loop (multiplayer-todo.md §3.3, task 3.3).
|
|
# Standalone RefCounted, same reason as input_jitter_buffer.gd — scene-free
|
|
# so it's directly unit-testable against scripted depth traces.
|
|
#
|
|
# §3.3's own rationale for why this is the CLIENT's job alone, not shared
|
|
# with any server-side adaptation: three control loops acting on one plant
|
|
# (buffer occupancy) with different time constants is a textbook
|
|
# oscillation, and on a jittery link it presents to the player as
|
|
# intermittent sticky controls that are nearly impossible to attribute.
|
|
# The server (InputJitterBuffer, §3.2) only ever reports input_buffer_depth
|
|
# — it does nothing adaptive with it.
|
|
#
|
|
# "Lead" is realized concretely as extra distance between this client's own
|
|
# outgoing sequence numbers and what the server has actually consumed:
|
|
# skipping a sequence number (jumping the client's own seq counter by more
|
|
# than 1 for one tick) buys the server one more tick of buffered depth
|
|
# before it would starve; duplicating one (not incrementing the seq counter
|
|
# for one tick — the same seq gets sent again) narrows that margin by one
|
|
# tick of latency. The server's own ring buffer doesn't need to know this
|
|
# happened: a skipped seq just means "the redundant copies of it never
|
|
# existed, it's an ordinary drop" (already handled), and a duplicated seq
|
|
# is a same-seq resend, already discarded harmlessly once consumed
|
|
# (InputJitterBuffer.ingest()'s "seq <= last_applied_seq" check).
|
|
#
|
|
# Fast attack, slow release — a symmetric ±1-per-N-ticks slew would take
|
|
# two full seconds to absorb a single wifi spike, during which the player
|
|
# steers and the ship does not turn, "the most rage-inducing failure mode
|
|
# in any netcode" per §3.3's own words.
|
|
|
|
const LEAD_MIN := 1
|
|
const LEAD_MAX := 12
|
|
# "Never change it more than once per 30 ticks" (§3.3) — the floor that
|
|
# binds the fast-attack side; slow-release's own 60-tick cadence already
|
|
# exceeds it, so this one constant covers both.
|
|
const MIN_CHANGE_INTERVAL_TICKS := 30
|
|
const RELEASE_INTERVAL_TICKS := 60
|
|
const CLEAN_SURPLUS_TICKS := 120 # 2s at 60Hz
|
|
# §3.3: "target_depth = 1 (16.7 ms), not 2." Release only fires when the
|
|
# server-reported depth is genuinely ABOVE this — see update()'s own
|
|
# comment for why gating on `lead` alone (an adversarial review's original
|
|
# finding here) was wrong.
|
|
const TARGET_DEPTH := 1
|
|
|
|
var lead := LEAD_MIN
|
|
|
|
var _ticks_since_change := 0
|
|
var _clean_surplus_ticks := 0
|
|
|
|
|
|
# Call once per client physics tick with the most recently known server-
|
|
# reported input_buffer_depth for THIS client's own slot (echoed in every
|
|
# snapshot, §3.2) — or -1 if no snapshot carrying that field has arrived
|
|
# yet. Returns the seq delta the caller should add for this tick's
|
|
# outgoing packet: ordinarily 1 (ship normally increments its send
|
|
# sequence by exactly one tick's worth), or 1+N / 0 on a tick where a lead
|
|
# change actually fires (skip N extra / duplicate the current one).
|
|
func update(input_buffer_depth: int) -> int:
|
|
_ticks_since_change += 1
|
|
if input_buffer_depth < 0:
|
|
return 1
|
|
|
|
if input_buffer_depth <= 0:
|
|
# A starve: the server's ring was empty for this player when it
|
|
# built that snapshot. React immediately, not after 2 seconds of
|
|
# evidence like release requires — but still debounced against
|
|
# MIN_CHANGE_INTERVAL_TICKS so a burst of consecutive starve
|
|
# reports doesn't compound into repeated, overlapping jumps.
|
|
_clean_surplus_ticks = 0
|
|
if _ticks_since_change >= MIN_CHANGE_INTERVAL_TICKS and lead < LEAD_MAX:
|
|
var new_lead := mini(lead + 3, LEAD_MAX)
|
|
var delta := new_lead - lead
|
|
lead = new_lead
|
|
_ticks_since_change = 0
|
|
return 1 + delta
|
|
return 1
|
|
|
|
# Release must react to the ACTUAL server-reported depth, not to this
|
|
# controller's own memory of past attacks. A first pass at this fix
|
|
# added the depth check above but left the OLD gate, `lead > LEAD_MIN`,
|
|
# still ANDed onto the final condition below — so a backlog this
|
|
# controller did NOT itself cause (a server hitch, persistent client/
|
|
# server clock drift, a ring resync) still could never be drained:
|
|
# with lead pinned at its starting floor, that clause always failed
|
|
# even while input_buffer_depth sat well above target. A second
|
|
# adversarial review caught it, confirmed by this file's own
|
|
# test_release_drains_a_backlog_it_never_caused_itself, whose original
|
|
# assertion text literally said "lead cannot release below its own
|
|
# floor even under large surplus" as if that were correct.
|
|
#
|
|
# The fix splits the one gate into two separate decisions: whether to
|
|
# duplicate this tick's seq (the only thing that actually narrows real
|
|
# buffered depth) follows the real signal alone, below; whether to
|
|
# keep decrementing `lead`'s own bookkeeping below its documented
|
|
# floor is a separate, cosmetic-only choice made inside that branch.
|
|
if input_buffer_depth > TARGET_DEPTH:
|
|
_clean_surplus_ticks += 1
|
|
else:
|
|
_clean_surplus_ticks = 0
|
|
|
|
if _clean_surplus_ticks >= CLEAN_SURPLUS_TICKS and _ticks_since_change >= RELEASE_INTERVAL_TICKS:
|
|
if lead > LEAD_MIN:
|
|
lead -= 1
|
|
_ticks_since_change = 0
|
|
return 0 # duplicate this tick's seq — one tick of latency recovered
|
|
return 1
|