mirror of
https://github.com/jcreek/CosmicClash.git
synced 2026-09-11 00:14:00 +00:00
fix(multiplayer): resolve composition regression from second adversarial review
A second adversarial review of the previous fix commit found two of its nine fixes silently defeated each other: the seq-range guard (fix for a MEDIUM epoch-mismatch finding) capped the exact variable the ring-overflow resync (fix for the original CRITICAL finding) depends on, making the resync unreachable in production and recreating permanent input death at a lower failure threshold, reachable via ordinary server tick loss alone. - CRITICAL: rebind the seq-range guard to InputJitterBuffer's own highest_ingested_seq (now public) instead of the consumer-side last_applied_seq, so it tracks the client's send epoch rather than a value that can lag arbitrarily far behind during a stall. - HIGH: InputLeadController's release logic still ANDed the old `lead > LEAD_MIN` gate onto the new depth-driven condition, so a backlog the controller never caused still couldn't drain. Split into two independent decisions: the seq-duplicate action follows real depth alone; lead's own bookkeeping separately never drops below its floor. - MEDIUM: widen the CI driver's movement/stalled sampling margin (run_seconds - 2.0, was - 0.5) and assert the peer is still in multiplayer.get_peers() at sample time, since the old margin let the check pass on residual starvation grace after a bot had already disconnected. - LOW: measure horizontal-only displacement in the human smoke test's movement check — the old 3D-distance bar was beatable by pure gravity settling with fully dead input. - LOW: fix a real "clean stderr" violation (match_net.gd broadcasting a departure notice to a peer whose ENet channels are already torn down, including a second peer disconnecting in the same poll batch) by deferring the notification to the next idle frame. - Wire the server's per-slot stalled bit into the client debug overlay for real — a prior commit message claimed this already reached the overlay when only the CI gate actually read it. Re-verified end-to-end against the real production RPC path (not just unit tests in isolation, which is how the composition bug got past the first round): a 2-bot CI match with a 1.5s host SIGSTOP freeze injected mid-run, well past the 0.6s threshold the review reproduced the bug at, now recovers cleanly on repeated runs with zero stderr noise.
This commit is contained in:
@@ -42,7 +42,19 @@ var _seeded := false
|
||||
# it, a backlog bigger than RING_SIZE (a host stall, or persistent client/
|
||||
# server clock drift) permanently zeroed a connected player's input for the
|
||||
# rest of the match.
|
||||
var _highest_ingested_seq := -1
|
||||
#
|
||||
# Deliberately public (no underscore), same as last_applied_seq: the
|
||||
# networked_match.gd caller's seq-range guard (§3.1 step 4) must bound
|
||||
# against THIS, not against last_applied_seq. A second adversarial review
|
||||
# found that bounding against last_applied_seq caps every accepted seq at
|
||||
# last_applied_seq + RING_SIZE, which in turn caps this field at the same
|
||||
# ceiling — making the resync condition below (which needs this field to
|
||||
# reach expected + RING_SIZE) arithmetically unreachable on the only call
|
||||
# path that exists in production. The two fixes looked independent but
|
||||
# shared a variable and silently cancelled each other out. highest_ingested
|
||||
# tracks the client's own send epoch instead, which the guard can safely
|
||||
# let run ahead of a lagging consumer.
|
||||
var highest_ingested_seq := -1
|
||||
|
||||
|
||||
func _init() -> void:
|
||||
@@ -72,8 +84,8 @@ func ingest(newest_seq: int, actions: Array) -> void:
|
||||
# "expected" with reality the moment real data first exists.
|
||||
last_applied_seq = newest_seq - actions.size()
|
||||
_seeded = true
|
||||
if newest_seq > _highest_ingested_seq:
|
||||
_highest_ingested_seq = newest_seq
|
||||
if newest_seq > highest_ingested_seq:
|
||||
highest_ingested_seq = newest_seq
|
||||
for i in actions.size():
|
||||
var seq: int = newest_seq - i
|
||||
if seq <= last_applied_seq:
|
||||
@@ -110,7 +122,7 @@ func consume() -> ShipAction:
|
||||
# ticks of not-yet-consumed data at once — if the caller has fallen
|
||||
# further behind the newest data actually arriving than that (a host
|
||||
# stall, or persistent client/server clock drift), every tick between
|
||||
# "expected" and "_highest_ingested_seq - RING_SIZE" has already been
|
||||
# "expected" and "highest_ingested_seq - RING_SIZE" has already been
|
||||
# irrecoverably overwritten by more recent arrivals landing on the same
|
||||
# ring slots. Waiting for it tick-by-tick would starve — and, past
|
||||
# STARVE_ZERO_TICKS, zero this player's ship — for the ENTIRE gap even
|
||||
@@ -119,8 +131,8 @@ func consume() -> ShipAction:
|
||||
# host freeze permanently zeroed a connected player's input for the
|
||||
# rest of the match, with no self-recovery). Skip the unrecoverable
|
||||
# span and resync directly to what the ring can still actually provide.
|
||||
if _highest_ingested_seq - expected >= RING_SIZE:
|
||||
last_applied_seq = _highest_ingested_seq - RING_SIZE
|
||||
if highest_ingested_seq - expected >= RING_SIZE:
|
||||
last_applied_seq = highest_ingested_seq - RING_SIZE
|
||||
expected = last_applied_seq + 1
|
||||
idx = expected % RING_SIZE
|
||||
|
||||
|
||||
Reference in New Issue
Block a user