multiplayer-next.md was a 1662-line mix of standing architecture spec and task-completion tracking, most of which was dense per-task DONE evidence for finished Phases 0-6. Split it: - MULTIPLAYER_SPEC.md (new): the locked architecture decisions, wire format, server-side input handling, prediction/reconciliation, latency/frame-rate budget, and match lifecycle state machine - standing design reference, not task-tracked. - multiplayer-next.md (trimmed 1662 -> ~370 lines): only outstanding work remains - §0 status, §7 Phase 7/8 task tables condensed to "what's left" per task, §8-11 reference material (refactoring notes, gotchas, testing, flagged items). Phases 0-6 collapsed to a pointer at git history instead of ~500 lines of DONE evidence. Also: - Repointed every `multiplayer-next.md §N` code comment (N 1-6) across Game/scripts, Game/tools and Game/tests to MULTIPLAYER_SPEC.md, since those sections moved. Task-number references (`task N.N`, §7-11) correctly still point at multiplayer-next.md. - Updated CLAUDE.md's doc index and docs/TECH_STACK.md's spec-section citations to match. - TODO.md: added a "what's left to actually finish multiplayer (human-actionable)" checklist pulled from multiplayer-next.md §0 and docs/MATCHMAKING.md - things that need a person (hardware, a design decision, a Steam App ID, hands on a controller), not more agent code.
42 KiB
Online multiplayer — task breakdown
The tracking document for the online multiplayer effort's outstanding work:
what's left, why, and the task breakdown. TODO.md points here for anything
multiplayer-related. The architecture decisions, wire format, input
handling, prediction, latency budget, and match lifecycle spec this work
assumes now live in MULTIPLAYER_SPEC.md as its
own sections 1–6 — read those before picking up work in Phase 2 or later.
How to use this doc: start at §0 for what's outstanding right now. Pick up a single numbered task, do it, verify it against its stated acceptance criterion, and stop. §9 is a running gotchas list — check it before debugging something that looks like a Godot/Jolt engine quirk, and add to it when you find a new one.
This revision keeps only outstanding work. Phases 0–6 are fully
implemented and verified locally; their task-by-task implementation evidence
has been trimmed from this document and lives in git history
(git log -- multiplayer-next.md) rather than here. Phase 7 (Steam) and
Phase 8 (matchmaking) are in progress — the tables below list only what
remains on each task, not what's already built. Phase 8 is a 1.0 launch
blocker, adds a component outside the Godot project (a Go backend
service), and has a critical open blocker: see §0.
0. Outstanding work — the short list
The one place to look before planning. Everything here is also written up where it belongs; this is the index, not the detail.
Phase 8 (matchmaking, ranked, per-match server autoscaling) is a 1.0 launch
blocker and is in progress. It is larger than anything below and adds a
backend service outside the Godot project. Tasks are in §7; the design is in
docs/MATCHMAKING.md.
The current root blocker: nothing in production ever publishes a
player's signed match assignment. store.SaveAssignment/SaveAssignments/
SaveVerifiedAssignmentRoster and allocator.Service.PublishRoster are
fully built and tested in isolation, but no real code path
(allocator/worker.go, cmd/allocator) ever calls them — only tests do, by
seeding the table directly rather than exercising the real write path. Since
AdvanceServerRegistration's SQL requires an assignments row per
participant before a match can reach ASSIGNMENT_READY, a real deployment
cannot advance any match past PROCESS_READY — no player can ever receive
a real assignment or connect, regardless of how correct every other piece
(including the client-side connect-wiring in task 8.41) is. See task 8.31
for the full detail. Closing it needs new security-relevant design (a
join-signing key shared between allocator and game server, roster-digest
computation, per-player authorisation construction) — flagged rather than
built, at the user's explicit direction, pending a decision on that design.
Blocking sign-off — the work exists, the verification does not
| # | What | Why it is not done | Detail |
|---|---|---|---|
| A | Phase 4 human playtest at ~100 ms RTT. Does the ship feel local? Does the ball? Do contact corrections read as bumps or as glitches? | Needs hands on a controller. Every numeric gate is green; feel is the milestone's actual subject and no percentile can answer it. | §5.7 |
| B | Phase 5 3v3 gate: a full start-to-finish match with 6 players, a mid-match disconnect, and a late joiner. | Needs a real multi-client session. Every scenario is verified at 1v1 plus a two-bot CI match; nothing has run at 3v3. | §6 |
These two are independent and can be done in either order, but B is the cheaper of the two to arrange and would also exercise A's conditions incidentally.
Known defects
| # | What | Severity | Detail |
|---|---|---|---|
| C | Slot reservation and takeover are keyed on display name alone, for direct/unauthenticated servers only. For allocated (signed-roster) matches this is resolved — reconnect reclaim and late-join promotion carry the verified PlayerID across peer-id changes. Direct/community servers with no Steam identity still resolve reclaim by display name; a peer connecting with a departed player's name inside the 30 s window claims their slot. |
Real, demonstrated, bounded to the unauthenticated direct-server path. | §11, Phase 7 task 7.4 |
| D | Input is still lost at the transport layer during a long server stall, variably — 7 of 8 runs measured 0.00 % of the sequence stream missing, the eighth 23.54 %. | Low. Distinct from the rate-limiter cause, which is fixed. The seq-guard resync visibly recovers it. | §9 gotchas 39, 48, 49 |
The residual half of C (direct/community servers) is fixed for free by task
7.4 (Steam auth tickets in hello) once Phase 7 lands; it has not been
given a bespoke solution for that reason.
Open architectural question
| # | What | Detail |
|---|---|---|
| F | A contact-cohort-only shadow world. The remaining known prediction weakness is the contact cohort. Whether it is worth a client-side shadow Jolt world scoped to contacts alone is undecided — and deliberately so until A supplies the felt evidence. | §5.7 |
Unstarted phases
- Phase 6 external gate: run the exported Docker server and clients from separate machines over the internet, then play a full match. This is a controlled test only until C is fully closed (i.e. Phase 7 lands).
- Phase 7 — Steam transport, browser, identity and production SDR (8 tasks, in progress): the optional bootstrap and
NetTransportfoundation now exist, but custom Steam client/server export templates have not yet been supplied. Browser, verified tickets, bans, production credentials and ticketed Hosted Dedicated Server SDR await a project-owned Steamworks App ID and Valve coordination. Carries the fix for the direct-server half of C and is the hard prerequisite for Phase 8's production Steam identity.
Phase 6's external gate has no dependency on Phase 7 for a controlled test, but Phase 7 is next in priority because Steam identity is required before public exposure.
Deferred by choice, not forgotten
120 Hz simulation, the latency-gap measurement (§5.7's acceptance
criterion), authored audio, split-screen — all in §11 with what each would
buy and cost. The procedural audio hooks are implemented; authored assets
and production mixing remain open in TODO.md.
Sections 1–6 (architecture, wire format, input handling, prediction,
latency budget, match lifecycle) live in
MULTIPLAYER_SPEC.md. A bare §N below refers to
that document for N 1–6, and to this one for N 7+.
7. Phase and task breakdown
[P] parallelisable within its phase · [D:x.y] hard dependency
Phases 0–6 — complete
Every task in Phases 0–6 is implemented and verified locally: non-networked
refactors, transport/connection/lobby, server-authoritative simulation with
a dumb client, input pipeline hardening, prediction and reconciliation for
ship and ball, match lifecycle, and dedicated-server productionisation
(Docker export, rotation/drain, CI). The two remaining gates on this work
are human verification, not code — see §0 gates A and B. Task-by-task
acceptance evidence for Phases 0–6 has been trimmed from this document;
git log -- multiplayer-next.md has the full history if a past task's
reasoning is needed.
Phase 7 — Steam transport, browser, identity
In progress. GodotSteam requires custom engine builds and export
templates — including for the headless server; budget for it. The
NetTransport boundary (ENet + feature-gated steam_transport.gd) is
already extracted so this phase adds a second implementation rather than
retrofitting one.
| # | Task | Remaining |
|---|---|---|
7.1 [D:1.2] |
GodotSteam integration and custom export templates, client and headless server | Awaiting the custom binaries/SDK access |
7.2 [D:7.1] |
NetTransport Steam implementation (SteamMultiplayerPeer, SDR) |
Server advertising waits for ISteamGameServer work |
7.3 [D:7.2] [P] |
Server-browser UI and ISteamMatchmakingServers adapter |
Unimplemented until real Steam SDK/API access is available; ENet direct-IP remains the supported browser-free path meanwhile |
7.4 [D:7.2] [P] |
Auth tickets in hello → BeginAuthSession, Steam identity in the roster, persistent ban list |
Real GodotSteam auth integration, server-side VAC state, durable ban storage remain. Fixes known defect C for direct/community servers once landed |
7.5 [D:7.2] [P] |
SteamBootstrap gating (stock builds keep ENet, explicit Steam selection fails closed) |
Custom Steam client/server export smoke remains blocked on pinned GodotSteam binaries |
7.6 [D:7.4] |
Backend AuthCoordinator, session persistence, ControlPlaneClient.login_steam() |
Real Steam BeginAuthSession/EndAuthSession adapter, login UI, live PostgreSQL/session integration remain |
7.7 [D:7.1] [P] |
Obtain the production App ID, publisher key, SDR coordinator SDK/signing approval, certificates and hosted-data-centre support from Valve | Not started |
7.8 [D:7.6,7.7] |
Ticketed Hosted Dedicated Server SDR: routing registration, coordinator-issued player→server relay tickets, client ticket installation, reconnect and expiry | Not started; depends on 7.6 and 7.7 |
Phase 8 — Matchmaking, ranked ladder, per-match server autoscaling
1.0 launch blocker. Full design and reasoning:
docs/MATCHMAKING.md. This is the first phase to add
a component outside the Godot project — a Go backend service — and that is
the largest architectural departure in the project's history; read the
design doc before picking up any task below. The local control-plane,
durable-store, allocated-server, and verification paths are substantially
implemented; every row below lists only what's still open, not what's
built. The critical path is task 8.31 — see §0's root blocker.
Hard dependency on 7.6 and 7.8. The local allocated path binds slot reclaim to a control-plane-signed player identity and locks its team/slot pair, but production Steam ticket verification is still required before a rating can be trusted. Production allocation also depends on the ticketed Hosted Dedicated Server SDR route; ENet remains the local/CI/community transport, not a silent production fallback.
This inverts the server model from Phases 1–7's community server (runs
forever, waits for --min-players, plays a match, rotates arena, repeats).
Matchmaking makes the player durable instead — queue, get grouped by
rating, and a server is allocated for that one match and destroyed
after. Both models ship; they are different playlists, not a replacement.
Tasks 8.1–8.4 (versioned contracts, state transitions, an ADR locking the Go/PostgreSQL/Redis/Agones stack, and launch SLOs) and 8.11 (threat model) are done; everything below is what's left on the tasks still open.
8A — Architecture, contracts and data
| # | Task | Remaining |
|---|---|---|
8.5 [D:8.4] |
PostgreSQL migrations 0001–0013 (idempotency, queue fencing, identities, ratings, matches, results, audits, outbox, allocator registry, proposal plans, leases, quotas) | New validations await a live database rerun — Docker storage exhausted locally |
8.6 [D:8.3,8.4] |
Allocated-mode ServerConfig fields |
Signed-authorisation admission, dynamic endpoint wiring, full manifest/runtime tests remain |
8B — Authentication and secure control plane
| # | Task | Remaining |
|---|---|---|
8.7 [D:7.6,8.3] |
Ticket policy binding expected App ID/identity | Real AuthenticateUserTicket backend adapter, bans, publisher secret store, real Steam verification remain |
8.8 [D:8.7] |
Session policy (opaque tokens, digests, revocation) | Distributed revocation coordination, live Steam/session integration remain |
8.9 [D:8.4,8.7] |
Join policy, durable reconnect leases | Live PostgreSQL/Godot process-restart and outage recovery verification remains |
8.10 [D:8.5,8.31] |
Workload credential policy (signed tokens, not Kubernetes JWTs), delivery channel, conflict alerting | Never run against a real Agones cluster; alert validated only statically, not against live Prometheus/Alertmanager traffic |
8.12 [D:8.11] |
Kubernetes hardening baseline, rate/quota limiting, degraded-mode gate | Private-store provisioning, distributed/global quotas, edge DDoS/WAF/origin shielding, encrypted backups, live policy/load tests remain |
8.13 [D:8.12] |
Digest-pinned images, supply-chain policy checker | Registry SBOM/scan/sign/admission execution and a concrete production overlay remain |
8C — Queueing, matchmaking, playlists and rating
| # | Task | Remaining |
|---|---|---|
8.14 [D:8.4,8.5,8.8] |
Queue policy (ownership, heartbeat/expiry, candidate projection) | Live Redis failover-under-load and worker integration remain |
8.15 [D:7.8,8.3] |
Probe validation (RTT, nonce/freshness/region, quarantine) | Steam coordinator, regional probe adapters, multi-region probe population remain |
8.16 [D:8.14,8.15] |
Candidate/team formation, matcher worker | Long-running worker integration remains |
8.17 [D:8.14,8.16] |
Proposal policy (response window, cooldowns, offender/innocent split) | Live PostgreSQL execution and allocation integration remain |
8.18 [D:8.5,8.14,8.17] |
Store layer (serializable retries, claim SQL, atomic promotion) | Allocation runtime integration remains |
8.19 [D:8.18] |
Casual lineup (2–6 humans, bot backfill) | Queue candidate selection, opt-in 10 s backfill proposals, reconnect/leave penalties, live integration remain |
8.20 [D:8.18] |
Ranked admission (six unique verified humans) | ArenaRegistry integration and allocation wiring remain |
8.21 [D:8.5,8.20] |
Rating core (Glicko-2, weights, transactional updates) | Live maintenance/DB execution remains |
8.22 [D:8.21] |
Ranked profile (provisional games, tiers) | Persisted tier policy, client UI, reconnect transport remain |
8.23 [D:8.21] |
Ranked season policy (compression, rollover) | Live maintenance/DB execution remains |
8.24 [D:8.9,8.20,8.21] |
Ranked connection policy, reconnect lease, abandon ladder | Live PostgreSQL/process-restart/outage execution remains, blocked by Docker storage |
8.25 [D:8.10,8.24] |
Result policy (workload-bound, idempotent, transactional) | Production credentials, Agones annotation persistence/reconciliation, integrity-evidence adapters remain |
8D — Agones, allocation and regional scaling
| # | Task | Remaining |
|---|---|---|
8.26 [D:8.1,8.6,8.12] |
Provider-neutral Fleet, EU/NA overlays, RBAC | Operator secret/image replacement, second-provider fixtures, edge/DNS, SDR POP/cert/public-UDP overlays remain |
8.27 [D:8.26] |
Supervisor package (Agones discovery, Ready transition) | Metadata watch, real Agones annotation/shutdown confirmation, emulator integration remain |
8.28 [D:8.6,8.27] |
Process-ready/Agones-Ready separation, control-plane registration | Remaining gates are live Agones annotation/shutdown behavior and production cluster readiness — see task 8.49 |
8.29 [D:8.26,8.27] |
Dynamic port/SDR env propagation | Real Agones dynamic/passthrough mapping, POP/cert/firewall/NAT, multi-match fixture remain |
8.30 [D:8.18,8.26,8.28,8.29] |
Allocation leasing, compatibility validation | Full unknown-outcome cluster recovery and signed roster metadata remain |
8.31 [D:8.9,8.30] |
Signed assignment/roster persistence, player recovery | This is the actual root blocker of the whole allocation-to-connect pipeline (see §0). store.SaveAssignment/SaveAssignments/SaveVerifiedAssignmentRoster and allocator.Service.PublishRoster are built and tested but never called from allocator/worker.go, cmd/allocator, or anywhere else in production — only tests seed the table directly. A real match cannot advance past PROCESS_READY. Closing it needs new security-relevant design: a join-signing key shared between the allocator (to sign) and the game server (fleet.yaml already mounts one for verification via --join-authorisations-key-file, but no control-plane binary has a matching signing flag), roster-digest computation, and per-player domain.JoinAuthorisation construction via the already-built domain.SignJoinAuthorisationHMAC. Flagged rather than fixed at the user's explicit direction, pending a decision on that design |
8.32 [D:8.2,8.26,8.30] |
FleetAutoscaler baseline, Ready buffer | Regional on-demand node pools/failure domains, pre-pull rollout, warm-allocation p95/p99, N+1 certification remain |
8.33 [D:8.26,8.32] |
Fleet scheduling, zone spread | Regional node pools, forced node-loss testing, measured N+1 headroom remain |
8.34 [D:8.28,8.29] |
Native x86_64 benchmark of boot-to-process-ready/assignment-ready, p99 CPU/RSS/network, node cap with 30% headroom | Not started |
8.35 [D:8.17,8.19,8.20,8.30,8.31] |
Admission lease, disconnect/reconnect generations, no-show/bot policy | Live PostgreSQL execution, allocated process termination evidence, real Agones multi-client verification remain |
8.36 [D:8.10,8.25,8.28,8.30] |
Authenticated drain, PodDisruptionBudget | Live 300 s/285 s lifecycle, PDB/Fleet drain, infrastructure-abort classification remain |
8.37 [D:8.5,8.10,8.25,8.26,8.31] |
Horizontally scaled primary control plane + warm standby, EU/NA fleets, RPO ≤5 m/RTO ≤30 m | Not started |
8.38 [D:7.7,8.26,8.36,8.37] |
Provider migration (restore, coordinator trust, switch allocations, drain old) | Not started; needs Valve approval for both providers' EU/NA POPs/certs and public UDP |
8E — Client experience and recovery
| # | Task | Remaining |
|---|---|---|
8.39 [D:8.3,8.14,8.17] |
MatchmakingState/ControlPlaneClient, queue/proposal UI, targeted revisioned events |
Live PostgreSQL-backed dispatcher/fan-out verification remains |
8.40 [D:8.3,8.14] |
Revisioned event stream, REST resync, outbox dispatcher | Allocator and Redis fan-out live verification remain |
8.41 [D:7.8,8.9,8.31,8.40] |
Player-scoped assignment API, connect_to_assignment() wiring, join-authorisation verification in MatchNet |
SDR relay-ticket installation and live Agones cluster integration remain |
8.42 [D:8.22,8.23,8.24,8.40] |
RankedProfileState, backend-authoritative rating/tier display |
Committed revision after reconnect, abandon status, season countdown remain dependent on live auth/backend events and Godot runtime verification |
8.43 [D:8.39,8.40,8.41] |
Error/expiry UX, generic mutation retry, version-mismatch and failed-reconnect messaging | Long-running worker integration (§8.16) remains |
8F — Observability, verification, cost and rollout
| # | Task | Remaining |
|---|---|---|
8.44 [D:8.3,8.4,8.28,8.31] |
Structured logging, redaction | Local complete; production gate open — production metrics/traces backend and dashboard/alert routing remain |
8.45 [D:8.2,8.44] |
SLO window checks, API latency histogram | Local complete; production gate open — production scrape configuration, alert routing, wait/MMR/proposal/flood/cost series, runbooks remain |
8.46 [D:8.5,8.7,8.9,8.10,8.14,8.18,8.21,8.23,8.25] |
Go unit/race/fuzz coverage, local verification gate | Live matcher-worker-under-load-during-failover integration remains |
8.47 [D:8.7,8.30] |
Offline testkit (fake Steam, fake allocation) | Live exhaustive matrix and production Steam remain |
8.48 [D:8.10,8.14,8.17,8.18,8.27,8.31,8.35,8.47] |
Allocated Compose end-to-end (queue → proposal → allocation → assignment → result) | Local complete; production gate open — real Agones/kind and production evidence remain open |
8.49 [D:8.25,8.26,8.28,8.29,8.30,8.31,8.35,8.36] |
Disposable kind+Agones cluster runner | CI/live evidence for production control-plane registration, roster/no-show, both readiness stages, races, multi-match node, result-pending reconciliation, drain, rollback remains open. Blocked locally on Docker storage/kind/Helm availability |
8.50 [D:8.25,8.37,8.43,8.49] |
Chaos recovery (stale allocation, no-penalty requeue) | Local complete; production gate open — 100 ms RTT/jitter/loss, matcher/client restart, game-pod death, node drain, Redis failover, control-plane loss, live chaos evidence remain |
8.51 [D:8.17,8.18,8.30,8.31,8.45] |
10,000-client API load gate | Local complete; production gate open — PostgreSQL saturation, durable matcher fencing under load, forecast launch concurrency ×2, replica scaling remain live infrastructure gates |
8.52 [D:8.32,8.34,8.45,8.51] |
Per-replica + shared regional allocator quota | Real image digest/secrets, measured regional cost model, threshold tuning, denial-of-wallet rehearsal remain |
8.53 [D:7.8,8.13,8.38,8.45,8.46,8.48,8.49,8.50,8.51,8.52] |
Fail-closed release-gate promotion validator | Actual reports, production rollback rehearsal, regional playtests, live promotion remain open |
Implementation invariants for every task above:
- Matchmade mode is opt-in; every new
ServerConfigdefault preserves the existing community-server path. compose.phase6-smoke.yml,make verify-phase6, andmake verify-enet-integrationare not repurposed or weakened.- Production uses ticketed Hosted Dedicated Server SDR; ENet remains the deterministic local/CI and direct-IP path.
- One process serves one match. Warm processes/nodes absorb startup variance; capacity and cost are determined from 8.34 measurements, not old estimates.
- Design changes first update
docs/MATCHMAKING.mdand dependencies.
8. What needs refactoring, not extending
Historical note: this table described Phase 0's non-networked refactors, all of which are now implemented (see §7's Phase 0 summary). Kept for the underlying reasoning where it's still relevant to Phase 7/8 work touching the same files.
On Engine.time_scale: replaced with camera-based effects in
single-player as well, so there is one code path and one game feel to
maintain rather than a networked variant that drifts away from the
single-player one.
What does not need surgery: the ShipController seam, the Arena/GameMode split, code-driven spawning, group-based discovery, and the dumb Goal sensor all extend cleanly. CLAUDE.md's claim about the three load-bearing seams is accurate — they hold. rl_ship_controller.gd is already the remote-input controller (a public action field that something else writes, pulled each tick), so no new class was needed for it.
9. Godot 4.7 + Jolt gotchas
ENetMultiplayerPeer.server_relaydefaults totrue— clients can RPC each other through your server. Set itfalse.MultiplayerAPI.poll()runs on the idle frame, so anrpc()from_physics_processwaits up to a full frame — andEngine.max_fps = 60on the server is what creates that delay on the return leg. Take manual control. ~16–33 ms of round-trip, for ~10 lines.- Jolt sleeps bodies. A ship corrected to near-zero velocity can sleep and then ignore
state.linear_velocitywrites.can_sleep = falseon Ship and Ball. - Teleporting a rigid body:
state.transforminside_integrate_forcesis the only path with no frame of lag.set_deferred("global_transform", …)lands between frames and interacts badly with Jolt's sleep/wake ordering. reset_physics_interpolation()is not automatic forstate.transformwrites (it is when you setglobal_transformdirectly). Call it explicitly, on the body and on$Visual.physics_jitter_fix = 0.0does not give you "a flat 60 Hz." You still get occasional 0-tick and 2-tick frames, because frame time is never exactly 16.667 ms. The real reason to set it to 0 is that you never want a tick's input delayed by the accumulator smoother. The send path must therefore transmit both ticks' actions on a 2-tick frame — redundancy-4 covers this, but only if you actually send both._integrate_forcesis not called on frozen bodies, so remote ships never pullget_action()— henceset_visual_action. UseFREEZE_MODE_KINEMATIC, notSTATIC, or contact velocity transfer breaks.- Never write
linear_velocityto a frozen body — Godot/Jolt zeroes and holds it. Engine.max_physics_steps_per_framedefaults to 8. If a server tick overruns 16.7 ms the accumulator backs up and the next frame runs multiple ticks, spiking CPU further. Log overruns.- ENet channel indices are offset by Godot's reserved system channels — verify the mapping empirically.
- ENet peer timeout defaults to ~5 s. Tune via
ENetPacketPeer.set_timeout()for faster drop detection. - Jolt is not bit-deterministic across platforms or across differing contact orderings. Never rely on it anywhere, including in "obviously safe" places like a client-side goal check.
dedicated_server=trueexports strip visual resources. Verify against a real stripped build.- MTU: ENet fragments above ~1400 B. At 219 B/snapshot there is ~6× headroom; recheck if per-body cosmetic state is ever added.
- RPC NodePath caching: the first
rpc()to a node sends the full path, later calls send a cached int. Routing hot paths through autoloads warms the cache once at connect and never invalidates it on scene change. - Physics tick rate is 60 for v1 — and must never be a literal. Every policy in
Game/bots/is tick-coupled throughship.gd:450's_tick_scaled(defined at a 60 Hz reference) andai_ship_controller.gd'sreaction_ticks, so raising it toward Rocket League's 120 invalidates every trained model and halves server density. But it is the largest single latency term left (§5.4), so it will be revisited: everything derives fromTICK_HZ, so that day is a config change plus a retrain. Node3D.get_global_transform_interpolated()is the only correct way to track a physics-interpolated body from_process.global_transformreturns the last physics tick's pose, so a per-frame camera reading it chases a 60 Hz staircase. Per the engine docs the method "creates an interpolation pump… the first time it is called" — call it once before anyreset_physics_interpolation()on that node, or the first hard snap streaks (§4.5).- Physics interpolation covers transforms only.
camera.fov, shader parameters, light energy and anything else written from_physics_processsteps at 60 Hz on a 240 Hz display. Either write them from_processor accept the stepping deliberately. display/window/vsync_modedefaults to enabled (FIFO) andmax_fpsto uncapped. Neither is set inproject.godot. FIFO present latency is 1.5–3 refresh intervals depending on swapchain image count (2 vs 3) and whether the present queue is full — §5's tables use the optimistic 1.5, which assumes the renderer is not GPU-bound. The model does not hold below refresh, where a missed vblank under strict FIFO halves the effective rate and roughly doubles present latency. Prefer Adaptive as the default, not Mailbox (§5.4). (Swapchain image count per platform needs empirical verification.)Engine.max_fpsis a throttle, not a frame pacer. It pads each frame with a post-frame sleep; it has no vblank phase lock. Caps that are not integer divisors of the refresh rate beat against scanout, and combining a cap with an active vsync paces worse than either alone (§5.4). Derive the offered caps fromDisplayServer.screen_get_refresh_rate().DisplayServer.window_get_vsync_mode()echoes your request, not the driver's grant. There is no GDScript API for the negotiatedVkPresentModeKHR, so a UI cannot honestly report what was applied. Show a live fps readout instead and let the player infer it.Engine.max_physics_steps_per_frame = 8is a client problem too, not just a server one (gotcha 9). A client hitching to 20 fps runs 3 ticks per frame, and each of those frames also runs the per-frame camera rig and remote-visual sampling. Set it to 4 client-side. On a multi-tick frame the send path must transmit every tick's action (gotcha 6) — §4.3's_physics_processsampling does this naturally, but nothing else guarantees it.hint_screen_textureforces a full-screen backbuffer copy on every frame the node is drawn, regardless of what the shader then does with it. Branching inside the shader saves taps, not the copy. Hide the node when the effect is at rest.physics_jitter_fixmatters less the higher the frame rate. Its purpose is smoothing when frame rate ≈ tick rate; at 240 fps against 60 Hz physics most frames run zero ticks and the accumulator is never near an edge. Gotcha 6's reasoning for setting it to0.0still holds, but do not expect a visible difference on a high-refresh machine — test that change at 60 fps.MultiplayerAPI.multiplayer_peer's default value is anOfflineMultiplayerPeersentinel, notnull. Resetting it withmultiplayer_peer = null(rather than a freshOfflineMultiplayerPeer.new()) leaves the API in a state distinct from its own default and is a known source of "the server never seespeer_connected,get_peers()stays empty" bugs (godotengine/godot#81540). Always reset to a realOfflineMultiplayerPeer.- Don't tear down a peer the instant its own connect signal fires.
connected_to_server(client-side) fires once the client's local view of the handshake completes, but the final ACK the server needs to consider its side complete may not have hit the wire yet — closing the peer or quitting the process in the same callback can drop it, and the other side then never seespeer_connected/connected_to_serverat all, even though your own side looked successful. This reproduced on every attempt until fixed and is easy to misdiagnose as a server-side bug (the server-side symptom —get_peers()staying empty — is identical to gotcha 25's). Give at least one frame (in practicetests/net_smoke.gduses 0.3 s) between a fresh connect signal and callingshutdown()/quit(). change_scene_to_file()must be called on (or from a descendant of) the actualget_tree().current_scene, and never synchronously from_ready(). (a) instantiating a scene as a plain child of a driver node, rather than loading it as the real current scene, breaks its own disconnect-handlingchange_scene_to_file()calls with a silent hang. (b) callingchange_scene_to_file()(oradd_child()onget_tree().root) synchronously from inside_ready()throws "Parent node is busy … Consider using.call_deferred()", because the tree is still mid-traversal adding the very node whose_ready()is running.ENetMultiplayerPeer'sconnection_failedsignal is not bounded to anything a UI should make a player wait for. Verified empirically: against a genuinely refused loopback connection,connection_failedhad still not fired 14 seconds in. Don't rely on it alone to end a "Connecting…" state — run your own app-level timeout.- A
MultiplayerPeer's "am I a client" flag turns true the instantjoin()/create_client()is called, not once the connection actually completes. Anything gated on that flag alone will try torpc_id()on a peer that's stillCONNECTING— or has already failed. Gate on the peer's actualget_connection_status() == MultiplayerPeer.CONNECTION_CONNECTED, not just the higher-level intent flag. load()on a.gdfile with a parse/compile error does not returnnull. It returns a non-null but uninstantiableGDScriptresource, soif script == nullsilently fails to catch the failure. The real guard isScript.can_instantiate().- An
@rpcmethod named_inputcollides withNode's built-in_input(event: InputEvent)virtual. On an autoload especially, treat any bare virtual-sounding method name (_input,_process,_ready,_unhandled_input, …) as reserved regardless of what you intend it to do — a signature mismatch there doesn't fail locally, it fails the whole autoload. - Disabling automatic multiplayer polling is global, not autoload-scoped — every scene that touches an RPC, not just
NetworkManager-adjacent code, must callNetworkManager.poll()itself every frame it wants traffic to move. If a scene sends or receives RPCs and nothing arrives with no errors at all, check for a missingpoll()before anything else. - A request/response fallback for a one-shot broadcast can double-deliver, and the receiving handler must be idempotent. Any handler reachable via both an original broadcast and a "resend on request" path needs its own guard rather than assuming "only sent once" from the RPC design alone.
- An
Area3D'sbody_enteredsignal fires as part of physics tick N's own step, strictly before tick N's_physics_processcallback — not "on the next frame." If you actually need next-tick-or-later semantics, compareEngine.get_physics_frames()against the tick the flag was set on and require strictly-greater, not just "check a boolean at the top of_physics_process." - A queued
queue_teleport()can take one tick longer to land than "the very next_integrate_forces" suggests, when the call originates from a signal handler mid-physics-step rather than from a_physics_processcallback. Don't assume queued-teleport timing without checking a real tick-by-tick log for your specific call site. NetworkManager.get_server_time_estimate_ms()(and anything derived from it) is not just imprecise before the first clock pong lands — it's actively wrong in a way that can persist far longer than the sync window implies. Always gate recording, not just consuming, anything derived from this estimate onrtt_ms >= 0.0.- Anything that deliberately delays an RPC dispatch must re-validate its target at fire time, not just at the moment it was scheduled. Any deliberate-delay layer needs to re-check both "do I still have a real peer at all" and "is this specific target still in
get_peers()" inside the delayed callback itself, immediately before dispatching. - A GDScript lambda captures an enclosing local variable BY VALUE at the moment the lambda is created, not by reference. Capture a container instead of a value —
var disconnected := [false]anddisconnected[0] = trueinside the lambda — since capturing anArray/Dictionary/Objectcaptures a reference to the same instance. - A fixed-size ring buffer fed by an unbounded-rate producer needs an explicit resync path, not just "wait for the next expected slot." If a ring's producer and consumer rates aren't provably bounded relative to each other, the consumer needs a way to detect "the data I'm waiting for no longer exists in the ring at all" and jump directly to what's still available, rather than assuming "keep waiting" is always eventually correct.
- A client-owned adaptive control loop must react to the actual ground-truth signal it's regulating, not to its own memory of past decisions. When a control loop's condition for acting can be satisfied or blocked by state the loop itself controls, rather than by the environment it's meant to respond to, it can silently stop responding to the environment.
- A "consecutive N over-budget windows" streak counter that hard-resets to 0 on any single clean window is trivially evaded by a duty-cycled attacker. A leaky-bucket accumulator is immune to the same evasion by construction, since it doesn't matter how the excess is distributed in time — only the sustained average matters.
- Two counters that don't share an epoch must never be compared directly, even when both are monotonically increasing integers that "look like" the same kind of thing. Bound a value against another value that shares its own actual epoch, not against a same-typed number from a conceptually different clock.
- A regression test that doesn't independently exercise the specific mechanism it claims to gate will pass even when that mechanism is completely broken. A CI gate's assertions should trace back to the specific claim in the task's own acceptance text, not just "the match ran and didn't crash."
- When adding a "is this connection still healthy" check to a test, sample it while the peer is still actively connected — not after its own end-of-run disconnect, which produces symptoms indistinguishable from the bug being checked for.
- Two fixes landed in the same commit, each individually correct in isolation, can share a variable and silently cancel each other out — and a fix's own unit test can miss it by testing the mechanism in isolation from the thing that defeats it. When two fixes in the same round touch the same subsystem, explicitly re-test the combination end-to-end — passing tests for each fix individually is not evidence the pair composes correctly.
- A guard that bounds an incoming value against the consumer's position, rather than against the producer's own epoch, re-introduces exactly the "consumer can never catch up past a stall" failure it's often added specifically to prevent. Bound against whichever side of a producer/consumer pair is not the one already known to be falling behind.
- A trace that holds its inputs steady cannot falsify anything about which sequence a prediction is filed under — and "we hold thrust for 60 seconds" describes almost every prediction test people write. When a test is meant to validate a label, an index, or a phase relationship rather than a magnitude, the trace has to change that quantity frequently; a steady-state trace validates the magnitude and silently asserts nothing about the label.
- A guard whose bound is derived from a value only the ACCEPTED path can advance is a latch, not a guard. The property to check when writing a guard like this is not "is the bound correct?" but "if this guard rejects everything from now on, what advances the bound?" If the answer is "an accepted packet", it needs an independent escape path regardless of how well-chosen the bound is.
- Advancing a consumer cursor past data that has not arrived is not a lossy shortcut — it is permanent, because the producer-side filter then rejects the very data being waited for. Only give up on an expected item once strictly newer data proves it lost; "it hasn't arrived yet" and "it will never arrive" are different states and must not share a code path.
- A metric that stops sampling during a failure will report that failure as healthy. Every rate-shaped assertion needs a companion assertion on the denominator, or an outage silently becomes an absence of evidence and then evidence of absence.
- An architectural blocker inherited from a previous session is a claim to verify, not a premise to build on. Reconstruct the failing invariant from the code and reproduce it against a control before accepting a scope estimate attached to it — especially when the recommendation arrives without the cheaper alternative recorded as tested.
10. Testing
Editor. Debug → Run Multiple Instances, 2–3 instances with per-instance args (-- --server, -- --connect 127.0.0.1:27015) and --position so windows don't stack.
CLI.
godot --headless --path Game res://scenes/server_boot.tscn -- --port 27015 --team-size 1 --auto-start
godot --path Game -- --connect 127.0.0.1:27015 --name Alice
CI smoke test. Headless server plus two headless --test-bot clients, driven by the existing AIShipController. Asserts:
- snapshots received ≥
N * snapshot_hz * 0.9 - own-ship prediction error p95 < 0.5 m, p99 < 2.0 m, hard-snap count < 3
- final score identical on the server and both clients
- no
push_erroremitted (scrape stderr)
Network conditions. net_sim.gd is first-class: seeded so failures reproduce, works in CI, needs no display, and can be applied asymmetrically — which OS tools make painful. tc netem / Network Link Conditioner / clumsy for a pre-release realism pass only. A real remote host once per phase from Phase 4 onward is the only true test of the jitter buffer's adaptivity.
Unit tests. godot --headless --path Game res://tests/test_runner.tscn. High-value targets, all zero-engine-state: codec quantise/dequantise round-trip and bounds; quaternion max error; snapshot pack→unpack identity; input packet framing; jitter-buffer policy against scripted arrival traces; ShipAction.copy() non-aliasing. These are exactly where a bug is invisible in play and catastrophic in aggregate.
11. Flagged, not solved
120 Hz simulation — deliberately deferred, not dismissed. §5.4 and §5.6 record what it would buy (≈21 ms of world response once L1 has taken the interpolation buffer out, plus ≈8 ms of own-ship feel — the difference between ≈127 ms and ≈107 ms), what it costs (a full bot retrain, half the server density, double the bandwidth), and the one rule that keeps the door open: TICK_HZ, never 60.
The latency gap to the reference has a plan but not yet an implementation. §5.2 lands at ≈174 ms as designed; §5.6 routes that to ≈127 ms and ≈103 ms (tasks L1–L4 plus 120 Hz simulation), against ~90–110 ms for the reference class at the same RTT. Every figure in §5.6 is arithmetic on the budget, not a measurement or a shipped change — L1–L4 remain unimplemented. Beyond that the residual is RTT, which is a server-siting problem (§6, Phase 8) rather than a code one and is worth more than every remaining code lever combined.
Audio. The runtime has dependency-free procedural placeholder hooks for UI, countdown, engine/thrust/turbo, impacts, wall contacts, goals, and camera/gameplay events. TODO.md tracks authored engine/turbo/impact/wall/goal/crowd/music assets and production mixing/QA as still open.
Split-screen. Tracked separately in TODO.md; unrelated to this effort, though the camera-outside-the-ship structure that enables it is the same structure this plan relies on.
Graphics: baked GI (task 0.26) and low/mid-tier hardware profiling. See §5.5 and §5.7 — real but smaller wins than originally assumed on reference-class desktop hardware; unmeasured on low-end/integrated GPUs.
Not locally certifiable from this workspace, and open prerequisites rather than done: Valve/GodotSteam credentials and hosted SDR (Phase 7 tasks 7.1–7.8), live Agones/kind lifecycle (tasks 8.30–8.38, 8.49), public-network chaos/load/cost/release gates (tasks 8.50–8.53), and real-hardware graphics profiling on low/mid-tier GPUs. make verify-kind-agones is the committed runner for 8.49; it has not yet completed a full run against a real cluster from this workspace (blocked on local Docker/kind/Helm resource availability, not a code gap). TODO.md's AI-training and presentation tasks remain separate from multiplayer.