Files
CosmicClash/docs/MATCHMAKING.md
T
Josh Creek 4ffa1543cc docs: design per-match server autoscaling, with measured boot time
Casual/ranked queues need servers allocated per match and shut down
afterwards, so cost is incurred only while a match runs - while the
existing Docker and CI gates keep passing unchanged.

Measured against the repo's own cosmicclash-server image rather than
estimated: the runtime image is ~148 MB of content, and boot to the
server_started line is ~870 ms on the container's own clock. That was
taken under x86_64 emulation on an arm64 host, so it is a pessimistic
bound and is recorded as one - it needs re-measuring on native Linux
before it sets any timeout.

Two findings that would each break a naive implementation, both hit
while taking that measurement:

- Godot's stdout is block-buffered off a TTY. A detached container logs
  nothing at all - server_started does not appear even after 35s - so an
  orchestrator readiness probe that greps the log hangs forever. Probe
  the UDP socket or flush explicitly.
- --port defaults to 7777 and the Dockerfile hardcodes EXPOSE 7777/udp,
  so several matches cannot share a host without a port range or an
  address per match. Being UDP, L7 ingress routing does not apply.

Also records the honest tension in 'only pay during a match': a server
must listen before players connect, and image pull plus scheduling can
dwarf 870 ms, so the recommendation is match-level scale-to-zero over a
small warm node pool rather than node-level scale-to-zero.

The rule for keeping verify-phase6 and verify-enet-integration green:
every allocation feature is opt-in via a ServerConfig flag defaulting to
current behaviour, with a second Compose file rather than mutating
compose.phase6-smoke.yml.
2026-08-31 18:45:05 +01:00

13 KiB
Raw Blame History

Matchmaking — casual and ranked queues

Design scope for online casual and ranked play. This is a 1.0 launch blocker, not a post-launch addition.

Nothing described here is implemented yet. This doc exists to record the decisions and the reasoning before code is written; per-task implementation evidence belongs in multiplayer-todo.md once work starts, and the live checklist lives in multiplayer-next.md.

The model change

The multiplayer that exists today is a community-server model. A dedicated server runs forever: it waits for --min-players by roster, counts down --start-countdown, loads the next arena from the rotation, plays a match, returns to the lobby, and repeats (server_match_loop.gd). Players reach it by direct IP, and after Phase 7 by a Steam server browser. The server is the durable thing and players come and go around it.

Queued matchmaking inverts that. Players are the durable thing: they enter a queue, a matchmaker groups them by rating and region, and a server is allocated for that one match and torn down afterwards. Both models can coexist — community servers via the browser, queues via the matchmaker — and they should, because the server browser is already most of the way to done.

Hard prerequisite: verified identity

Ranked cannot ship before Phase 7's Steam auth tickets.

Slot reclaim is currently keyed by display name (see --slot-reservation-seconds, and the known-issues list in multiplayer-next.md). A rating attached to a spoofable identity is worse than no rating at all: it is trivially farmed, and it invites players to invest in a ladder that cannot be defended. "Ranked is critical" therefore raises the priority of Steam identity rather than routing around it.

Casual queueing has a weaker requirement — it still needs stable identity for abandon penalties and ban enforcement, but the cost of a compromise is lower.

Architecture

Decided: Steam for identity, a project-owned backend for everything else.

This reverses the "no backend" position stated in TECH_STACK.md and README.md, which described the state of the project before matchmaking was scoped. The dedicated server remains a Godot export; the new service is separate from it.

The alternative — Steam-native matchmaking (lobbies plus Leaderboards or User Stats as the rating store) — was rejected on two grounds. Steam lobby matchmaking has no real concept of a skill distribution to match against, and Leaderboards are a display surface rather than a rating store with the transactional guarantees a ladder needs. It would also permanently bind the game to Steam, foreclosing other platforms.

Components

Component Runs where Responsibility
Steam auth ticket validation backend Turn a client-supplied ticket into a verified SteamID via the Steamworks Web API. The only trusted source of identity.
Queue / matchmaker backend Hold queued players per playlist and region; form matches on rating proximity with a widening tolerance over wait time.
Rating store backend (DB) Per-identity, per-playlist rating and match history. Written only by the backend, never by a game client.
Server allocator backend Start a dedicated-server instance per formed match, hand its address to the matched clients, reclaim it on exit.
Dedicated server Godot export Unchanged simulation. Gains a mode where the roster is assigned rather than open, and reports a result at the end.
Game client Godot Queue UI, estimated wait, accept/decline, connect-on-assignment, post-match rating delta.

What already exists and gets reused

The server side needs less new work than it looks:

  • --max-matches=1 already makes the server drain and exit(0) after a single match. That is precisely the lifecycle a per-match allocator wants; it was built for CI, and it generalises for free.
  • ServerConfig is a single-source-of-truth flag table with strict validation — new allocation flags are declared in one place and are automatically parsed, type-checked, config-file-backed and documented.
  • --min-players / --start-countdown / --slot-reservation-seconds are the match-formation primitives, and they already count roster members rather than raw peers.
  • MatchNet's roster already survives the lobby→match transition, which is the structure an assigned roster slots into.
  • MatchState already has a legal-transition table with wire-stable integer values, so new lifecycle states append cleanly.

What is genuinely new

  • The backend service itself (process, deploy, DB, ops) — nothing like it exists in this repo today.
  • Server-authoritative match results: the dedicated server must report the outcome to the backend over a channel a client cannot forge. This is the first non-ENet/SDR network path in the project (see TECH_STACK's "no HTTP layer" note, which this supersedes).
  • An assigned-roster server mode: only the matched SteamIDs may take a slot, replacing the current first-come model.
  • Client-side queue UI and the accept/decline flow.

Server orchestration and autoscaling

Requirement: game servers scale horizontally and automatically, spin up fast on demand, serve exactly one match, and shut down — so cost is incurred only while a match is being played. Docker and the existing CI gates must keep working unchanged.

Why this is achievable: the server is already shaped for it

Two properties of the current build make per-match allocation practical rather than aspirational:

  • The container is small. The server image target is a slim ubuntu:24.04 runtime with three shared libraries and the exported binary — about 148 MB of content, not the ~2.6 GB godot-ci build image. Pulling it onto a fresh node is cheap.
  • Boot to listening is sub-second. Measured on this repo's cosmicclash-server:latest: ~870 ms from container start to the server_started log line, averaged over three runs, read from the container's own clock. That measurement was taken under x86_64 emulation on an arm64 host, so it is a pessimistic bound — native x86_64 Linux should be faster. Re-measure on the real target before setting timeouts.

Combined with --max-matches=1, which already drains and exit(0)s after a single match, the lifecycle the allocator needs mostly exists: start container → serve one match → process exits → orchestrator reclaims.

The cold-start tension, stated honestly

"Only pay during a match" and "a player never waits" are in tension. A server must be listening before the matched players connect, so some cost always precedes the match. Sub-second boot makes the gap small enough that a pure scale-to-zero design is plausible — but the risk is not the container, it is everything around it: image pull on a cold node, scheduler placement, and network/port programming can each dwarf 870 ms.

Recommendation: scale to zero at the node level is the wrong target; scale to zero at the match level is the right one. Keep a small warm pool of nodes sized to the current queue depth, and start a per-match container on demand within it. The per-match process genuinely exists only for the match; the node pool absorbs the cold-start variance. Revisit only if measured allocation latency on real infrastructure shows the warm pool is unnecessary.

Findings that block a naive implementation

Readiness cannot be detected from the log line. Godot's stdout is block-buffered when it is not attached to a TTY. Run the server image detached without -t and docker logs shows nothing at all — the server_started line does not appear even after 35 seconds, because the buffer never flushes. An orchestrator readiness probe that greps for that line will hang forever, and this was reproduced directly while measuring the boot time above. Either probe the UDP socket instead, or make the server flush explicitly. This also means container logs are not a reliable observability channel for a short-lived match server; treat log shipping as a separate problem.

One fixed port per container does not scale on a shared host. --port defaults to 7777 and the Dockerfile hardcodes EXPOSE 7777/udp. Packing several matches onto one node needs either a port range allocated per container, or one address per container. This is a UDP service, so the usual HTTP ingress/L7 routing answers do not apply — the allocator must hand the client a concrete host:port.

The match cannot start on a schedule the players do not control. Today the loop waits for --min-players then counts down. An allocated server is told which identities to expect, and needs a no-show timeout: if a matched player never connects, the server must abandon and exit rather than sit idle burning the cost this design is trying to avoid.

Keeping Docker and CI green

The existing gates must not regress. make verify-phase6 builds the export, runs it in Compose, joins two headless clients and asserts both saw both goals and that the arena rotated between matches; make verify-enet-integration runs the source-build ENet matrix. Both depend on current behaviour: compose.phase6-smoke.yml hardcodes --port=7777, relies on first-come slot assignment, and uses --max-matches=2 to prove rotation.

The rule that keeps them passing: every allocation feature is opt-in via a new ServerConfig flag whose default reproduces today's behaviour. An assigned roster, a no-show timeout and result reporting must each be inert unless explicitly enabled. ServerConfig is built for exactly this — a flag declared once is parsed, validated, type-checked, config-file-backed and documented — and tests/cases/ can cover the new parsing without a live server. A second Compose file should cover the allocated-match path rather than mutating the Phase 6 one, so the community-server model stays tested alongside the matchmade one.

Open questions

  • Orchestrator. Kubernetes (with Agones, which exists for precisely this game-server lifecycle), Nomad, or direct cloud-API container starts. Not chosen. Agones is the strongest default because it models allocation, readiness and per-match lifetime natively.
  • Port strategy — port range per node versus one IP per match.
  • Bin-packing. SERVER.md's Phase 1 sizing estimate is 610 match processes per modern core and 150250 MB RSS each. That estimate predates any allocation work and should be re-measured under real concurrency before it sizes a bill.
  • Draining and deploys. How a server version rolls out without killing matches in flight.

Casual vs ranked

They are different playlists, not a difficulty toggle, and their rules diverge in ways that affect the server:

Casual Ranked
Rating Hidden, used only for matching Visible, with tiers
Backfill on disconnect Yes — keep the match playable No — the match is rating-bearing and must not change shape mid-way
Bots filling empty slots Acceptable (--fill-bots exists) Never
Abandon penalty Light (short queue cooldown) Real (rating loss, escalating cooldown)
Arena selection Full rotation Restricted set, so a variant nobody has practised can't decide a ladder match
Party / premade Unrestricted Constrained by size and rating spread

Note the arena constraint interacts with an existing rule: elevated-goal variants are Free-Play-only until a checkpoint trained on training_elevated.tscn is promoted (arena_registry.gd). Ranked's arena set should be drawn from "random": true arenas only.

Open questions

  • Rating algorithm. Glicko-2 is the default recommendation over plain Elo: it models rating uncertainty, which matters enormously for a small launch population where most players have few games. Not yet decided.
  • Team rating from individual ratings. How a 3v3 match's outcome distributes across six players is a separate design problem from the rating system itself.
  • Server cost. Allocated servers cost real money per match, unlike community servers that players host themselves. README.md's original note about a subscription to fund servers is suddenly load-bearing again. Population size and match length set the bill; this needs a number before launch, not after.
  • Region / ping policy. How much rating tolerance to trade for latency, and whether cross-region is ever allowed at low population.
  • Placement matches and whether ranked has a soft reset per season.
  • Backend language and hosting. Not chosen. It does not have to be C# despite the original README framing — that framing was aspirational and predates every real decision in this project.

Explicitly out of scope

Tournaments, in-game leaderboards beyond a personal rank display, cross-platform play with non-Steam identity providers, and spectator/observer tooling for ranked matches. None are precluded by this design; none are launch scope.