feat(multiplayer): wire allocator metrics discovery

This commit is contained in:
Josh Creek
2026-09-01 18:46:31 +01:00
parent 9a72dd5eab
commit 95e82cc719
5 changed files with 50 additions and 1 deletions
+15
View File
@@ -0,0 +1,15 @@
apiVersion: v1
kind: Service
metadata:
name: allocator
namespace: cosmic-clash
labels:
app.kubernetes.io/name: allocator
app.kubernetes.io/component: allocator
spec:
selector:
app.kubernetes.io/name: allocator
ports:
- name: metrics
port: 9091
targetPort: metrics
+1
View File
@@ -7,6 +7,7 @@ resources:
- network-policies.yaml
- control-plane-deployment.yaml
- control-plane-service.yaml
- allocator-service.yaml
- fleet.yaml
- fleet-autoscaler.yaml
- game-server-pdb.yaml
@@ -0,0 +1,20 @@
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
name: cosmic-clash-allocator
namespace: cosmic-clash
labels:
app.kubernetes.io/name: cosmic-clash
app.kubernetes.io/component: observability
spec:
selector:
matchLabels:
app.kubernetes.io/name: allocator
namespaceSelector:
matchNames:
- cosmic-clash
endpoints:
- port: metrics
path: /metrics
interval: 15s
scrapeTimeout: 5s
+1 -1
View File
@@ -1251,7 +1251,7 @@ the local/CI/community transport, not a silent production fallback.
| 8.49 `[D:8.25,8.26,8.28,8.29,8.30,8.31,8.35,8.36]` | **IN PROGRESS.** `scripts/verify_kind_agones.sh` creates a disposable kind cluster, installs pinned Agones, loads the real `game-server` image, applies the Fleet in an explicitly separate Agones-only supervisor/UDP readiness mode, and verifies readiness plus allocation of a dynamic UDP endpoint; `.github/workflows/agones-integration.yml` runs it for infrastructure changes and on demand | The cloud-free runner is committed and fails clearly when Docker/kind/Helm are unavailable. CI/live evidence for production control-plane registration, roster/no-show, both readiness stages, races, multi-match node, result-pending reconciliation, drain, and rollback remains open |
| 8.50 `[D:8.25,8.37,8.43,8.49]` | **IN PROGRESS.** `make verify-chaos-recovery` provides a disposable PostgreSQL + real testkit API + real maintenance flow: it restarts the API, injects a stale allocation, and verifies no-penalty requeue plus a durable participant-targeted lifecycle event | The API-restart/stalled-allocation slice is implemented and documented; 100 ms RTT/jitter/loss, matcher/client restart, game-pod death, node drain, Redis failover, control-plane loss, and live chaos evidence remain |
| 8.51 `[D:8.17,8.18,8.30,8.31,8.45]` | **IN PROGRESS.** The opt-in `make verify-multiplayer-load` gate drives 10,000 real HTTP queue-create requests through the service with 256 in flight and records p95/p99, plus 100 concurrent proposal formations through the real matcher/domain path; the handler and in-process ownership boundary are exercised without weakening normal tests | Local API load passes at p95 <250 ms in normal and race runs, and the matcher forms 100 unique proposals; PostgreSQL saturation, durable matcher fencing under load, forecast launch concurrency x2, and replica scaling remain live infrastructure gates |
| 8.52 `[D:8.32,8.34,8.45,8.51]` | **IN PROGRESS.** Allocator supports both an opt-in per-replica fixed-window quota (`--allocation-quota` / `--allocation-quota-window`) and an optional PostgreSQL-backed EU/NA quota table consumed inside the serializable allocation transaction before any provider call; idempotent replays do not double-count. The allocator now exposes bounded EU/NA Prometheus counters at an explicit `/metrics` listener (`--metrics-addr`), including quota denials, and the checked-in rule warns on regional denial activity | Normal/race/vet tests cover local quota exhaustion, window reset, region isolation, invalid input, atomic concurrent consumption, bounded metrics labels, and read-only endpoint behavior; migration/SQL coverage defines the shared quota boundary; production scrape wiring, measured regional cost model, threshold tuning, and denial-of-wallet rehearsal remain |
| 8.52 `[D:8.32,8.34,8.45,8.51]` | **IN PROGRESS.** Allocator supports both an opt-in per-replica fixed-window quota (`--allocation-quota` / `--allocation-quota-window`) and an optional PostgreSQL-backed EU/NA quota table consumed inside the serializable allocation transaction before any provider call; idempotent replays do not double-count. The allocator now exposes bounded EU/NA Prometheus counters at an explicit `/metrics` listener (`--metrics-addr`), including quota denials; the checked-in rule warns on regional denial activity, and Kubernetes Service/ServiceMonitor resources provide the discovery contract | Normal/race/vet tests cover local quota exhaustion, window reset, region isolation, invalid input, atomic concurrent consumption, bounded metrics labels, and read-only endpoint behavior; migration/SQL and manifest coverage define the shared quota/metrics boundaries; allocator Deployment provisioning, measured regional cost model, threshold tuning, and denial-of-wallet rehearsal remain |
| 8.53 `[D:7.8,8.13,8.38,8.45,8.46,8.48,8.49,8.50,8.51,8.52]` | **IN PROGRESS.** `scripts/verify_release_gate.py` provides a fail-closed promotion check for the ordered development → internal → casual canary → casual → provisional ranked → ranked stages, requiring an evidence report for SLO, security, cost, rollback, EU+NA playtests, and both legacy gates | Validator and adversarial tests cover skipped stages, unknown stages, missing gates, non-boolean gate values, and blank release IDs; the actual reports, production rollback rehearsal, regional playtests, and live promotion remain open |
Implementation invariants for every task above:
+13
View File
@@ -17,6 +17,8 @@ def verify(directory: Path, service_path: Path) -> None:
monitor = (directory / "prometheus-service-monitor.yaml").read_text()
rules = (directory / "prometheus-rules.yaml").read_text()
service = service_path.read_text()
allocator_monitor = (directory / "prometheus-allocator-service-monitor.yaml").read_text()
allocator_service = (ROOT / "deploy/k8s/base/allocator-service.yaml").read_text()
if "kind: ServiceMonitor" not in monitor:
raise ValueError("ServiceMonitor resource is missing")
@@ -33,6 +35,17 @@ def verify(directory: Path, service_path: Path) -> None:
if "interval: 15s" not in monitor or "scrapeTimeout: 5s" not in monitor:
raise ValueError("ServiceMonitor interval/timeout contract changed")
if "kind: ServiceMonitor" not in allocator_monitor or "name: allocator" not in allocator_monitor:
raise ValueError("allocator ServiceMonitor is missing")
if not re.search(r"(?m)^ - port: metrics$", allocator_monitor) or not re.search(r"(?m)^ path: /metrics$", allocator_monitor):
raise ValueError("allocator ServiceMonitor endpoint is invalid")
if "namespace: cosmic-clash" not in allocator_monitor or " - cosmic-clash" not in allocator_monitor:
raise ValueError("allocator ServiceMonitor namespace is not restricted")
if "kind: Service" not in allocator_service or "name: allocator" not in allocator_service:
raise ValueError("allocator metrics Service is missing")
if "name: metrics" not in allocator_service or "port: 9091" not in allocator_service:
raise ValueError("allocator metrics Service port is missing")
if "kind: Service" not in service or "name: control-plane" not in service:
raise ValueError("control-plane Service is missing")
if not re.search(r"(?m)^ - name: http$", service):