Two things that made the integration gate untrustworthy.
The retry budget was too small for expected contention.
TestPostgreSQLConcurrentIdenticalResultSubmission fires five identical
concurrent submissions and requires all five to succeed; it failed 4 runs
in 20. The error was retryable and retries did fire -- three attempts
simply was not enough. Contention here is normal rather than
exceptional: several game servers can submit results, and several
matchers can claim candidates, against the same rows at once. Raised to
five attempts, which is 0 failures in 40 runs.
Also jittered the backoff, but measured rather than assumed: my first
theory was a thundering herd, since the delay was exactly
RetryBackoff*(attempt+1) and every loser of a race woke at the same
instant. Isolating the two changes showed jitter alone moved 4/20 to
3/20, while the budget alone reached 0/20. The budget was the real
constraint. Jitter is kept because it costs nothing and its benefit
grows with the number of contending writers -- production is not capped
at five -- but the comment now says plainly that it is the smaller half,
so nobody inherits my wrong explanation.
Second, the integration scripts leaked one throwaway database volume per
run. --rm does reclaim anonymous volumes on a normal exit, but these
scripts force-remove the container from a trap, and `docker rm -f`
without -v keeps the volume. Sixty-four accumulated during this branch
until PostgreSQL stopped starting, surfacing only as the scripts' own
readiness timeout rather than as a disk error -- which is what the
"Docker storage exhausted locally" notes were really describing.
Measured at one volume per run before, zero after, across all five
scripts.