2-vCPU production test: 150 passes, 200 lags, 300 collapses
The numbers up front: Tencent Cloud SA2.MEDIUM4, 2 vCPU, 3.6 GiB and 2 Mbps fixed public egress, with Postgres and Caddy on the same host. 150 players pass the latency budget, 200 miss it, and 300 hit a 4.1-second p99; that is the starting point for isolating outbound delivery, optimizing the code, and rerunning the same ladder.
The answer first: 150 passes, 200 starts lagging, 300 collapses
On 21 August 2026 we drove app.bluffking.ai through the public internet, real DNS, Caddy and TLS. The load shape was one human owning a table with five bots: one session task per human, the costliest form of solo practice. Here, “no lag” means a server-side action round-trip p99 no higher than 150 ms.
The direct result: 150 players / 150 tables passed the budget; 200 / 200 missed it; 300 / 300 collapsed to a p99 of about 4.1 seconds. This was a stepped ladder, so the defensible claim is that the healthy range reaches at least 150 and the knee lies between 150 and 200. It is not a claim that player 151 always lags.
| Test subject | Production value | What it means |
|---|---|---|
| Cloud VM | Tencent Cloud CVM SA2.MEDIUM4, Hong Kong | KVM; not a local simulation |
| CPU | 2 vCPU · AMD EPYC 7K62 | No separate app CPU quota; app, Postgres and Caddy share both cores |
| Memory | 3.6 GiB usable | Hard 3 GiB app-container limit; Postgres and Caddy use the remainder |
| Network | 2 Mbps fixed public egress | The public-internet test client had about 45 ms RTT; theoretical egress cap is about 250 kB/s |
| Network gap | Actual egress / utilization during each test window was not sampled | Knowing the plan cap does not tell us whether the 300-table collapse saturated it |
The 2 Mbps figure is this production VM’s purchased public-egress plan, not a NIC line rate. Split evenly across 300 connections, it leaves about 6.7 kbps, or 0.83 kB/s per connection before TLS, TCP and WebSocket overhead. We did not record per-second egress, so this experiment tells us when table latency broke, but not whether app fan-out scheduling failed first or the 2 Mbps public uplink saturated first.
First: decide what “too slow” means, per surface
A single latency target is the most common mistake, and it is a mistake in both directions at once. Our own starting assumption was “keep everything under 500 ms.” That is simultaneously too loose and too strict.
The same millisecond costs a different amount depending on what the person is waiting for:
| What the user is waiting for | Budget we set | Anchored to |
|---|---|---|
| I tap CALL → my chips move | 150 ms server-side | 1 % of the 15 s action clock; the thing that decides whether the table feels alive |
| Someone else acts → I see it | 500 ms | Nobody is holding a button down waiting for it |
| A list of past hands loads | 500 ms | Our web client sets no fetch timeout at all, so a slow read hangs the UI rather than failing it |
| Login (memory-hard password hash) | 1500 ms | Paid once per session, and the hash parameters are at the security floor |
Note where each budget comes from: a constant already in the codebase, not taste. If you cannot point at the thing your budget is derived from, you do not have a budget, you have a preference — and preferences lose arguments at 2 a.m.
For us the practical upshot was blunt: 500 ms is about right for an API read and roughly 3× too loose for the action round-trip. A ceiling measured against the wrong budget is not a conservative estimate. It is a different number about a different question.
Second: the three ways we measured the wrong machine
We measured the ramp, not the steady state
The first run connected N clients and started timing immediately. That measures N users arriving, which is a real and interesting event, but it is not the same as N users being online — and it is the second one that decides how many people fit.
Once each rung warmed up for 20 s and then reset its samples, the difference was visible inside a single run. At 200 concurrent tables, the ramp window and the steady window disagreed by half:
200 tables, same run:
ramp window p99 379.9 ms
steady window p99 246.4 ms
Only the steady-window value is recorded in the ADR-105 ladder. The ramp-window figure is what the harness reported for the same run’s warm-up window, and that run’s raw output was not kept as a published artifact — read it as the size of the gap, not as a citable rung.
If you report the first number you will under-provision; if you report the second without saying so, you will be surprised by every marketing push. Measure both, label both.
We measured our own laptop
Halfway through the first ladder the tail started climbing while the server’s CPU sat still. The load generator was a single Node process on a developer machine — and on that machine, a stuck helper process was burning a full core.
The fix is not “be careful.” The fix is that a load generator must report its own health as a first-class output. Ours now records its event-loop delay every run:
generator_event_loop_delay_ms: { mean: …, p50: …, p99: 11.6, max: … }
// empty-box 300-table control run, 2026-08-21 (ADR-105)
That line is what lets you say the numbers are about the server. Flat at about 12 ms p99 across every rung from 50 to 300 clients, well under the reported p99 — so the queueing being measured is the server’s. Without it, every tail number is an accusation you cannot support.
We measured a machine that was still full of ghosts
When a player closes their tab, our server keeps their table alive for 10 to 20 minutes before an idle reaper collects it. That is a reasonable product decision — people reconnect. It is also a measurement trap: after four ladder runs the box was holding ~800 abandoned tables, quietly costing about 7 % of both cores.
And then do the thing that turns a suspicion into a fact: we re-ran the collapsing rung once on a completely empty box — zero live tables, checked in the database seconds before the run. It collapsed identically. The ghosts were real, and they were not the cause. A control run is cheap; an unexamined confounder gets quoted for years.
Two lessons, and the second is the useful one. Obviously: check what is already running before you attribute a number to your load. Less obviously: that ghost load is real in production too. A busy hour leaves a residue of tables nobody is sitting at, and they consume the same runtime as the live ones. Any capacity limit that counts only active users will be wrong during exactly the churny hour when it matters.
The full ladder: 150 passes, 200 misses, 300 collapses
Here is the ladder, solo tables (one human plus five bots — our most expensive shape per person), measured end-to-end over TLS from a client ~45 ms away:
| live tables | action p50 | action p99 | server-side p99 | app CPU |
|---|---|---|---|---|
| 50 | 43.0 ms | 115.9 ms | 70 ms | 15 % |
| 100 | 43.2 ms | 115.0 ms | 72 ms | 24 % |
| 150 | 43.9 ms | 161.4 ms | 118 ms | 34 % |
| 200 | 44.3 ms | 246.4 ms | 201 ms | 41 % |
| 300 | 85.6 ms | 4116 ms | 4071 ms | — collapse |
Read the last two rows together. Between 200 and 300 the p99 went up by a factor of seventeen, and the tables stopped making progress: the same 60-second window that produced 697 finished hands at 200 tables produced 148 at 300, and the run collected 90 latency samples where the rung below it collected 2994. All of that while roughly a core and a half sat idle.
Memory was nowhere near the machine limit either. In a separate host-sampled validation, app RSS peaked at 84 MB with 150 tables and 97 MB with 200; the empty-box 300-table control reached 53 MB before it aborted. Those values should not be turned into a cross-run linear memory model, but they do rule out “the 3.6 GiB was exhausted” as the explanation.
The logs narrow the failure to outbound delivery: the 300-table collapse window produced 199 fanout_backpressure_evict events. That counter proves the 64-frame outbound queue reached its backpressure path; by itself it does not prove what recovery the client observed during that run. The contemporaneous test report described frames being dropped, while today’s server path evicts the connection, requests a close, and has the client reconnect and resync from an authoritative Snapshot. We did not separately capture the version-specific close path in the load window, so we do not rewrite that historical observation as a proven reconnect. Database-pool waits peaked around 100 ms, Postgres stayed below roughly 0.1 core, and the generator event loop remained near 12 ms p99. They do not narrow it to one root cause. App fan-out scheduling can fill the queue; a saturated 2 Mbps public uplink can stop the same queue from draining, and this run did not sample egress utilization. Backpressure is the proven symptom; on the test date, scheduler contention and public-egress throttling were two candidates that had not been A/B-isolated.
Two consequences worth stealing:
- An autoscaling or alerting rule on average CPU will not fire before your users are in the collapse. Ours would have been comfortably green at the last healthy rung and still green while the tail was seconds long. Alert on the tail, or on a queue depth, or on shed count — not on how busy the box looks.
- Extrapolating from a healthy point is worthless near the cliff. 41 % CPU at 200 tables does not mean 480 tables is possible. The relationship is not linear and it does not degrade gracefully; it holds and then it falls over.
A load test is only halfway done: move the knee
If the work stops at “cap new connections at 48,” we have protection, not a performance improvement. A load test is not supposed to label a machine “good for 150 players.” It should locate the transition from stable latency to collapse, then give engineering a target to move. The real goal is to support more players on the same machine, under the same workload and latency budget, without weakening game correctness or hole-card secrecy.
There is code left to optimize, but guessing is not the method. Every connection currently has a bounded 64-frame outbound queue (server/src/ws.rs:758). The session task uses non-blocking try_send; when the queue is full it evicts that connection so the client can reconnect and recover from an authoritative Snapshot (server/src/session.rs:13008-13034). Changing 64 to 128 or 256 would merely hold more old messages. It would not make the writer drain faster or enlarge a 2 Mbps uplink. It may postpone the alarm, but it does not remove the bottleneck.
The hot path is not wholly unoptimized either. Viewer-invariant events such as ActionApplied and PotUpdated are already translated and serialized once, then distributed to every connection (server/src/session.rs:13116-13141). The eight-seat microbenchmark fell from 1.382 µs to 182 ns, about 7.6× faster. That means “reduce JSON serialization” is no longer a useful generic prescription. The next pass has to explain why messages stop draining after they leave the session task.
- Instrument the missing path. Record frame count and bytes by
ServerMsgvariant, each connection’s queue high-water mark,ws_tx.sendtime, and actual host egress bytes per second. The writer currently projects and writes one WebSocket frame at a time (server/src/ws.rs:1087-1177). Without these measurements we cannot distinguish too many frames, oversized frames, projection scheduling, and public-egress throttling. - Run two causal A/Bs. First, run the same binary at 300 tables over a high-bandwidth or private-network path. If the collapse disappears, optimize public egress and payload bytes first. Second, use an experimental branch that carries the adjacent
ActionAppliedandPotUpdatedupdates from one action in fewer transport frames. If backpressure falls, frame count and scheduling are the useful code targets. Both are proposed experiments, not shipped results. That was true when this section was added on 24 August 2026. Later that day the same code was run locally over loopback, on two Tokio workers, at 1,000 solo tables with no collapse — 1.2 ms action p99, zero early closes — while emitting about 4.07 Mbit/s of application JSON, more than the 2 Mbps production link can drain; and a compact wire profile that cuts outbound bytes 22.3% and frames 18.2% under that same local load shipped to production as web-2026.08.24.7 (ADR-106; see Why WebSocket delivery backed up). Neither local run is the production A/B above, and the production ladder has not yet been rerun. - Optimize only what the data names. If bytes bind first, identify the largest message variants. Keep Snapshot as the authoritative reconnect state, and inspect only ordinary transitions where a full Snapshot can safely become a narrower delta. If frame count or scheduling binds first, batch viewer-invariant events; never bypass the per-viewer projection that protects private hole cards.
- Repeat the same ladder. Run the same 150 / 200 / 300 tables, RTT, warm-up, and 150 ms budget. An optimization counts only if the healthy bound rises, p99 falls,
fanout_backpressure_evictfalls, and game-correctness plus card-secrecy regressions remain green.
That is the complete loop: load testing finds the knee, instrumentation narrows the symptom to a path, code changes target that path, and the original experiment proves whether the knee moved. A faster microbenchmark, lower CPU, or a larger queue is not a result by itself. More player load inside the latency budget is.
How many players did production admit on the test date?
The machine ceiling and the production operating point are different numbers. The defaults deployed on 21 August 2026 were 48 new WebSockets, with 12 additional slots reserved for reconnects or joins to existing tables. Normal arrivals that day therefore never reached the measured edge at 150, much less made players discover the 300-player collapse. There were separate caps of 12 connections per IP and four per account. These values are a historical snapshot of this load test and the ADR-105 rollout; changing a production constant later does not rewrite the experiment.
Why not take 80% of the solo result and admit 120? Six-max is a different shape: fewer session tasks, but every action fans out to five tablemates. The first two public-origin rungs measured:
| Six-max load | Action p99 | Server p99 | CPU | App RSS | Verdict |
|---|---|---|---|---|---|
| 60 players / 10 tables | 203.0 ms | 157 ms 69 ms in another run | 30.2% | 94 MB | Two runs straddled the budget |
| 120 players / 20 tables | 384.1 ms | 339 ms | 27.2% | 91 MB | Clearly over budget |
The 60-player rung actually ran twice: the host-sampled run in the table had a 157 ms server p99, while another run had 69 ms. It is not a stable result that says “60 always lags.” The current 48-new-socket operating point conservatively uses the worse 157 ms run and takes 80% of 60: production protection, not a claim that the hardware can serve only 48 people. The other default, ADMISSION_MAX_LIVE_TABLES=2000, is not a player-capacity claim either. It counts abandoned tables during their 10–20-minute reap window and exists as a memory fuse against a bug or script creating tables forever.
The numbers, one last time
- Machine: Tencent Cloud SA2.MEDIUM4; 2 vCPU, 3.6 GiB; 3 GiB app cap; Postgres and Caddy co-located; 2 Mbps fixed public egress.
- Solo measurement: 150 passes; 200 misses the budget; 300 collapses to a 4.1-second p99. The proven knee is a 150–200 interval, not an invented exact integer.
- Six-max measurement: two 60-player / 10-table runs had server p99s of 69 ms and 157 ms, straddling the budget; 120 / 20 reached 339 ms and are clearly over it.
- Production point on 2026-08-21: 48 new connections plus a 12-connection returning reserve. That was the day’s protection line, not the hardware ceiling.
- Still unmeasured: actual public egress / utilization during the collapse window, and capacity under a real geographic, RTT and workload mix. Missing data stays labelled missing.
These numbers are not a closing report. They are the start of the optimization work: machine, load shape, budget, healthy rung, over-budget rung, collapse rung and enforced operating point, plus the observations and causal tests still missing. The load test has done its job only when a code change moves the knee under the same ladder.