Run window (UTC) 2026-07-31T18:00:00Z to 2026-07-31T19:00:00Z · artifacts aaaaaaaaaaaaaaaa
WAIVED
At least one gate was explicitly WAIVED: it was not enforced for this run, and somebody recorded why. This is NOT a pass. The requirement was not withdrawn, it was set aside, and the reason is printed on the gate row so the risk that was accepted is legible rather than absent.
The verdict is the worst gate below (PASS < WAIVED < WARN < UNKNOWN < FAIL), decided when the run executed, not when this report was exported.
Per test
Test
Peak concurrency
Duration
Established
Peak CPU
Peak memory
Failed
By cause
ramp-500
492
60.0 min
99.900%
68.0%
61.2%
15
timeout 3, unexpected_sip 9
Peak CPU and memory are the worst node’s peak over each test’s window, correlated by time overlap. They describe the fleet during the test, not any particular call.
Capacity
Sized from 150 calls per node: highest concurrency sustained at or under 70% CPU
Target concurrency
Nodes for load
Spare
Total nodes
500
4
1
5
Machine type c7i.2xlarge for SIP, quoted from sizing-runbook.md:115. Nothing here derives a machine type.
Gates
Status
Gate
What it found
Metric
Value
Threshold
WAIVED
establishment ramp-500
fewer nodes were funded
not recorded
not measured
not measured
A gate reading UNKNOWN did not evaluate. It is not a pass with a caveat: nothing was compared against anything, and the value it would have reported is unknown.
Reproducible test assets
None recorded. Without the commands and flag semantics behind them, the numbers above cannot be reproduced by anybody else, which is most of what makes them evidence.
What this report does not measure
Read this section before acting on anything above. A load run measures a fleet from outside, through a generator placing real calls, and these are its structural limits, not this run’s bad luck.
RTP-port headroom IS measured, from the media port range's own occupancy against its configured size. The range size is a DECLARED configuration value rather than a measurement, so a node publishing occupancy without it reports unmeasured rather than being divided by a guess.
Network headroom is measured against a DECLARED baseline, not a measured capacity. The ENA driver reports no link speed, so there is no capacity to measure; the denominator is the instance type's published baseline, supplied by an operator at import. The instance can burst above it, so that figure is a floor rather than a ceiling and the ratio over-reads rather than under-reads. Inbound and outbound are judged separately because the hypervisor maintains separate credit buckets.
Packets-per-second headroom is NOT measured and never will be. No per-instance-type PPS allowance is published by anyone, including AWS, so there is no denominator to divide by. PPS saturation is still detected as an EVENT by the allowance gate; what cannot be quantified is how much headroom remained.
Per-call packet loss is NOT measured. No server-side surface attributes loss to one participant's leg, so no loss figure appears per call anywhere in this report.
Node samples are correlated to a test by TIME WINDOW OVERLAP. They say what the fleet was doing while a test ran, never that a node served any particular call: there is no per-call identity at that layer.
Peak CPU and memory are the WORST node's peak over the window, not a fleet average. A tier that looks healthy on average can contain one node that breached.
The calls-per-node figure is only meaningful if the generator actually reached each step. A ramp holding its arrival rate fixed while raising the target concurrency plateaus, and that plateau is the generator's ceiling rather than the node's.
Every call here was placed from ONE vantage point: the host that ran the generator. The establishment rate is what that host achieved against this deployment over that path, not what a caller on another network would see.