mcgyvr · serving evidence · 2026-08-19 → 2026-08-28

Which setup earns the slot

Every decode-throughput measurement taken in August across both rigs, normalised to one score and ranked. Each cell names the run it came from.

score=tok/s×total params (B)×n — aggregate throughput, 475-token replies, ignore_eos, temp 0, one fixed prompt
rows extracted2,308
rankable1,301
refusals recorded196
directories mined15
rig one

srv1

gpu
GTX 1660 SUPER · 6 GB · sm75
cpu
i5-9600K · 6c/6t · 4.6 GHz
ram
48 GB DDR4-3200 · 19.6 GB/s
trait
big RAM, small card → expert offload
rig two

srv2

gpu
RTX 3060 · 12 GB · sm86
cpu
i9-10900F · 10c/20t · turbo off
ram
16 GB DDR4-2667 · 20.3 GB/s
trait
big card, small RAM → fits on GPU
table 1 & 2 — n = 1

One request, one user

Nobody waiting behind you. The score collapses to tok/s × size, so it asks one question: how much model can the rig push at interactive speed. Both rigs answer with a mixture-of-experts model — few active params to compute, many total params to count. srv2's answer is North-Mini-Code-1.0, added to the evidence base on 2026-08-27: 30B total on 3B active, whole model on the card at IQ2_M, 90.3 tok/s. It beats the incumbent GPT-OSS-20B by 36%.

srv1 · GTX 1660 SUPER · 6 GB · 48 GB DDR4
#modelquantenginetotal Btok/sscore
1Qwen3-Coder-Next-80B-A3BMoEQ3_K_XLllamacpp80.019.01,520
2026-08-282026-08-28-setup-selection/rows.jsonlline:82np=8 ctx_slot=1024 c=8192 ncmoe=50
2Qwen3.6-35B-A3BMoEIQ3_XXSllamacpp35.029.31,026
2026-08-252026-08-25-moe-expert-offload/width-sweep/srv1-35B-IQ3XXS-ncmoe35.txtline:7np=8 ctx_slot=1024 c=8192 ncmoe=35
3Qwen3-Coder-30B-A3BMoEQ4_K_Mllamacpp30.525.9790
2026-08-262026-08-26-capability-boundaries/srv1-llamacpp.txtline:10np=8 ctx_slot=1024 c=8192 ncmoe=38
4North-Mini-Code-1.0MoEQ4_K_Mllamacpp30.023.7711
2026-08-282026-08-28-north-mini-code/results-srv1.txtline:3np=8 ctx_slot=1024 c=8192 ncmoe=40
5GPT-OSS-4BdenseQ4_K_Mllamacpp4.299.0416
2026-08-282026-08-28-setup-selection/rows.jsonlline:298np=32 ctx_slot=1024 c=32768 ncmoe=0
6Qwen2.5-Coder-7BdenseIQ4_XSllamacpp7.6154.5415
2026-08-252026-08-25-moe-expert-offload/width-sweep/srv1-7B-IQ4XS.txtline:3np=1 ctx_slot=1024 c=1024 ncmoe=0
7Nemotron-7BdenseQ4_K_Mllamacpp7.649.6377
2026-08-282026-08-28-setup-selection/rows.jsonlline:52np=16 ctx_slot=1024 c=16384 ncmoe=0
8DeepSeek-Coder-V2-Lite-16BMoEQ4_0llamacpp15.720.3319
2026-08-282026-08-28-setup-selection/rows.jsonlline:93np=8 ctx_slot=1024 c=8192 ncmoe=30
9Qwen3-4BdenseQ4_K_Mllamacpp4.0276.7308
2026-08-282026-08-28-setup-selection/rows.jsonlline:355np=16 ctx_slot=1024 c=16384 ncmoe=0
10Qwen2.5-Coder-3BdenseQ4_K_Mllamacpp3.0996.8299
2026-08-282026-08-28-setup-selection/rows.jsonlline:29np=32 ctx_slot=1024 c=32768 ncmoe=0
srv2 · RTX 3060 · 12 GB · 16 GB DDR4
#modelquantenginetotal Btok/sscore
1North-Mini-Code-1.0MoEIQ2_Mllamacpp30.090.32,709
2026-08-282026-08-28-north-mini-code/results-srv2.txtline:3np=8 ctx_slot=1024 c=8192 ncmoe=0
2GPT-OSS-20BMoEMXFP4llamacpp20.597.01,988
2026-08-282026-08-28-setup-selection/rows.jsonlline:330np=8 ctx_slot=1024 c=8192 ncmoe=0
3Qwen3.6-35B-A3BMoEIQ3_XXSllamacpp35.044.91,572
2026-08-252026-08-25-moe-expert-offload/width-sweep/srv2-35B-IQ3XXS-ncmoe25.txtline:15np=32 ctx_slot=1024 c=32768 ncmoe=25
4DeepSeek-Coder-V2-Lite-16BMoEQ4_0ollama15.794.41,482
2026-08-19calibration-2026-08-19/d7-survey.json$.hosts.srv2.measured['deepseek-coder-v2-16b'].concurrency.levels[0] (n=1)ollama default (declared_slots=1 from llama-server /props total_slots; ctx=4096) entry=dee
5Nemotron-3-Nano-30B-A3BMoEIQ2_XXSllamacpp30.044.31,329
2026-08-282026-08-28-setup-selection/rows.jsonlline:237np=8 ctx_slot=1024 c=8192 ncmoe=30
6Qwen3-Coder-30B-A3BMoEIQ3_XXSllamacpp30.537.81,153
2026-08-282026-08-28-setup-selection/rows.jsonlline:274np=32 ctx_slot=1024 c=32768 ncmoe=25
7Qwen3-Coder-30B-A3BMoEQ4_K_Mollama30.522.7692
2026-08-19calibration-2026-08-19/d7-survey.json$.hosts.srv2.measured['qwen3-coder-30b'].concurrency.levels[0] (n=1)ollama default (declared_slots=1 from llama-server /props total_slots; ctx=4096) entry=qwe
8GPT-OSS-20BMoEMXFP4ollama20.532.4664
2026-08-19calibration-2026-08-19/d7-survey.json$.hosts.srv2.measured['gpt-oss-20b'].concurrency.levels[0] (n=1)ollama default (declared_slots=1 from llama-server /props total_slots; ctx=4096) entry=gpt
9Qwen3-Coder-Next-80B-A3BMoEQ3_K_XLllamacpp80.07.5600
2026-08-282026-08-28-setup-selection/rows.jsonlline:249np=8 ctx_slot=1024 c=8192 ncmoe=60
10GPT-OSS-4BdenseQ4_K_Mllamacpp4.2130.5548
2026-08-282026-08-28-setup-selection/rows.jsonlline:267np=32 ctx_slot=1024 c=32768 ncmoe=0
table 3 & 4 — n = argmax

Many requests, many users

Each setup is taken at the concurrency that maximises its own score, over the measured ladder {1 … 384}. This is where the two rigs stop resembling each other: srv1's best fleet row is a 35B MoE crawling at 4.0 tok/s per stream, srv2's is a 1.5B on vLLM serving 384 streams at once.

srv1 · GTX 1660 SUPER · 6 GB · 48 GB DDR4
#modelquantenginetotal Btok/snscore ×nscore ÷n
1Qwen3.6-35B-A3BMoEIQ3_XXSllamacpp35.0128.14.0/stream32143,4724,484
2026-08-262026-08-26-capability-boundaries/srv1-llamacpp.txtline:7np=32 ctx_slot=1024 c=32768 ncmoe=40
2Qwen2.5-Coder-1.5BdenseAWQvllm1.54294.71.2/stream256116,183454
2026-08-242026-08-24-config-sweep/srv1-1.5B-stage2.jsonlline:4 $.levels[5]cell=s2-noeager-len512-seqs256; --gpu-memory-utilization 0.85 --max-model-len 512 --max-nu
3Qwen2.5-Coder-1.5BdenseQ4_K_Mllamacpp1.54467.53.7/stream12892,154720
2026-08-282026-08-28-setup-selection/rows.jsonlline:27np=128 ctx_slot=1024 c=131072 ncmoe=0
4Qwen2.5-Coder-3BdenseQ4_K_Mllamacpp3.09268.74.2/stream6453,138830
2026-08-282026-08-28-setup-selection/rows.jsonlline:42np=64 ctx_slot=1024 c=65536 ncmoe=0
5Qwen3-Coder-30B-A3BMoEQ4_K_Mllamacpp30.549.61.6/stream3248,4101,513
2026-08-262026-08-26-capability-boundaries/srv1-llamacpp.txtline:15np=32 ctx_slot=1024 c=32768 ncmoe=44
6Qwen3-Coder-Next-80B-A3BMoEQ3_K_XLllamacpp80.029.01.8/stream1637,1202,320
2026-08-282026-08-28-setup-selection/rows.jsonlline:91np=16 ctx_slot=1024 c=16384 ncmoe=60
7North-Mini-Code-1.0MoEQ4_K_Mllamacpp30.066.24.1/stream1631,7761,986
2026-08-282026-08-28-north-mini-code/results-srv1.txtline:12np=16 ctx_slot=1024 c=16384 ncmoe=44
8Qwen2.5-Coder-3BdenseAWQvllm3.09155.32.4/stream6430,712480
2026-08-282026-08-28-setup-selection/rows.jsonlline:125util=0.85 len=1024 seqs=64 kv=auto
9GPT-OSS-4BdenseQ4_K_Mllamacpp4.2215.36.7/stream3228,936904
2026-08-282026-08-28-setup-selection/rows.jsonlline:303np=32 ctx_slot=1024 c=32768 ncmoe=0
10Qwen3-4BdenseAWQvllm4.0287.91.4/stream6422,615353
2026-08-282026-08-28-setup-selection/rows.jsonlline:133util=0.85 len=1024 seqs=64 kv=auto
srv2 · RTX 3060 · 12 GB · 16 GB DDR4
#modelquantenginetotal Btok/snscore ×nscore ÷n
1Qwen2.5-Coder-1.5BdenseAWQvllm1.546,038.715.7/stream3843,571,0469,300
2026-08-242026-08-24-config-sweep/srv2-1.5B-ceiling.jsonlline:4 $.levels[6]cell=g-1.5B-seqs384; --gpu-memory-utilization 0.85 --max-model-len 1024 --max-num-seqs 384
2Qwen2.5-Coder-7BdenseAWQvllm7.611,674.36.5/stream2563,261,80412,741
2026-08-242026-08-24-config-sweep/srv2-7B.jsonlline:6 $.levels[5]cell=g-7B-noeager-kvfp8-len1024-seqs256; --gpu-memory-utilization 0.85 --max-model-len 102
3Qwen2.5-Coder-3BdenseAWQvllm3.093,497.313.7/stream2562,766,50410,807
2026-08-262026-08-26-capability-boundaries/srv2-vllm-3b.txtline:6util=0.85 len=1024 seqs=256 kv=fp8
4Qwen3-4BdenseAWQvllm4.022,325.09.1/stream2562,392,7049,346
2026-08-242026-08-24-config-sweep/srv2-q3-4B.jsonlline:6 $.levels[5]cell=g-q3-4B-noeager-kvfp8-len1024-seqs256; --gpu-memory-utilization 0.85 --max-model-len
5Qwen2.5-Coder-14BdenseAWQvllm14.7313.51.2/stream2561,179,7634,608
2026-08-262026-08-26-capability-boundaries/srv2-vllm-14b.txtline:7util=0.95 len=1024 seqs=256 kv=fp8
6Qwen2.5-Coder-7BdenseIQ4_XSllamacpp7.611,107.68.7/stream1281,078,8918,429
2026-08-282026-08-28-setup-selection/rows.jsonlline:207np=128 ctx_slot=1024 c=131072 ncmoe=0
7Qwen2.5-Coder-1.5BdenseQ4_K_Mllamacpp1.541,684.56.6/stream256664,0972,594
2026-08-282026-08-28-setup-selection/rows.jsonlline:157np=256 ctx_slot=1024 c=262144 ncmoe=0
8Qwen2.5-Coder-3BdenseQ4_K_Mllamacpp3.091,361.510.6/stream128538,5004,207
2026-08-282026-08-28-setup-selection/rows.jsonlline:174np=128 ctx_slot=1024 c=131072 ncmoe=0
9Qwen3.6-35B-A3BMoEIQ3_XXSllamacpp35.0258.98.1/stream32289,9689,062
2026-08-282026-08-28-setup-selection/rows.jsonlline:235np=32 ctx_slot=1024 c=32768 ncmoe=25
10Qwen3-4BdenseQ4_K_Mllamacpp4.021,010.815.8/stream64260,0594,063
2026-08-282026-08-28-setup-selection/rows.jsonlline:182np=64 ctx_slot=1024 c=65536 ncmoe=0
the formula

Why n is counted twice

The formula is right. The number it is fed is the wrong one.

what the record stores at concurrency n:
agg_tok_s = total tokens from all n streams ÷ wall = per_stream × n

so the key, evaluated on the stored number, becomes:
score = agg_tok_s × size × n = per_stream × size × n²

Written as tok/s × size × n with tok/s meaning one stream's rate, the key is exactly correct — it is model-mass delivered per second across the fleet. The record stores the aggregate instead, which already contains n. Feeding the aggregate in makes the score quadratic in width, so it does not rank setups: it ranks how many slots each one will accept.

Both readings collapse to the same simple thing. per_stream × size × n = agg_tok_s × size. That is the ÷n column added to tables 3 and 4 — the same cell, with the second n removed. No new measurement, just the arithmetic the key intended.

It matters. Under the quadratic score the argmax always drifts to the widest level that has not yet collapsed; under the corrected one it lands where aggregate throughput actually peaks. Below is the fleet top ten re-ranked, each setup taken at its own new argmax n.

srv1 · corrected fleet score — tok/s × total B
#modelquantenginetotal Btok/snscore ÷nwas
1Qwen3.6-35B-A3BMoEIQ3_XXSllamacpp35.0128.14.0/stream324,484#1 =
2026-08-262026-08-26-capability-boundaries/srv1-llamacpp.txtline:7np=32 ctx_slot=1024 c=32768 ncmoe=40
2Qwen3-Coder-Next-80B-A3BMoEQ3_K_XLllamacpp80.029.01.8/stream162,320#6 ↑4
2026-08-282026-08-28-setup-selection/rows.jsonlline:91np=16 ctx_slot=1024 c=16384 ncmoe=60
3Qwen3-Coder-30B-A3BMoEQ4_K_Mllamacpp30.567.88.5/stream82,068#5 ↑2
2026-08-282026-08-28-setup-selection/rows.jsonlline:63np=8 ctx_slot=1024 c=8192 ncmoe=38
4North-Mini-Code-1.0MoEQ4_K_Mllamacpp30.066.24.1/stream161,986#7 ↑3
2026-08-282026-08-28-north-mini-code/results-srv1.txtline:12np=16 ctx_slot=1024 c=16384 ncmoe=44
5Qwen2.5-Coder-7BdenseIQ4_XSllamacpp7.61128.716.1/stream8979#12 ↑7
2026-08-252026-08-25-moe-expert-offload/width-sweep/srv1-7B-IQ4XS.txtline:12np=16 ctx_slot=1024 c=16384 ncmoe=0
6GPT-OSS-4BdenseQ4_K_Mllamacpp4.2215.36.7/stream32904#9 ↑3
2026-08-282026-08-28-setup-selection/rows.jsonlline:303np=32 ctx_slot=1024 c=32768 ncmoe=0
7Nemotron-7BdenseQ4_K_Mllamacpp7.6115.67.2/stream16879#11 ↑4
2026-08-282026-08-28-setup-selection/rows.jsonlline:56np=16 ctx_slot=1024 c=16384 ncmoe=0
8Qwen2.5-Coder-3BdenseQ4_K_Mllamacpp3.09268.74.2/stream64830#4 ↓4
2026-08-282026-08-28-setup-selection/rows.jsonlline:42np=64 ctx_slot=1024 c=65536 ncmoe=0
9Qwen2.5-Coder-1.5BdenseQ4_K_Mllamacpp1.54467.53.7/stream128720#3 ↓6
2026-08-282026-08-28-setup-selection/rows.jsonlline:27np=128 ctx_slot=1024 c=131072 ncmoe=0
10Qwen3-4BdenseQ4_K_Mllamacpp4.02174.410.9/stream16701#13 ↑3
2026-08-282026-08-28-setup-selection/rows.jsonlline:359np=16 ctx_slot=1024 c=16384 ncmoe=0
srv2 · corrected fleet score — tok/s × total B
#modelquantenginetotal Btok/snscore ÷nwas
1North-Mini-Code-1.0MoEIQ2_Mllamacpp30.0518.432.4/stream1615,552#12 ↑11
2026-08-282026-08-28-north-mini-code/results-srv2.txtline:12np=16 ctx_slot=1024 c=16384 ncmoe=0
2Qwen2.5-Coder-7BdenseAWQvllm7.611,674.36.5/stream25612,741#2 =
2026-08-242026-08-24-config-sweep/srv2-7B.jsonlline:6 $.levels[5]cell=g-7B-noeager-kvfp8-len1024-seqs256; --gpu-memory-utilization 0.85 --max-model-len 102
3Qwen2.5-Coder-3BdenseAWQvllm3.093,497.313.7/stream25610,807#3 =
2026-08-262026-08-26-capability-boundaries/srv2-vllm-3b.txtline:6util=0.85 len=1024 seqs=256 kv=fp8
4Qwen2.5-Coder-1.5BdenseAWQvllm1.546,480.625.3/stream2569,980#1 ↓3
2026-08-242026-08-24-engine-sweep/srv2/rows.jsonlline:136 (cell=R2, n=256, run=1)image=vllm/vllm-openai:v0.26.0;--gpu-memory-utilization 0.85 --max-model-len 1024 --max-nu
5Qwen3-4BdenseAWQvllm4.022,378.618.6/stream1289,562#4 ↓1
2026-08-282026-08-28-setup-selection/rows.jsonlline:386util=0.9 len=1024 seqs=128 kv=fp8
6Qwen3.6-35B-A3BMoEIQ3_XXSllamacpp35.0258.98.1/stream329,062#9 ↑3
2026-08-282026-08-28-setup-selection/rows.jsonlline:235np=32 ctx_slot=1024 c=32768 ncmoe=25
7Qwen2.5-Coder-14BdenseAWQvllm14.7610.19.5/stream648,968#5 ↓2
2026-08-282026-08-28-setup-selection/rows.jsonlline:353util=0.9 len=1024 seqs=64 kv=fp8
8Qwen2.5-Coder-7BdenseIQ4_XSllamacpp7.611,107.68.7/stream1288,429#6 ↓2
2026-08-282026-08-28-setup-selection/rows.jsonlline:207np=128 ctx_slot=1024 c=131072 ncmoe=0
9Qwen3-Coder-30B-A3BMoEIQ3_XXSllamacpp30.5264.78.3/stream328,073#11 ↑2
2026-08-282026-08-28-setup-selection/rows.jsonlline:279np=32 ctx_slot=1024 c=32768 ncmoe=25
10Nemotron-7BdenseQ4_K_Mllamacpp7.6778.024.3/stream325,913#13 ↑3
2026-08-282026-08-28-setup-selection/rows.jsonlline:214np=32 ctx_slot=1024 c=32768 ncmoe=0

srv2's answer changes hands entirely. The 1.5B that wins on the quadratic score falls to fourth; North-Mini-Code-1.0 at n=16 takes it on 518.4 tok/s × 30B, and Qwen2.5-Coder-7B AWQ at n=256 — the setup the 08-28 record picked before the driver capped it at n=128 — comes second. A 30B MoE that never accepted more than 16 streams was invisible under the quadratic score and is the best fleet setup on the rig under the corrected one. On srv1 the mixture-of-experts models climb three to seven places each: they were being punished for not accepting 128 streams, which was never the question.

speculative decoding

What a draft model actually buys

Nine measured target/draft pairs, each against its own no-draft baseline from the same run and the same build. These are not on the 475-token protocol — reply lengths were 60, 150 and 256 tokens — so the ratio transfers to the ladder above, the absolute rates do not.

enginerigtargetdraftnno draftwith draft×note
llama.cppsrv2Qwen2.5-Coder-7B IQ4_XSqwen2.5-coder-1.5b172.780.91.11best of four NMAX settings (76.0 / 79.7 / 80.2 / 80.9)150-token replies
llama.cppsrv2Qwen2.5-Coder-7B IQ4_XSqwen2.5-coder-1.5b4212.6226.21.06gain survives batching, but shrinks150-token replies
llama.cppsrv1Qwen3-Coder-30B-A3BQwen3-1.7B122.823.51.03offload-bound at --n-cpu-moe 44; the draft cannot help what the PCIe bus gates60-token replies
vLLMsrv2Qwen2.5-Coder-7B AWQQwen2.5-Coder-1.5B AWQ168.2369.331.02the only vLLM pairing that is not a loss256-token replies
vLLMsrv2Qwen2.5-Coder-7B AWQQwen2.5-Coder-1.5B AWQ8 cuda-graphs475.37418.390.88256-token replies
vLLMsrv2Qwen2.5-Coder-7B AWQQwen2.5-Coder-1.5B AWQ8 FLASH_ATTN475.32411.510.87256-token replies
vLLMsrv2Qwen2.5-Coder-7B AWQQwen2.5-Coder-1.5B AWQ8 eager320.99240.230.75256-token replies
vLLMsrv1Qwen2.5-Coder-3B AWQQwen2.5-Coder-0.5B AWQ8 eager64.8743.370.67256-token replies
vLLMsrv1Qwen2.5-Coder-3B AWQQwen2.5-Coder-0.5B AWQ8 cuda-graphs64.8537.970.59worst case measured: a third of throughput given away256-token replies

The split is by engine, not by rig. llama.cpp gains: +11% on the srv2 7B at n=1, still +6% at n=4. vLLM loses under load: every batched pairing measured is negative, down to 0.59× on srv1 — the draft steals the compute the batch was using. At n=1 on srv2 vLLM it is a wash at 1.02×.

It changes no table above. Applied to the two ranked setups that have a measured pairing, srv1's Qwen3-Coder-30B-A3B goes 790 → 814 (holds rank 3) and srv2's 7B IQ4_XS goes 548 → 610 (rank 10 → 8). Nothing else moves, and no fleet row improves at all. Two SD configurations were never tested: 35B-A3B native MTP and external-draft on srv1 are narrative-only refusals with no measurement, and vLLM rejects the 1.5B→7B pair outright without use_heterogeneous_vocab.
time to first token

How long before it starts talking

TTFT was never recorded under the 475-token protocol — neither the llama.cpp nor the vLLM driver logs it, and it cannot be recovered from p50 and wall alone. It exists in exactly one place: 334 per-request ollama records in the 08-24 engine sweep, which report prefill, decode and total separately. TTFT here is derived as total − decode, so it includes queue wait — which is the honest number, because that is what the caller feels.

rigmodelnreqsTTFT p50prefillqueuesplit
srv1qwen2.5-coder:1.5b1342ms1128
srv1qwen2.5-coder:1.5b26101ms4233
srv1qwen2.5-coder:1.5b824430ms8810
srv1qwen2.5-coder:1.5b3264655ms51843
srv2qwen2.5-coder:1.5b1321ms710
srv2qwen2.5-coder:1.5b82466ms319
srv2qwen2.5-coder:1.5b3232100ms6326
srv2qwen2.5-coder:1.5b128128809ms284523
srv2qwen2.5-coder:7b1242ms1720
srv2qwen2.5-coder:7b816162ms4434
srv2qwen2.5-coder:7b3232488ms46419
prefill queue wait bar length ∝ TTFT

At n=1 srv2 answers in 20.8 ms against srv1's 42.1 ms — the same 2× that shows up everywhere else between these two rigs. The shape changes with width: through n=32 prefill dominates and TTFT tracks the model, but at n=128 on srv2 queue wait is 523 ms of the 809 ms. Past that point TTFT stops measuring the engine and starts measuring the backlog, which is precisely the cost the fleet tables charge nothing for.

These are ollama numbers on 1.5B and 7B only. They set the floor and the shape, not the value for a vLLM or llama.cpp setup in the tables above. Measuring TTFT on the ranked setups is an unrun experiment.
before you act on these

Six things the tables do not say

The headline score multiplies n twice.

agg_tok_s is already summed across all n streams — verified against n × 475 / wall to a 0.13% median error across 295 levels, with the per-stream reading rejected at 87% error. The ÷n column in tables 3 and 4, and the re-ranked tables above, carry the corrected reading. Both are shown because the quadratic version is the key as written and as the 08-28 record's own Key 2 applies it.

Anything inside 2.6% on srv1 or 5.2% on srv2 is noise.

Those are the across-reload spreads, not the 0.04%/0.2% within-service figures the raw files quote. srv1 also drifts about 2.6% downward over six minutes of load, so any A-then-B contrast is biased against B. Ranks 9 and 10 of a table are frequently the same setup twice.

Peak throughput is not free.

srv1's top fleet row holds 128.1 tok/s at n=32 with a 118.6-second p50 — each caller waits two minutes. srv2's 1.5B ceiling run needs 30.2 s of wall at n=384. The score has no latency term.

Four claims in the record were falsified on re-measurement.

The srv2 --no-mmap win is +2.1% cold, not +63%. The srv1 ncmoe=38 peak does not exist — 37 is 1.7% above it, inside the noise bar. srv1's memory bandwidth is 19.6 GB/s, not 26.8, which inverts the srv1/srv2 bandwidth ratio. Rows resting on those readings are flagged or dropped.

Total params is the smarts proxy, not active.

The owner's lean. On active params the podium inverts — GPT-OSS-20B's 97.0 tok/s scores 1,988 on 20.5B total but 349 on 3.6B active, behind GPT-OSS-4B's 548. Active params are carried in the dataset if you want to re-score.

Two big directories hold no throughput at all.

2026-08-23-phase0-footprint (2.6 MB survey) and phase0-refit set collect.concurrency=false by design — they measure VRAM residency and weight digests. 4.4 MB of evidence, zero cells.

evidence base

What was dropped, and why

2,308 measured rows came out of the August directories. 1,301 were eligible to rank. The gap is not noise — it is rows that answer a different question.

427not on the 475-token protocolollama and LMDeploy stop on EOS (replies ran 294–390 tokens); the whole 2026-08-25 expert-offload campaign is a 128-token run.
284restatement of a row already countedPost-swap summary tables and survey repeat arrays re-print earlier runs. levels[i] is the max of its repeats, not a fresh sample.
184secondary baseline minebaseline-2026-08-23..27.jsonl is labelled reference-only, and its implied reply length varies 302–475 tokens. Every directory it mines is covered by a primary read.
64co-resident, not soloMeasured with a second model still holding VRAM — not the number you get running it alone.
17offline harnessllama-batched-bench is not a server; it does not measure serving throughput.
12speculative decodingTarget+draft pairs are a different quantity. Their matched no-draft baselines are kept and do rank.
11contaminated or supersededsrv1 gpt-oss-4b np=32 was re-run as contaminated; its two attempts disagree 22% at n=16.
4retracted by the recordThe docker --memory=15g cells never bound: the GGUF sat in host page cache outside the cgroup.
4control / bridge runb10481 cells exist to bridge builds, not to be chosen.
walls

Empty cells, with cause

A refusal is a recorded result. 195 of them are in the dataset with the engine's own sentence; these are the ones that shape the tables above.

14B dense on llama.cppboth8.9 GB weights + KV exceeds the budget at every np that fits
32B densesrv1 / srv2OOM on the 6 GB card; srv2's 15 GB RAM is under the 19.9 GB file
gpt-oss-20b MXFP4srv1cudaMalloc OOM — 12.1 GB file, 6 GB card
gpt-oss-20b Q3_K_Mbothunknown model architecture: 'gptoss' — the unsloth tag; only the official MXFP4 conversion loads
7B AWQ on vLLMsrv1OOM at util 0.85–0.95 — cc 7.5, no FA2. 100% refusal across 8 attempts; there is no srv1 7B vLLM number anywhere
nemotron-4b fp8bothMinimum capability: 89. Current: 75/86
vLLM seqs=256 on srv2srv2Engine core initialization failed on driver 595.84 — but the 08-24 runs on driver 580 reached n=384
speculative decoding, 08-24bothrecorded as refused, but the cause was a shell-quoting bug that split the JSON — it was never actually tested