Every decode-throughput measurement taken in August across both rigs, normalised to one score and ranked. Each cell names the run it came from.
Nobody waiting behind you. The score collapses to tok/s × size, so it asks one question: how much model can the rig push at interactive speed. Both rigs answer with a mixture-of-experts model — few active params to compute, many total params to count. srv2's answer is North-Mini-Code-1.0, added to the evidence base on 2026-08-27: 30B total on 3B active, whole model on the card at IQ2_M, 90.3 tok/s. It beats the incumbent GPT-OSS-20B by 36%.
| # | model | quant | engine | total B | tok/s | score |
|---|---|---|---|---|---|---|
| 1 | Qwen3-Coder-Next-80B-A3BMoE | Q3_K_XL | llamacpp | 80.0 | 19.0 | 1,520 |
| 2026-08-282026-08-28-setup-selection/rows.jsonlline:82np=8 ctx_slot=1024 c=8192 ncmoe=50 | ||||||
| 2 | Qwen3.6-35B-A3BMoE | IQ3_XXS | llamacpp | 35.0 | 29.3 | 1,026 |
| 2026-08-252026-08-25-moe-expert-offload/width-sweep/srv1-35B-IQ3XXS-ncmoe35.txtline:7np=8 ctx_slot=1024 c=8192 ncmoe=35 | ||||||
| 3 | Qwen3-Coder-30B-A3BMoE | Q4_K_M | llamacpp | 30.5 | 25.9 | 790 |
| 2026-08-262026-08-26-capability-boundaries/srv1-llamacpp.txtline:10np=8 ctx_slot=1024 c=8192 ncmoe=38 | ||||||
| 4 | North-Mini-Code-1.0MoE | Q4_K_M | llamacpp | 30.0 | 23.7 | 711 |
| 2026-08-282026-08-28-north-mini-code/results-srv1.txtline:3np=8 ctx_slot=1024 c=8192 ncmoe=40 | ||||||
| 5 | GPT-OSS-4Bdense | Q4_K_M | llamacpp | 4.2 | 99.0 | 416 |
| 2026-08-282026-08-28-setup-selection/rows.jsonlline:298np=32 ctx_slot=1024 c=32768 ncmoe=0 | ||||||
| 6 | Qwen2.5-Coder-7Bdense | IQ4_XS | llamacpp | 7.61 | 54.5 | 415 |
| 2026-08-252026-08-25-moe-expert-offload/width-sweep/srv1-7B-IQ4XS.txtline:3np=1 ctx_slot=1024 c=1024 ncmoe=0 | ||||||
| 7 | Nemotron-7Bdense | Q4_K_M | llamacpp | 7.6 | 49.6 | 377 |
| 2026-08-282026-08-28-setup-selection/rows.jsonlline:52np=16 ctx_slot=1024 c=16384 ncmoe=0 | ||||||
| 8 | DeepSeek-Coder-V2-Lite-16BMoE | Q4_0 | llamacpp | 15.7 | 20.3 | 319 |
| 2026-08-282026-08-28-setup-selection/rows.jsonlline:93np=8 ctx_slot=1024 c=8192 ncmoe=30 | ||||||
| 9 | Qwen3-4Bdense | Q4_K_M | llamacpp | 4.02 | 76.7 | 308 |
| 2026-08-282026-08-28-setup-selection/rows.jsonlline:355np=16 ctx_slot=1024 c=16384 ncmoe=0 | ||||||
| 10 | Qwen2.5-Coder-3Bdense | Q4_K_M | llamacpp | 3.09 | 96.8 | 299 |
| 2026-08-282026-08-28-setup-selection/rows.jsonlline:29np=32 ctx_slot=1024 c=32768 ncmoe=0 | ||||||
| # | model | quant | engine | total B | tok/s | score |
|---|---|---|---|---|---|---|
| 1 | North-Mini-Code-1.0MoE | IQ2_M | llamacpp | 30.0 | 90.3 | 2,709 |
| 2026-08-282026-08-28-north-mini-code/results-srv2.txtline:3np=8 ctx_slot=1024 c=8192 ncmoe=0 | ||||||
| 2 | GPT-OSS-20BMoE | MXFP4 | llamacpp | 20.5 | 97.0 | 1,988 |
| 2026-08-282026-08-28-setup-selection/rows.jsonlline:330np=8 ctx_slot=1024 c=8192 ncmoe=0 | ||||||
| 3 | Qwen3.6-35B-A3BMoE | IQ3_XXS | llamacpp | 35.0 | 44.9 | 1,572 |
| 2026-08-252026-08-25-moe-expert-offload/width-sweep/srv2-35B-IQ3XXS-ncmoe25.txtline:15np=32 ctx_slot=1024 c=32768 ncmoe=25 | ||||||
| 4 | DeepSeek-Coder-V2-Lite-16BMoE | Q4_0 | ollama | 15.7 | 94.4 | 1,482 |
| 2026-08-19calibration-2026-08-19/d7-survey.json$.hosts.srv2.measured['deepseek-coder-v2-16b'].concurrency.levels[0] (n=1)ollama default (declared_slots=1 from llama-server /props total_slots; ctx=4096) entry=dee | ||||||
| 5 | Nemotron-3-Nano-30B-A3BMoE | IQ2_XXS | llamacpp | 30.0 | 44.3 | 1,329 |
| 2026-08-282026-08-28-setup-selection/rows.jsonlline:237np=8 ctx_slot=1024 c=8192 ncmoe=30 | ||||||
| 6 | Qwen3-Coder-30B-A3BMoE | IQ3_XXS | llamacpp | 30.5 | 37.8 | 1,153 |
| 2026-08-282026-08-28-setup-selection/rows.jsonlline:274np=32 ctx_slot=1024 c=32768 ncmoe=25 | ||||||
| 7 | Qwen3-Coder-30B-A3BMoE | Q4_K_M | ollama | 30.5 | 22.7 | 692 |
| 2026-08-19calibration-2026-08-19/d7-survey.json$.hosts.srv2.measured['qwen3-coder-30b'].concurrency.levels[0] (n=1)ollama default (declared_slots=1 from llama-server /props total_slots; ctx=4096) entry=qwe | ||||||
| 8 | GPT-OSS-20BMoE | MXFP4 | ollama | 20.5 | 32.4 | 664 |
| 2026-08-19calibration-2026-08-19/d7-survey.json$.hosts.srv2.measured['gpt-oss-20b'].concurrency.levels[0] (n=1)ollama default (declared_slots=1 from llama-server /props total_slots; ctx=4096) entry=gpt | ||||||
| 9 | Qwen3-Coder-Next-80B-A3BMoE | Q3_K_XL | llamacpp | 80.0 | 7.5 | 600 |
| 2026-08-282026-08-28-setup-selection/rows.jsonlline:249np=8 ctx_slot=1024 c=8192 ncmoe=60 | ||||||
| 10 | GPT-OSS-4Bdense | Q4_K_M | llamacpp | 4.2 | 130.5 | 548 |
| 2026-08-282026-08-28-setup-selection/rows.jsonlline:267np=32 ctx_slot=1024 c=32768 ncmoe=0 | ||||||
Each setup is taken at the concurrency that maximises its own score, over the measured ladder {1 … 384}. This is where the two rigs stop resembling each other: srv1's best fleet row is a 35B MoE crawling at 4.0 tok/s per stream, srv2's is a 1.5B on vLLM serving 384 streams at once.
| # | model | quant | engine | total B | tok/s | n | score ×n | score ÷n |
|---|---|---|---|---|---|---|---|---|
| 1 | Qwen3.6-35B-A3BMoE | IQ3_XXS | llamacpp | 35.0 | 128.14.0/stream | 32 | 143,472 | 4,484 |
| 2026-08-262026-08-26-capability-boundaries/srv1-llamacpp.txtline:7np=32 ctx_slot=1024 c=32768 ncmoe=40 | ||||||||
| 2 | Qwen2.5-Coder-1.5Bdense | AWQ | vllm | 1.54 | 294.71.2/stream | 256 | 116,183 | 454 |
| 2026-08-242026-08-24-config-sweep/srv1-1.5B-stage2.jsonlline:4 $.levels[5]cell=s2-noeager-len512-seqs256; --gpu-memory-utilization 0.85 --max-model-len 512 --max-nu | ||||||||
| 3 | Qwen2.5-Coder-1.5Bdense | Q4_K_M | llamacpp | 1.54 | 467.53.7/stream | 128 | 92,154 | 720 |
| 2026-08-282026-08-28-setup-selection/rows.jsonlline:27np=128 ctx_slot=1024 c=131072 ncmoe=0 | ||||||||
| 4 | Qwen2.5-Coder-3Bdense | Q4_K_M | llamacpp | 3.09 | 268.74.2/stream | 64 | 53,138 | 830 |
| 2026-08-282026-08-28-setup-selection/rows.jsonlline:42np=64 ctx_slot=1024 c=65536 ncmoe=0 | ||||||||
| 5 | Qwen3-Coder-30B-A3BMoE | Q4_K_M | llamacpp | 30.5 | 49.61.6/stream | 32 | 48,410 | 1,513 |
| 2026-08-262026-08-26-capability-boundaries/srv1-llamacpp.txtline:15np=32 ctx_slot=1024 c=32768 ncmoe=44 | ||||||||
| 6 | Qwen3-Coder-Next-80B-A3BMoE | Q3_K_XL | llamacpp | 80.0 | 29.01.8/stream | 16 | 37,120 | 2,320 |
| 2026-08-282026-08-28-setup-selection/rows.jsonlline:91np=16 ctx_slot=1024 c=16384 ncmoe=60 | ||||||||
| 7 | North-Mini-Code-1.0MoE | Q4_K_M | llamacpp | 30.0 | 66.24.1/stream | 16 | 31,776 | 1,986 |
| 2026-08-282026-08-28-north-mini-code/results-srv1.txtline:12np=16 ctx_slot=1024 c=16384 ncmoe=44 | ||||||||
| 8 | Qwen2.5-Coder-3Bdense | AWQ | vllm | 3.09 | 155.32.4/stream | 64 | 30,712 | 480 |
| 2026-08-282026-08-28-setup-selection/rows.jsonlline:125util=0.85 len=1024 seqs=64 kv=auto | ||||||||
| 9 | GPT-OSS-4Bdense | Q4_K_M | llamacpp | 4.2 | 215.36.7/stream | 32 | 28,936 | 904 |
| 2026-08-282026-08-28-setup-selection/rows.jsonlline:303np=32 ctx_slot=1024 c=32768 ncmoe=0 | ||||||||
| 10 | Qwen3-4Bdense | AWQ | vllm | 4.02 | 87.91.4/stream | 64 | 22,615 | 353 |
| 2026-08-282026-08-28-setup-selection/rows.jsonlline:133util=0.85 len=1024 seqs=64 kv=auto | ||||||||
| # | model | quant | engine | total B | tok/s | n | score ×n | score ÷n |
|---|---|---|---|---|---|---|---|---|
| 1 | Qwen2.5-Coder-1.5Bdense | AWQ | vllm | 1.54 | 6,038.715.7/stream | 384 | 3,571,046 | 9,300 |
| 2026-08-242026-08-24-config-sweep/srv2-1.5B-ceiling.jsonlline:4 $.levels[6]cell=g-1.5B-seqs384; --gpu-memory-utilization 0.85 --max-model-len 1024 --max-num-seqs 384 | ||||||||
| 2 | Qwen2.5-Coder-7Bdense | AWQ | vllm | 7.61 | 1,674.36.5/stream | 256 | 3,261,804 | 12,741 |
| 2026-08-242026-08-24-config-sweep/srv2-7B.jsonlline:6 $.levels[5]cell=g-7B-noeager-kvfp8-len1024-seqs256; --gpu-memory-utilization 0.85 --max-model-len 102 | ||||||||
| 3 | Qwen2.5-Coder-3Bdense | AWQ | vllm | 3.09 | 3,497.313.7/stream | 256 | 2,766,504 | 10,807 |
| 2026-08-262026-08-26-capability-boundaries/srv2-vllm-3b.txtline:6util=0.85 len=1024 seqs=256 kv=fp8 | ||||||||
| 4 | Qwen3-4Bdense | AWQ | vllm | 4.02 | 2,325.09.1/stream | 256 | 2,392,704 | 9,346 |
| 2026-08-242026-08-24-config-sweep/srv2-q3-4B.jsonlline:6 $.levels[5]cell=g-q3-4B-noeager-kvfp8-len1024-seqs256; --gpu-memory-utilization 0.85 --max-model-len | ||||||||
| 5 | Qwen2.5-Coder-14Bdense | AWQ | vllm | 14.7 | 313.51.2/stream | 256 | 1,179,763 | 4,608 |
| 2026-08-262026-08-26-capability-boundaries/srv2-vllm-14b.txtline:7util=0.95 len=1024 seqs=256 kv=fp8 | ||||||||
| 6 | Qwen2.5-Coder-7Bdense | IQ4_XS | llamacpp | 7.61 | 1,107.68.7/stream | 128 | 1,078,891 | 8,429 |
| 2026-08-282026-08-28-setup-selection/rows.jsonlline:207np=128 ctx_slot=1024 c=131072 ncmoe=0 | ||||||||
| 7 | Qwen2.5-Coder-1.5Bdense | Q4_K_M | llamacpp | 1.54 | 1,684.56.6/stream | 256 | 664,097 | 2,594 |
| 2026-08-282026-08-28-setup-selection/rows.jsonlline:157np=256 ctx_slot=1024 c=262144 ncmoe=0 | ||||||||
| 8 | Qwen2.5-Coder-3Bdense | Q4_K_M | llamacpp | 3.09 | 1,361.510.6/stream | 128 | 538,500 | 4,207 |
| 2026-08-282026-08-28-setup-selection/rows.jsonlline:174np=128 ctx_slot=1024 c=131072 ncmoe=0 | ||||||||
| 9 | Qwen3.6-35B-A3BMoE | IQ3_XXS | llamacpp | 35.0 | 258.98.1/stream | 32 | 289,968 | 9,062 |
| 2026-08-282026-08-28-setup-selection/rows.jsonlline:235np=32 ctx_slot=1024 c=32768 ncmoe=25 | ||||||||
| 10 | Qwen3-4Bdense | Q4_K_M | llamacpp | 4.02 | 1,010.815.8/stream | 64 | 260,059 | 4,063 |
| 2026-08-282026-08-28-setup-selection/rows.jsonlline:182np=64 ctx_slot=1024 c=65536 ncmoe=0 | ||||||||
The formula is right. The number it is fed is the wrong one.
Written as tok/s × size × n with tok/s meaning one stream's rate,
the key is exactly correct — it is model-mass delivered per second across the fleet. The record
stores the aggregate instead, which already contains n. Feeding the aggregate in makes the score
quadratic in width, so it does not rank setups: it ranks how many slots each one will accept.
per_stream × size × n = agg_tok_s × size. That is the
÷n column added to tables 3 and 4 — the same cell, with the second n removed.
No new measurement, just the arithmetic the key intended.It matters. Under the quadratic score the argmax always drifts to the widest level that has not yet collapsed; under the corrected one it lands where aggregate throughput actually peaks. Below is the fleet top ten re-ranked, each setup taken at its own new argmax n.
| # | model | quant | engine | total B | tok/s | n | score ÷n | was |
|---|---|---|---|---|---|---|---|---|
| 1 | Qwen3.6-35B-A3BMoE | IQ3_XXS | llamacpp | 35.0 | 128.14.0/stream | 32 | 4,484 | #1 = |
| 2026-08-262026-08-26-capability-boundaries/srv1-llamacpp.txtline:7np=32 ctx_slot=1024 c=32768 ncmoe=40 | ||||||||
| 2 | Qwen3-Coder-Next-80B-A3BMoE | Q3_K_XL | llamacpp | 80.0 | 29.01.8/stream | 16 | 2,320 | #6 ↑4 |
| 2026-08-282026-08-28-setup-selection/rows.jsonlline:91np=16 ctx_slot=1024 c=16384 ncmoe=60 | ||||||||
| 3 | Qwen3-Coder-30B-A3BMoE | Q4_K_M | llamacpp | 30.5 | 67.88.5/stream | 8 | 2,068 | #5 ↑2 |
| 2026-08-282026-08-28-setup-selection/rows.jsonlline:63np=8 ctx_slot=1024 c=8192 ncmoe=38 | ||||||||
| 4 | North-Mini-Code-1.0MoE | Q4_K_M | llamacpp | 30.0 | 66.24.1/stream | 16 | 1,986 | #7 ↑3 |
| 2026-08-282026-08-28-north-mini-code/results-srv1.txtline:12np=16 ctx_slot=1024 c=16384 ncmoe=44 | ||||||||
| 5 | Qwen2.5-Coder-7Bdense | IQ4_XS | llamacpp | 7.61 | 128.716.1/stream | 8 | 979 | #12 ↑7 |
| 2026-08-252026-08-25-moe-expert-offload/width-sweep/srv1-7B-IQ4XS.txtline:12np=16 ctx_slot=1024 c=16384 ncmoe=0 | ||||||||
| 6 | GPT-OSS-4Bdense | Q4_K_M | llamacpp | 4.2 | 215.36.7/stream | 32 | 904 | #9 ↑3 |
| 2026-08-282026-08-28-setup-selection/rows.jsonlline:303np=32 ctx_slot=1024 c=32768 ncmoe=0 | ||||||||
| 7 | Nemotron-7Bdense | Q4_K_M | llamacpp | 7.6 | 115.67.2/stream | 16 | 879 | #11 ↑4 |
| 2026-08-282026-08-28-setup-selection/rows.jsonlline:56np=16 ctx_slot=1024 c=16384 ncmoe=0 | ||||||||
| 8 | Qwen2.5-Coder-3Bdense | Q4_K_M | llamacpp | 3.09 | 268.74.2/stream | 64 | 830 | #4 ↓4 |
| 2026-08-282026-08-28-setup-selection/rows.jsonlline:42np=64 ctx_slot=1024 c=65536 ncmoe=0 | ||||||||
| 9 | Qwen2.5-Coder-1.5Bdense | Q4_K_M | llamacpp | 1.54 | 467.53.7/stream | 128 | 720 | #3 ↓6 |
| 2026-08-282026-08-28-setup-selection/rows.jsonlline:27np=128 ctx_slot=1024 c=131072 ncmoe=0 | ||||||||
| 10 | Qwen3-4Bdense | Q4_K_M | llamacpp | 4.02 | 174.410.9/stream | 16 | 701 | #13 ↑3 |
| 2026-08-282026-08-28-setup-selection/rows.jsonlline:359np=16 ctx_slot=1024 c=16384 ncmoe=0 | ||||||||
| # | model | quant | engine | total B | tok/s | n | score ÷n | was |
|---|---|---|---|---|---|---|---|---|
| 1 | North-Mini-Code-1.0MoE | IQ2_M | llamacpp | 30.0 | 518.432.4/stream | 16 | 15,552 | #12 ↑11 |
| 2026-08-282026-08-28-north-mini-code/results-srv2.txtline:12np=16 ctx_slot=1024 c=16384 ncmoe=0 | ||||||||
| 2 | Qwen2.5-Coder-7Bdense | AWQ | vllm | 7.61 | 1,674.36.5/stream | 256 | 12,741 | #2 = |
| 2026-08-242026-08-24-config-sweep/srv2-7B.jsonlline:6 $.levels[5]cell=g-7B-noeager-kvfp8-len1024-seqs256; --gpu-memory-utilization 0.85 --max-model-len 102 | ||||||||
| 3 | Qwen2.5-Coder-3Bdense | AWQ | vllm | 3.09 | 3,497.313.7/stream | 256 | 10,807 | #3 = |
| 2026-08-262026-08-26-capability-boundaries/srv2-vllm-3b.txtline:6util=0.85 len=1024 seqs=256 kv=fp8 | ||||||||
| 4 | Qwen2.5-Coder-1.5Bdense | AWQ | vllm | 1.54 | 6,480.625.3/stream | 256 | 9,980 | #1 ↓3 |
| 2026-08-242026-08-24-engine-sweep/srv2/rows.jsonlline:136 (cell=R2, n=256, run=1)image=vllm/vllm-openai:v0.26.0;--gpu-memory-utilization 0.85 --max-model-len 1024 --max-nu | ||||||||
| 5 | Qwen3-4Bdense | AWQ | vllm | 4.02 | 2,378.618.6/stream | 128 | 9,562 | #4 ↓1 |
| 2026-08-282026-08-28-setup-selection/rows.jsonlline:386util=0.9 len=1024 seqs=128 kv=fp8 | ||||||||
| 6 | Qwen3.6-35B-A3BMoE | IQ3_XXS | llamacpp | 35.0 | 258.98.1/stream | 32 | 9,062 | #9 ↑3 |
| 2026-08-282026-08-28-setup-selection/rows.jsonlline:235np=32 ctx_slot=1024 c=32768 ncmoe=25 | ||||||||
| 7 | Qwen2.5-Coder-14Bdense | AWQ | vllm | 14.7 | 610.19.5/stream | 64 | 8,968 | #5 ↓2 |
| 2026-08-282026-08-28-setup-selection/rows.jsonlline:353util=0.9 len=1024 seqs=64 kv=fp8 | ||||||||
| 8 | Qwen2.5-Coder-7Bdense | IQ4_XS | llamacpp | 7.61 | 1,107.68.7/stream | 128 | 8,429 | #6 ↓2 |
| 2026-08-282026-08-28-setup-selection/rows.jsonlline:207np=128 ctx_slot=1024 c=131072 ncmoe=0 | ||||||||
| 9 | Qwen3-Coder-30B-A3BMoE | IQ3_XXS | llamacpp | 30.5 | 264.78.3/stream | 32 | 8,073 | #11 ↑2 |
| 2026-08-282026-08-28-setup-selection/rows.jsonlline:279np=32 ctx_slot=1024 c=32768 ncmoe=25 | ||||||||
| 10 | Nemotron-7Bdense | Q4_K_M | llamacpp | 7.6 | 778.024.3/stream | 32 | 5,913 | #13 ↑3 |
| 2026-08-282026-08-28-setup-selection/rows.jsonlline:214np=32 ctx_slot=1024 c=32768 ncmoe=0 | ||||||||
srv2's answer changes hands entirely. The 1.5B that wins on the quadratic score falls to fourth; North-Mini-Code-1.0 at n=16 takes it on 518.4 tok/s × 30B, and Qwen2.5-Coder-7B AWQ at n=256 — the setup the 08-28 record picked before the driver capped it at n=128 — comes second. A 30B MoE that never accepted more than 16 streams was invisible under the quadratic score and is the best fleet setup on the rig under the corrected one. On srv1 the mixture-of-experts models climb three to seven places each: they were being punished for not accepting 128 streams, which was never the question.
Nine measured target/draft pairs, each against its own no-draft baseline from the same run and the same build. These are not on the 475-token protocol — reply lengths were 60, 150 and 256 tokens — so the ratio transfers to the ladder above, the absolute rates do not.
| engine | rig | target | draft | n | no draft | with draft | × | note |
|---|---|---|---|---|---|---|---|---|
| llama.cpp | srv2 | Qwen2.5-Coder-7B IQ4_XS | qwen2.5-coder-1.5b | 1 | 72.7 | 80.9 | 1.11 | best of four NMAX settings (76.0 / 79.7 / 80.2 / 80.9)150-token replies |
| llama.cpp | srv2 | Qwen2.5-Coder-7B IQ4_XS | qwen2.5-coder-1.5b | 4 | 212.6 | 226.2 | 1.06 | gain survives batching, but shrinks150-token replies |
| llama.cpp | srv1 | Qwen3-Coder-30B-A3B | Qwen3-1.7B | 1 | 22.8 | 23.5 | 1.03 | offload-bound at --n-cpu-moe 44; the draft cannot help what the PCIe bus gates60-token replies |
| vLLM | srv2 | Qwen2.5-Coder-7B AWQ | Qwen2.5-Coder-1.5B AWQ | 1 | 68.23 | 69.33 | 1.02 | the only vLLM pairing that is not a loss256-token replies |
| vLLM | srv2 | Qwen2.5-Coder-7B AWQ | Qwen2.5-Coder-1.5B AWQ | 8 cuda-graphs | 475.37 | 418.39 | 0.88 | 256-token replies |
| vLLM | srv2 | Qwen2.5-Coder-7B AWQ | Qwen2.5-Coder-1.5B AWQ | 8 FLASH_ATTN | 475.32 | 411.51 | 0.87 | 256-token replies |
| vLLM | srv2 | Qwen2.5-Coder-7B AWQ | Qwen2.5-Coder-1.5B AWQ | 8 eager | 320.99 | 240.23 | 0.75 | 256-token replies |
| vLLM | srv1 | Qwen2.5-Coder-3B AWQ | Qwen2.5-Coder-0.5B AWQ | 8 eager | 64.87 | 43.37 | 0.67 | 256-token replies |
| vLLM | srv1 | Qwen2.5-Coder-3B AWQ | Qwen2.5-Coder-0.5B AWQ | 8 cuda-graphs | 64.85 | 37.97 | 0.59 | worst case measured: a third of throughput given away256-token replies |
The split is by engine, not by rig. llama.cpp gains: +11% on the srv2 7B at n=1, still +6% at n=4. vLLM loses under load: every batched pairing measured is negative, down to 0.59× on srv1 — the draft steals the compute the batch was using. At n=1 on srv2 vLLM it is a wash at 1.02×.
use_heterogeneous_vocab.TTFT was never recorded under the 475-token protocol — neither the llama.cpp nor the vLLM driver logs it, and it cannot be recovered from p50 and wall alone. It exists in exactly one place: 334 per-request ollama records in the 08-24 engine sweep, which report prefill, decode and total separately. TTFT here is derived as total − decode, so it includes queue wait — which is the honest number, because that is what the caller feels.
| rig | model | n | reqs | TTFT p50 | prefill | queue | split |
|---|---|---|---|---|---|---|---|
| srv1 | qwen2.5-coder:1.5b | 1 | 3 | 42ms | 11 | 28 | |
| srv1 | qwen2.5-coder:1.5b | 2 | 6 | 101ms | 42 | 33 | |
| srv1 | qwen2.5-coder:1.5b | 8 | 24 | 430ms | 88 | 10 | |
| srv1 | qwen2.5-coder:1.5b | 32 | 64 | 655ms | 518 | 43 | |
| srv2 | qwen2.5-coder:1.5b | 1 | 3 | 21ms | 7 | 10 | |
| srv2 | qwen2.5-coder:1.5b | 8 | 24 | 66ms | 31 | 9 | |
| srv2 | qwen2.5-coder:1.5b | 32 | 32 | 100ms | 63 | 26 | |
| srv2 | qwen2.5-coder:1.5b | 128 | 128 | 809ms | 284 | 523 | |
| srv2 | qwen2.5-coder:7b | 1 | 2 | 42ms | 17 | 20 | |
| srv2 | qwen2.5-coder:7b | 8 | 16 | 162ms | 44 | 34 | |
| srv2 | qwen2.5-coder:7b | 32 | 32 | 488ms | 464 | 19 |
At n=1 srv2 answers in 20.8 ms against srv1's 42.1 ms — the same 2× that shows up everywhere else between these two rigs. The shape changes with width: through n=32 prefill dominates and TTFT tracks the model, but at n=128 on srv2 queue wait is 523 ms of the 809 ms. Past that point TTFT stops measuring the engine and starts measuring the backlog, which is precisely the cost the fleet tables charge nothing for.
agg_tok_s is already summed across all n streams — verified against n × 475 / wall to a 0.13% median error across 295 levels, with the per-stream reading rejected at 87% error. The ÷n column in tables 3 and 4, and the re-ranked tables above, carry the corrected reading. Both are shown because the quadratic version is the key as written and as the 08-28 record's own Key 2 applies it.
Those are the across-reload spreads, not the 0.04%/0.2% within-service figures the raw files quote. srv1 also drifts about 2.6% downward over six minutes of load, so any A-then-B contrast is biased against B. Ranks 9 and 10 of a table are frequently the same setup twice.
srv1's top fleet row holds 128.1 tok/s at n=32 with a 118.6-second p50 — each caller waits two minutes. srv2's 1.5B ceiling run needs 30.2 s of wall at n=384. The score has no latency term.
The srv2 --no-mmap win is +2.1% cold, not +63%. The srv1 ncmoe=38 peak does not exist — 37 is 1.7% above it, inside the noise bar. srv1's memory bandwidth is 19.6 GB/s, not 26.8, which inverts the srv1/srv2 bandwidth ratio. Rows resting on those readings are flagged or dropped.
The owner's lean. On active params the podium inverts — GPT-OSS-20B's 97.0 tok/s scores 1,988 on 20.5B total but 349 on 3.6B active, behind GPT-OSS-4B's 548. Active params are carried in the dataset if you want to re-score.
2026-08-23-phase0-footprint (2.6 MB survey) and phase0-refit set collect.concurrency=false by design — they measure VRAM residency and weight digests. 4.4 MB of evidence, zero cells.
2,308 measured rows came out of the August directories. 1,301 were eligible to rank. The gap is not noise — it is rows that answer a different question.
| 427 | not on the 475-token protocol | ollama and LMDeploy stop on EOS (replies ran 294–390 tokens); the whole 2026-08-25 expert-offload campaign is a 128-token run. |
| 284 | restatement of a row already counted | Post-swap summary tables and survey repeat arrays re-print earlier runs. levels[i] is the max of its repeats, not a fresh sample. |
| 184 | secondary baseline mine | baseline-2026-08-23..27.jsonl is labelled reference-only, and its implied reply length varies 302–475 tokens. Every directory it mines is covered by a primary read. |
| 64 | co-resident, not solo | Measured with a second model still holding VRAM — not the number you get running it alone. |
| 17 | offline harness | llama-batched-bench is not a server; it does not measure serving throughput. |
| 12 | speculative decoding | Target+draft pairs are a different quantity. Their matched no-draft baselines are kept and do rank. |
| 11 | contaminated or superseded | srv1 gpt-oss-4b np=32 was re-run as contaminated; its two attempts disagree 22% at n=16. |
| 4 | retracted by the record | The docker --memory=15g cells never bound: the GGUF sat in host page cache outside the cgroup. |
| 4 | control / bridge run | b10481 cells exist to bridge builds, not to be chosen. |
A refusal is a recorded result. 195 of them are in the dataset with the engine's own sentence; these are the ones that shape the tables above.
| 14B dense on llama.cpp | both | 8.9 GB weights + KV exceeds the budget at every np that fits |
| 32B dense | srv1 / srv2 | OOM on the 6 GB card; srv2's 15 GB RAM is under the 19.9 GB file |
| gpt-oss-20b MXFP4 | srv1 | cudaMalloc OOM — 12.1 GB file, 6 GB card |
| gpt-oss-20b Q3_K_M | both | unknown model architecture: 'gptoss' — the unsloth tag; only the official MXFP4 conversion loads |
| 7B AWQ on vLLM | srv1 | OOM at util 0.85–0.95 — cc 7.5, no FA2. 100% refusal across 8 attempts; there is no srv1 7B vLLM number anywhere |
| nemotron-4b fp8 | both | Minimum capability: 89. Current: 75/86 |
| vLLM seqs=256 on srv2 | srv2 | Engine core initialization failed on driver 595.84 — but the 08-24 runs on driver 580 reached n=384 |
| speculative decoding, 08-24 | both | recorded as refused, but the cause was a shell-quoting bug that split the JSON — it was never actually tested |