[MOUNT@a7b0bcb1-arbi/bf16] === arbi bf16 | /models/Qwen3.5-0.8B | ISL=16384 OSL=1024 | lock=2700MHz bits=n/a(TKV_BITS) | TP=1 gmu=0.99 mbt=2048 | mode=mount code=a7b0bcb1 img=registry.arbi.work/arbi-serve:latest@sha256:cda88574d7be1cccaebbe79a566407da745aa0ed4d0daf4f5dcfd9405ce99a98 | max_running_req=? ===
[MOUNT@a7b0bcb1-arbi/bf16] [c=1] Request throughput (req/s):              0.33      
[MOUNT@a7b0bcb1-arbi/bf16] [c=1] Output token throughput (tok/s):         333.20    
[MOUNT@a7b0bcb1-arbi/bf16] [c=1] Peak output token throughput (tok/s):    372.00    
[MOUNT@a7b0bcb1-arbi/bf16] [c=1] Median TTFT (ms):                        298.68    
[MOUNT@a7b0bcb1-arbi/bf16] [c=1] P95 TTFT (ms):                           368.18    
[MOUNT@a7b0bcb1-arbi/bf16] [c=1] Mean TPOT (ms):                          2.70      
[MOUNT@a7b0bcb1-arbi/bf16] [c=1] Median TPOT (ms):                        2.70      
[MOUNT@a7b0bcb1-arbi/bf16] [c=1] P95 TPOT (ms):                           2.70      
[rig] GPU 0: lock HELD @ 2700MHz across 7 working samples (peak host load1 14.11, peak mem 22818 MiB, no throttle).
[rig] host=dev-gpu load1=14.11/32c | gpu0: sm=2700MHz util=100% mem=1MiB throttle=0x0000000000000000 pcie=gen4x16
[MOUNT@a7b0bcb1-arbi/bf16] === arbi bf16 | /models/Qwen3.5-0.8B | ISL=16384 OSL=1024 | lock=2700MHz bits=n/a(TKV_BITS) | TP=1 gmu=0.99 mbt=2048 | mode=mount code=a7b0bcb1 img=registry.arbi.work/arbi-serve:latest@sha256:cda88574d7be1cccaebbe79a566407da745aa0ed4d0daf4f5dcfd9405ce99a98 | max_running_req=? ===
[MOUNT@a7b0bcb1-arbi/bf16] [c=16] Request throughput (req/s):              1.35      
[MOUNT@a7b0bcb1-arbi/bf16] [c=16] Output token throughput (tok/s):         1380.49   
[MOUNT@a7b0bcb1-arbi/bf16] [c=16] Peak output token throughput (tok/s):    2158.00   
[MOUNT@a7b0bcb1-arbi/bf16] [c=16] Median TTFT (ms):                        670.03    
[MOUNT@a7b0bcb1-arbi/bf16] [c=16] P95 TTFT (ms):                           3928.48   
[MOUNT@a7b0bcb1-arbi/bf16] [c=16] Mean TPOT (ms):                          10.33     
[MOUNT@a7b0bcb1-arbi/bf16] [c=16] Median TPOT (ms):                        10.78     
[MOUNT@a7b0bcb1-arbi/bf16] [c=16] P95 TPOT (ms):                           10.97     
[rig] GPU 0: lock HELD @ 2700MHz across 6 working samples (peak host load1 12.19, peak mem 23574 MiB, no throttle).
[rig] host=dev-gpu load1=9.71/32c | gpu0: sm=2700MHz util=100% mem=1MiB throttle=0x0000000000000000 pcie=gen4x16
[MOUNT@a7b0bcb1-arbi/bf16] === arbi bf16 | /models/Qwen3.5-0.8B | ISL=16384 OSL=1024 | lock=2700MHz bits=n/a(TKV_BITS) | TP=1 gmu=0.99 mbt=8192 | mode=mount code=a7b0bcb1 img=registry.arbi.work/arbi-serve:latest@sha256:cda88574d7be1cccaebbe79a566407da745aa0ed4d0daf4f5dcfd9405ce99a98 | max_running_req=? ===
[MOUNT@a7b0bcb1-arbi/bf16] [c=64] Request throughput (req/s):              1.73      
[MOUNT@a7b0bcb1-arbi/bf16] [c=64] Output token throughput (tok/s):         1772.10   
[MOUNT@a7b0bcb1-arbi/bf16] [c=64] Peak output token throughput (tok/s):    3007.00   
[MOUNT@a7b0bcb1-arbi/bf16] [c=64] Median TTFT (ms):                        939.88    
[MOUNT@a7b0bcb1-arbi/bf16] [c=64] P95 TTFT (ms):                           13677.27  
[MOUNT@a7b0bcb1-arbi/bf16] [c=64] Mean TPOT (ms):                          32.70     
[MOUNT@a7b0bcb1-arbi/bf16] [c=64] Median TPOT (ms):                        34.79     
[MOUNT@a7b0bcb1-arbi/bf16] [c=64] P95 TPOT (ms):                           35.33     
[rig] GPU 0: lock HELD @ 2700MHz across 24 working samples (peak host load1 9.71, peak mem 23140 MiB, SW power-cap in 1 samples — clock held).
[rig] host=dev-gpu load1=2.10/32c | gpu0: sm=2700MHz util=100% mem=1MiB throttle=0x0000000000000000 pcie=gen4x16
[MOUNT@a7b0bcb1-arbi/tkv] === arbi tkv | /models/Qwen3.5-0.8B | ISL=16384 OSL=1024 | lock=2700MHz bits=4(TKV_BITS) | TP=1 gmu=0.99 mbt=2048 | mode=mount code=a7b0bcb1 img=registry.arbi.work/arbi-serve:latest@sha256:cda88574d7be1cccaebbe79a566407da745aa0ed4d0daf4f5dcfd9405ce99a98 | max_running_req=? ===
[MOUNT@a7b0bcb1-arbi/tkv] [c=1] Request throughput (req/s):              0.33      
[MOUNT@a7b0bcb1-arbi/tkv] [c=1] Output token throughput (tok/s):         338.31    
[MOUNT@a7b0bcb1-arbi/tkv] [c=1] Peak output token throughput (tok/s):    382.00    
[MOUNT@a7b0bcb1-arbi/tkv] [c=1] Median TTFT (ms):                        318.37    
[MOUNT@a7b0bcb1-arbi/tkv] [c=1] P95 TTFT (ms):                           381.03    
[MOUNT@a7b0bcb1-arbi/tkv] [c=1] Mean TPOT (ms):                          2.64      
[MOUNT@a7b0bcb1-arbi/tkv] [c=1] Median TPOT (ms):                        2.65      
[MOUNT@a7b0bcb1-arbi/tkv] [c=1] P95 TPOT (ms):                           2.65      
[rig] GPU 0: lock HELD @ 2700MHz across 5 working samples (peak host load1 2.31, peak mem 22566 MiB, no throttle).
[rig] host=dev-gpu load1=1.93/32c | gpu0: sm=2700MHz util=100% mem=1MiB throttle=0x0000000000000000 pcie=gen4x16
[MOUNT@a7b0bcb1-arbi/tkv] === arbi tkv | /models/Qwen3.5-0.8B | ISL=16384 OSL=1024 | lock=2700MHz bits=4(TKV_BITS) | TP=1 gmu=0.99 mbt=2048 | mode=mount code=a7b0bcb1 img=registry.arbi.work/arbi-serve:latest@sha256:cda88574d7be1cccaebbe79a566407da745aa0ed4d0daf4f5dcfd9405ce99a98 | max_running_req=? ===
[MOUNT@a7b0bcb1-arbi/tkv] [c=16] Request throughput (req/s):              1.55      
[MOUNT@a7b0bcb1-arbi/tkv] [c=16] Output token throughput (tok/s):         1583.84   
[MOUNT@a7b0bcb1-arbi/tkv] [c=16] Peak output token throughput (tok/s):    2920.00   
[MOUNT@a7b0bcb1-arbi/tkv] [c=16] Median TTFT (ms):                        750.98    
[MOUNT@a7b0bcb1-arbi/tkv] [c=16] P95 TTFT (ms):                           4349.61   
[MOUNT@a7b0bcb1-arbi/tkv] [c=16] Mean TPOT (ms):                          8.70      
[MOUNT@a7b0bcb1-arbi/tkv] [c=16] Median TPOT (ms):                        9.19      
[MOUNT@a7b0bcb1-arbi/tkv] [c=16] P95 TPOT (ms):                           9.45      
[rig] GPU 0: lock HELD @ 2700MHz across 7 working samples (peak host load1 1.98, peak mem 23356 MiB, no throttle).
[rig] host=dev-gpu load1=1.90/32c | gpu0: sm=2700MHz util=0% mem=1MiB throttle=0x0000000000000000 pcie=gen4x16
[MOUNT@a7b0bcb1-arbi/tkv] === arbi tkv | /models/Qwen3.5-0.8B | ISL=16384 OSL=1024 | lock=2700MHz bits=4(TKV_BITS) | TP=1 gmu=0.99 mbt=8192 | mode=mount code=a7b0bcb1 img=registry.arbi.work/arbi-serve:latest@sha256:cda88574d7be1cccaebbe79a566407da745aa0ed4d0daf4f5dcfd9405ce99a98 | max_running_req=? ===
[MOUNT@a7b0bcb1-arbi/tkv] [c=64] Request throughput (req/s):              2.16      
[MOUNT@a7b0bcb1-arbi/tkv] [c=64] Output token throughput (tok/s):         2207.57   
[MOUNT@a7b0bcb1-arbi/tkv] [c=64] Peak output token throughput (tok/s):    4809.00   
[MOUNT@a7b0bcb1-arbi/tkv] [c=64] Median TTFT (ms):                        1156.28   
[MOUNT@a7b0bcb1-arbi/tkv] [c=64] P95 TTFT (ms):                           14120.29  
[MOUNT@a7b0bcb1-arbi/tkv] [c=64] Mean TPOT (ms):                          25.43     
[MOUNT@a7b0bcb1-arbi/tkv] [c=64] Median TPOT (ms):                        27.59     
[MOUNT@a7b0bcb1-arbi/tkv] [c=64] P95 TPOT (ms):                           28.18     
RIG-PREFLIGHT FATAL: GPU 0 left the 2700+/-60MHz band in 3/17 WORKING samples (min 2625MHz) — the clock MOVED mid-run; DISCARD.
[rig] host=dev-gpu load1=1.40/32c | gpu0: sm=2700MHz util=100% mem=1MiB throttle=0x0000000000000000 pcie=gen4x16
[MOUNT@fcf3fca7-vllm/bf16] === vllm bf16 | /models/Qwen3.5-0.8B | ISL=16384 OSL=1024 | lock=2700MHz bits=n/a(TKV_BITS) | TP=1 gmu=0.95 mbt=2048 | mode=mount code=fcf3fca7 img=registry.arbi.work/vllm-tkv:latest@sha256:895c53bbedc4f4ab9dafa3e1608adfe73b16b2602f842210defc3fe288d57121 | max_running_req=? ===
[MOUNT@fcf3fca7-vllm/bf16] [c=1] Request throughput (req/s):              0.34      
[MOUNT@fcf3fca7-vllm/bf16] [c=1] Output token throughput (tok/s):         343.85    
[MOUNT@fcf3fca7-vllm/bf16] [c=1] Peak output token throughput (tok/s):    376.00    
[MOUNT@fcf3fca7-vllm/bf16] [c=1] Median TTFT (ms):                        257.57    
[MOUNT@fcf3fca7-vllm/bf16] [c=1] P95 TTFT (ms):                           265.95    
[MOUNT@fcf3fca7-vllm/bf16] [c=1] Mean TPOT (ms):                          2.66      
[MOUNT@fcf3fca7-vllm/bf16] [c=1] Median TPOT (ms):                        2.66      
[MOUNT@fcf3fca7-vllm/bf16] [c=1] P95 TPOT (ms):                           2.66      
[rig] GPU 0: lock HELD @ 2700MHz across 10 working samples (peak host load1 1.42, peak mem 20692 MiB, no throttle).
[rig] host=dev-gpu load1=1.17/32c | gpu0: sm=2700MHz util=100% mem=1MiB throttle=0x0000000000000000 pcie=gen4x16
[MOUNT@fcf3fca7-vllm/bf16] === vllm bf16 | /models/Qwen3.5-0.8B | ISL=16384 OSL=1024 | lock=2700MHz bits=n/a(TKV_BITS) | TP=1 gmu=0.95 mbt=2048 | mode=mount code=fcf3fca7 img=registry.arbi.work/vllm-tkv:latest@sha256:895c53bbedc4f4ab9dafa3e1608adfe73b16b2602f842210defc3fe288d57121 | max_running_req=? ===
[MOUNT@fcf3fca7-vllm/bf16] [c=16] Request throughput (req/s):              1.46      
[MOUNT@fcf3fca7-vllm/bf16] [c=16] Output token throughput (tok/s):         1494.13   
[MOUNT@fcf3fca7-vllm/bf16] [c=16] Peak output token throughput (tok/s):    2304.00   
[MOUNT@fcf3fca7-vllm/bf16] [c=16] Median TTFT (ms):                        603.21    
[MOUNT@fcf3fca7-vllm/bf16] [c=16] P95 TTFT (ms):                           3525.92   
[MOUNT@fcf3fca7-vllm/bf16] [c=16] Mean TPOT (ms):                          9.55      
[MOUNT@fcf3fca7-vllm/bf16] [c=16] Median TPOT (ms):                        10.05     
[MOUNT@fcf3fca7-vllm/bf16] [c=16] P95 TPOT (ms):                           10.11     
[rig] GPU 0: lock HELD @ 2700MHz across 7 working samples (peak host load1 1.79, peak mem 20692 MiB, SW power-cap in 1 samples — clock held).
[rig] host=dev-gpu load1=1.25/32c | gpu0: sm=2700MHz util=90% mem=1MiB throttle=0x0000000000000000 pcie=gen4x16
[MOUNT@fcf3fca7-vllm/bf16] === vllm bf16 | /models/Qwen3.5-0.8B | ISL=16384 OSL=1024 | lock=2700MHz bits=n/a(TKV_BITS) | TP=1 gmu=0.95 mbt=8192 | mode=mount code=fcf3fca7 img=registry.arbi.work/vllm-tkv:latest@sha256:895c53bbedc4f4ab9dafa3e1608adfe73b16b2602f842210defc3fe288d57121 | max_running_req=? ===
[MOUNT@fcf3fca7-vllm/bf16] [c=64] Request throughput (req/s):              1.69      
[MOUNT@fcf3fca7-vllm/bf16] [c=64] Output token throughput (tok/s):         1733.61   
[MOUNT@fcf3fca7-vllm/bf16] [c=64] Peak output token throughput (tok/s):    3300.00   
[MOUNT@fcf3fca7-vllm/bf16] [c=64] Median TTFT (ms):                        2336.07   
[MOUNT@fcf3fca7-vllm/bf16] [c=64] P95 TTFT (ms):                           19767.52  
[MOUNT@fcf3fca7-vllm/bf16] [c=64] Mean TPOT (ms):                          29.90     
[MOUNT@fcf3fca7-vllm/bf16] [c=64] Median TPOT (ms):                        32.68     
[MOUNT@fcf3fca7-vllm/bf16] [c=64] P95 TPOT (ms):                           33.01     
RIG-PREFLIGHT FATAL: GPU 0 left the 2700+/-60MHz band in 3/25 WORKING samples (min 2610MHz) — the clock MOVED mid-run; DISCARD.
[rig] host=dev-gpu load1=0.61/32c | gpu0: sm=2700MHz util=39% mem=1MiB throttle=0x0000000000000000 pcie=gen4x16
[MOUNT@fcf3fca7-vllm/tkv] === vllm tkv | /models/Qwen3.5-0.8B | ISL=16384 OSL=1024 | lock=2700MHz bits=4(TKV_BITS) | TP=1 gmu=0.95 mbt=2048 | mode=mount code=fcf3fca7 img=registry.arbi.work/vllm-tkv:latest@sha256:895c53bbedc4f4ab9dafa3e1608adfe73b16b2602f842210defc3fe288d57121 | max_running_req=? ===
[MOUNT@fcf3fca7-vllm/tkv] [c=1] Request throughput (req/s):              0.35      
[MOUNT@fcf3fca7-vllm/tkv] [c=1] Output token throughput (tok/s):         353.51    
[MOUNT@fcf3fca7-vllm/tkv] [c=1] Peak output token throughput (tok/s):    389.00    
[MOUNT@fcf3fca7-vllm/tkv] [c=1] Median TTFT (ms):                        267.32    
[MOUNT@fcf3fca7-vllm/tkv] [c=1] P95 TTFT (ms):                           276.92    
[MOUNT@fcf3fca7-vllm/tkv] [c=1] Mean TPOT (ms):                          2.57      
[MOUNT@fcf3fca7-vllm/tkv] [c=1] Median TPOT (ms):                        2.57      
[MOUNT@fcf3fca7-vllm/tkv] [c=1] P95 TPOT (ms):                           2.57      
RIG-PREFLIGHT FATAL: GPU 0 left the 2700+/-60MHz band in 1/8 WORKING samples (min 2565MHz) — the clock MOVED mid-run; DISCARD.
[rig] host=dev-gpu load1=1.30/32c | gpu0: sm=2715MHz util=0% mem=1MiB throttle=0x0000000000000000 pcie=gen4x16
[MOUNT@fcf3fca7-vllm/tkv] === vllm tkv | /models/Qwen3.5-0.8B | ISL=16384 OSL=1024 | lock=2700MHz bits=4(TKV_BITS) | TP=1 gmu=0.95 mbt=2048 | mode=mount code=fcf3fca7 img=registry.arbi.work/vllm-tkv:latest@sha256:895c53bbedc4f4ab9dafa3e1608adfe73b16b2602f842210defc3fe288d57121 | max_running_req=? ===
[MOUNT@fcf3fca7-vllm/tkv] [c=16] Request throughput (req/s):              1.82      
[MOUNT@fcf3fca7-vllm/tkv] [c=16] Output token throughput (tok/s):         1863.25   
[MOUNT@fcf3fca7-vllm/tkv] [c=16] Peak output token throughput (tok/s):    3136.00   
[MOUNT@fcf3fca7-vllm/tkv] [c=16] Median TTFT (ms):                        633.89    
[MOUNT@fcf3fca7-vllm/tkv] [c=16] P95 TTFT (ms):                           3368.62   
[MOUNT@fcf3fca7-vllm/tkv] [c=16] Mean TPOT (ms):                          7.46      
[MOUNT@fcf3fca7-vllm/tkv] [c=16] Median TPOT (ms):                        7.86      
[MOUNT@fcf3fca7-vllm/tkv] [c=16] P95 TPOT (ms):                           8.04      
[rig] GPU 0: lock HELD @ 2700MHz across 8 working samples (peak host load1 3.71, peak mem 20786 MiB, no throttle).
[rig] host=dev-gpu load1=1.95/32c | gpu0: sm=2700MHz util=100% mem=1MiB throttle=0x0000000000000000 pcie=gen4x16
[MOUNT@fcf3fca7-vllm/tkv] === vllm tkv | /models/Qwen3.5-0.8B | ISL=16384 OSL=1024 | lock=2700MHz bits=4(TKV_BITS) | TP=1 gmu=0.95 mbt=8192 | mode=mount code=fcf3fca7 img=registry.arbi.work/vllm-tkv:latest@sha256:895c53bbedc4f4ab9dafa3e1608adfe73b16b2602f842210defc3fe288d57121 | max_running_req=? ===
[MOUNT@fcf3fca7-vllm/tkv] [c=64] Request throughput (req/s):              2.34      
[MOUNT@fcf3fca7-vllm/tkv] [c=64] Output token throughput (tok/s):         2395.69   
[MOUNT@fcf3fca7-vllm/tkv] [c=64] Peak output token throughput (tok/s):    5184.00   
[MOUNT@fcf3fca7-vllm/tkv] [c=64] Median TTFT (ms):                        1123.28   
[MOUNT@fcf3fca7-vllm/tkv] [c=64] P95 TTFT (ms):                           13063.76  
[MOUNT@fcf3fca7-vllm/tkv] [c=64] Mean TPOT (ms):                          23.35     
[MOUNT@fcf3fca7-vllm/tkv] [c=64] Median TPOT (ms):                        25.36     
[MOUNT@fcf3fca7-vllm/tkv] [c=64] P95 TPOT (ms):                           25.56     
RIG-PREFLIGHT FATAL: GPU 0 left the 2700+/-60MHz band in 7/20 WORKING samples (min 2595MHz) — the clock MOVED mid-run; DISCARD.
[rig] host=dev-gpu load1=1.52/32c | gpu0: sm=2700MHz util=0% mem=1MiB throttle=0x0000000000000000 pcie=gen4x16
RIG-PREFLIGHT FATAL: no sample caught a GPU actually working — the lock was never exercised; the run is unverifiable.
[rig] host=dev-gpu load1=1.72/32c | gpu0: sm=240MHz util=0% mem=1MiB throttle=0x0000000000000001 pcie=gen1x16
RIG-PREFLIGHT FATAL: no sample caught a GPU actually working — the lock was never exercised; the run is unverifiable.
[rig] host=dev-gpu load1=1.66/32c | gpu0: sm=210MHz util=0% mem=1MiB throttle=0x0000000000000001 pcie=gen1x16
RIG-PREFLIGHT FATAL: no sample caught a GPU actually working — the lock was never exercised; the run is unverifiable.
[rig] host=dev-gpu load1=1.77/32c | gpu0: sm=210MHz util=0% mem=1MiB throttle=0x0000000000000001 pcie=gen1x16
RIG-PREFLIGHT FATAL: no sample caught a GPU actually working — the lock was never exercised; the run is unverifiable.
[rig] host=dev-gpu load1=3.15/32c | gpu0: sm=210MHz util=0% mem=1MiB throttle=0x0000000000000001 pcie=gen1x16
RIG-PREFLIGHT FATAL: no sample caught a GPU actually working — the lock was never exercised; the run is unverifiable.
[rig] host=dev-gpu load1=4.02/32c | gpu0: sm=210MHz util=0% mem=1MiB throttle=0x0000000000000001 pcie=gen1x16
RIG-PREFLIGHT FATAL: no sample caught a GPU actually working — the lock was never exercised; the run is unverifiable.
[rig] host=dev-gpu load1=4.74/32c | gpu0: sm=210MHz util=0% mem=1MiB throttle=0x0000000000000001 pcie=gen1x16
