LiteInfer Benchmark Dashboard

2 run(s) in history  ·  3 engine(s) compared  ·  green = best per column red = worst per column  ·  x ratio is vs first (baseline) engine

throughput

All requests submitted at once — engine queues them, processes B=1 at a time. E2E includes queue wait.

📅 2026-05-07T23:09:05v0_baseline🤖 Llama-3.2-1B-Instructbatch_size=132 prompts · all submitted at once
Engine (B=1)req/stok/sE2E p50E2E p99
liteinfer1.69base70.06base11956 msbase18913 msbase
liteinfer-kvcache1.771.04x70.611.01x10480 ms1.14x18114 ms1.04x
vllm2.891.71x180.132.57x5786 ms2.07x11042 ms1.71x

latency

Sequential, no queue — each request sent only after previous finishes. Pure per-request engine latency.

📅 2026-05-07T23:10:48v0_baseline🤖 Llama-3.2-1B-Instructbatch_size=120 prompts · sequential, no queue
Engine (B=1)TTFT p50TTFT p99E2E p50tok/s
liteinfer14 msbase16 msbase1726 msbase71.67base
liteinfer-kvcache15 ms0.89x17 ms0.90x1541 ms1.12x72.091.01x
vllm26 ms0.53x31 ms0.50x694 ms2.49x182.122.54x