LiteInfer Benchmark Dashboard

4 run(s) in history  ·  5 engine(s) compared  ·  green = best per column red = worst per column  ·  x ratio is vs first (baseline) engine

throughput

All requests submitted at once — engine queues them, processes B=1 at a time. E2E includes queue wait.

📅 2026-05-09T14:57:20static-batching-b4🤖 Llama-3.2-1B-Instruct32 prompts · all submitted at once
EngineBreq/stok/sE2E p50E2E p99
liteinfer11.81base75.02base11192.5 msbase17661.7 msbase
liteinfer-kvcache11.871.03x74.801.00x9865.2 ms1.13x17099.8 ms1.03x
liteinfer-b444.422.44x181.652.42x4095.7 ms2.73x7233.5 ms2.44x
vllm12.891.60x180.452.41x5785.3 ms1.93x11023.2 ms1.60x
vllm-b4410.675.89x649.588.66x1656.0 ms6.76x2963.3 ms5.96x

latency

Sequential, no queue — each request sent only after previous finishes. Pure per-request engine latency.

📅 2026-05-09T15:01:00static-batching-b4🤖 Llama-3.2-1B-Instruct20 prompts · sequential, no queue
EngineBTTFT p50TTFT p99E2E p50tok/s
liteinfer113.4 msbase13.7 msbase1653.9 msbase74.83base
liteinfer-kvcache114.6 ms0.92x16.3 ms0.84x1489.6 ms1.11x75.131.00x
vllm125.5 ms0.53x27.2 ms0.51x692.8 ms2.39x182.552.44x