pip install quantcost · MIT

Quantization makes it 4× smaller. Whether it's faster depends on your hardware.

On an Apple M4, int8 generates tokens at 0.47× the fp32 rate while being 4× smaller — and it lands on 0.47 for both a 135M and a 500M model, which makes it look like a property of the runtime rather than one odd benchmark. Smaller: yes. Less memory: yes. Faster: no. One command measures speed, size, peak memory and perplexity on your own CPU, where the answer may well differ. No PyTorch, no GPU.

quantcost — on an Apple M4
$ pip install quantcost
$ quantcost run

# fp32 / int8 / int4, measured on this machine
  fp32  515 MB  83.7 tok/s  baseline  ppl 23.1
  int8  131 MB  39.4 tok/s  0.47x     ppl 25.0
  q4    174 MB  57.4 tok/s  0.69x     ppl 28.4

→ 4× smaller, 26% less RAM, 2.1× slower.
→ Quantize to fit in RAM, not to go fast.
Measured, not claimed

What quantization actually bought

SmolLM2-135M-Instruct, pre-quantized ONNX from the Hub, on an Apple M4 (arm64, macOS): 4 performance cores, 32 new tokens, greedy, 2 warmup / 5 measured runs, each precision in its own process. Perplexity over 8 × 512-token windows of a pinned WikiText-2 slice. No instability warning on this run.

3.9×
Smaller on disk
131 MB int8 vs 515 MB fp32 ONNX
0.47×
Throughput — slower, not faster
39.4 vs 83.7 tok/s · the same 0.47× on Qwen2.5-0.5B
26%
Peak RAM saved
1002 MB vs 1360 MB resident high-water mark
+8.5%
Perplexity cost
23.07 → 25.02 — quality mostly held

This is one machine, and that is the point: on hardware with native int8 kernels the speed column flips. The leaderboard is where that spread becomes visible.

Leaderboard

What quantization costs, across real machines

Every row below was measured by someone on their own hardware and submitted as a result card. Speedup is relative to that machine's own fp32 baseline, so the numbers compare quantization choices rather than hardware budgets.

Loading results…

add your machine
$ pip install quantcost
$ quantcost run
$ quantcost submit     # opens the PR for you
Deep dive · Qwen2.5-0.5B

One model, taken apart

A different measurement from the leaderboard above, and worth keeping separate: this is Qwen2.5-0.5B-Instruct through the authoring path — locally exported and quantized ONNX, driven by Optimum and by the C++ harness, rather than the pre-quantized artifacts the leaderboard measures. Different model, different runtime, so the numbers are not comparable with the table above and are not meant to be.

These runs predate a methodology fix. Peak RAM was measured with every precision in one process, which reports the union of their footprints, and the file cache was not warmed before timing — which penalises whichever artifact is largest, here the 2.4 GB fp32 baseline. Treat the relative throughput here with suspicion and the leaderboard as the trustworthy source; the size and perplexity columns are unaffected.

Model size on disk
ONNX artifacts · MB · lower is better
FP32
2405
ONNX FP32 export · 2404.97 MB · perplexity 19.00
INT4
771
ONNX INT4 block-wise · 770.93 MB · perplexity 24.58
INT8
605
ONNX INT8 per-channel · 604.98 MB · perplexity 20.10

INT4 is bigger than INT8 here — weight-only quantization compresses MatMul but not the ~544 MB FP32 token-embedding Gather of a 152k-token vocab.

Throughput — Python
ONNX Runtime CPU · tokens / second · higher is better
FP32
19.7
ort-cpu fp32 · 19.71 tok/s · 3.25 ± 0.09 s · 2911 MB RAM
INT4
20.0
ort-cpu int4 · 19.97 tok/s · 3.21 ± 0.07 s · 3194 MB RAM
INT8
38.3
ort-cpu int8 · 38.27 tok/s · 1.67 ± 0.04 s · 2862 MB RAM
PT INT8
15.9
PyTorch dynamic per-tensor INT8 · 15.92 tok/s · perplexity 58.42

Same bit width, different scheme: PyTorch's per-tensor dynamic INT8 is both slower and accuracy-destroying on ARM.

Decode speed — C++17 harness
tokens / second, KV cache · higher is better
FP32
18.0
edgellm_infer fp32 · 18.0 tok/s decode · 257 ms prefill
INT4
22.9
edgellm_infer int4 · 22.9 tok/s decode · 219 ms prefill
INT8
69.9
edgellm_infer int8 · 69.9 tok/s decode · 66 ms prefill

The C++ harness beats the Python ORT wrapper (69.9 vs 38.3 tok/s) on the same INT8 model — lower per-step overhead, identical weights.

Full benchmark table

Model: Qwen/Qwen2.5-0.5B-Instruct · 64 new tokens, greedy · warmup 2 / measured 5 · perplexity on WikiText-2.

Backend Precision Device Size (MB) Latency (s) Throughput (tok/s) Peak RAM (MB) Perplexity
pytorch fp32 mps 942.32 2.96 ± 0.02 21.60 2495.7 19.00
ort-cpu fp32 cpu 2404.97 3.25 ± 0.09 19.71 2910.8 19.00
ort-cpu best int8 cpu 604.98 1.67 ± 0.04 38.27 2861.8 20.10
ort-cpu int4 cpu 770.93 3.21 ± 0.07 19.97 3194.2 24.58
pytorch int8 cpu n/a¹ 4.02 ± 0.08 15.92 5396.5 58.42

¹ PyTorch dynamic INT8 is an in-memory transform with no standalone on-disk artifact. The PyTorch FP32 row measures the hub's bf16 safetensors (942 MB); the ONNX FP32 export is true fp32 (2405 MB) — the INT8/INT4 rows are the meaningful size comparison. Correctness anchor: PyTorch and ORT FP32 give identical perplexity (19.00), confirming the export is numerically faithful before any quantization.

Pipeline · authoring path

From Hugging Face checkpoint to on-device tokens

This is the pipeline that produces quantized artifacts, behind the .[quantize] extra. Benchmarking does not need any of it — quantcost run consumes ONNX already published on the Hub, which is what keeps that install torch-free. Reach for these commands when you want to quantize a model yourself.

01

Load

Resolve model + device (CPU / MPS / CUDA) and generate text with a swappable loader.

edgellm generate
02

Export

Optimum-based ONNX FP32 export, verified numerically faithful against PyTorch.

edgellm export
03

Quantize

INT8 per-channel dynamic and INT4 block-wise weight-only, with honest backend availability.

edgellm quantize
04

Benchmark

Latency, tokens/sec, peak RAM, on-disk size and perplexity across every backend.

edgellm benchmark
05

Report

Regenerate the results table and bar chart straight from the measurements JSON.

edgellm report
Runtimes

Four ways to run the quantized model

🐍

Python + ONNX Runtime

PyTorch and ORT runners behind one interface, so every precision is measured the same way.

✓ measured
⚙️

C++17 harness

edgellm_infer does autoregressive greedy decoding with a KV cache through the ORT C++ API, self-configuring from config.json.

✓ 69.9 tok/s INT8
📱

Android (ORT Mobile)

Kotlin app with the same KV-cache greedy decode, showing on-device tok/s in the UI.

✓ scaffold builds in Android Studio
🔲

Snapdragon NPU

Two paths: Qualcomm AI Hub device-farm profiling, and the ORT QNN execution provider.

⏳ needs AI Hub token
Findings

What the numbers actually taught me

Including the results that did not go the textbook way — kept in, not hidden.

The headline

"Quantize it, it'll be faster" is half wrong

Smaller is reliable. Faster is not. On an Apple M4, running the pre-quantized ONNX everyone actually downloads, int8 generates tokens at 0.47× the fp32 rate — and at 0.47× again on Qwen2.5-0.5B. The same figure twice, across a 135M and a 500M model, looks like a property of the runtime rather than one odd benchmark.

The cause is not slow integer arithmetic. Give the same artifacts a 512-token prefill instead of one token at a time and the int8 penalty roughly halves — 0.43× → 0.79× on SmolLM2, 0.59× → 0.97× on Qwen. That is the signature of a fixed per-step cost: ONNX Runtime unpacks the weights back to float inside each matmul, and one token cannot amortise unpacking a whole weight matrix. fp32 decode is bandwidth-bound on reading weights; int8 decode is compute-bound on unpacking them, so it does strictly more work while being 4× smaller.

Format, not bit width

int8 and q4 trade places

q4 is the better choice for generation and much the worse for prompt processing — exactly inverted from int8:

prefill (512 tok)decode (1 tok/step)
int80.79× / 0.97×0.43× / 0.59×
q40.31× / 0.36×0.68× / 1.02×

SmolLM2 / Qwen. They run through different kernels — MatMulNBits for q4, dynamic-quantize + MatMulInteger for int8 — with opposite strengths. So the right format depends on whether your workload is prompt-heavy or generation-heavy, which is not a question a single "speedup" number can answer.

Methodology

Efficiency cores cost int8 two-thirds of its speed

This M4 has 4 performance and 6 efficiency cores. Pinning ONNX Runtime to all ten physical cores — the obvious default, and the one this tool originally shipped — measured int8 at 24.1 tok/s. Pinning the 4 performance cores measured 40.0. Same model, same bytes, 66% of the throughput thrown away, and run-to-run spread doubled.

An ORT parallel region ends on a barrier, so one thread scheduled onto an efficiency core gates the whole region, and the OS places threads differently each run. fp32 barely noticed — it is bandwidth-bound — so the bad default penalised precisely the quantized formats under test and made the headline look worse than reality. Every card now records which rule chose its thread count.

Methodology

A cold file cache cost 4× on the baseline

ONNX Runtime memory-maps weights and faults them in lazily. Measured straight after download, the fp32 baseline read 17 tok/s; warm, the same machine and build read 64. Every speedup is divided by that number, so a cold run does not mis-measure one row — it corrupts the whole comparison.

The benchmark now reads each artifact through once before timing, runs every precision in its own process so peak RAM is not the union of all of them, and says so out loud when run-to-run variance exceeds 15%.

Scheme > bit width

Not all INT8 is equal

ONNX Runtime's per-channel INT8 holds perplexity at 20.10. PyTorch's default per-tensor dynamic INT8 — same 8 bits — collapses to 58.42, because on ARM/qnnpack it also ignores reduce_range.

That row stays in the benchmark table as a demonstration of the failure mode rather than a result to bury.

Model-specific

INT4 didn't beat INT8 on size

771 MB vs 605 MB. Qwen2.5-0.5B has a ~152k-token vocab, so its tied embedding / lm_head weight dominates the file. Weight-only INT4 compresses the MatMuls but not the FP32 embedding Gather, and the tied head adds a 4-bit copy.

Perplexity degrades further too (24.58). On a smaller-vocab model, INT4 wins.

Runtime overhead

The same weights are faster in C++

Identical INT8 ONNX model: 38.3 tok/s through the Python wrapper, 69.9 tok/s through the C++ harness, with prefill down from 257 ms to 66 ms. The gap is per-step Python overhead, not the model.

Authoring-path measurement on Qwen2.5-0.5B, taken before the thread-pinning and file-cache fixes above. The direction is solid — a C++ decode loop avoids per-token interpreter overhead — but treat the magnitude as unverified until it is re-measured.

Honest caveat

The compiler is a strong baseline

A hand-written NEON SDOT INT8 GEMM hits 59.2 GOPS vs 3.63 for a scalar loop with vectorization disabled — 16.3×. But at -O3 the compiler auto-vectorizes that same naive loop to ~74 GOPS.

The real lesson: beating a modern compiler on simple INT8 dot products takes register and cache blocking, not just intrinsics.

Quickstart

Run it yourself

  • Python 3.10–3.14. The default install is ONNX Runtime, NumPy and a tokenizer — no PyTorch, no GPU, about ten seconds.
  • quantcost models -m <repo> lists which precisions a model actually publishes before anything downloads.
  • Quantizing your own artifacts needs the heavier .[quantize] extra — PyTorch, Optimum and the C++ harness. Benchmarking does not.
  • CI runs ruff and pytest across Python 3.10–3.13, and asserts the default install never pulls in torch.
zsh
# 1 · benchmark your own machine
$ pip install quantcost
$ quantcost run
$ quantcost submit     # opens the PR for you

# 2 · try another model
$ quantcost run -m onnx-community/Qwen2.5-0.5B-Instruct

# 3 · author your own quantized artifacts (heavy extra)
$ pip install -e ".[quantize]"
$ edgellm export && edgellm quantize
$ cmake -S cpp -B cpp/build -DCMAKE_BUILD_TYPE=Release && cmake --build cpp/build

Run it on your machine

pip install quantcost — about ten seconds, no PyTorch, no GPU. Then quantcost submit puts your hardware on the leaderboard. The quantization pipeline, C++ harness, SIMD kernel and Android app are all in the same MIT-licensed repository.