Quantization makes it 4× smaller. Whether it's faster depends on your hardware.
On an Apple M4, int8 generates tokens at 0.47× the fp32 rate while being 4× smaller — and it lands on 0.47 for both a 135M and a 500M model, which makes it look like a property of the runtime rather than one odd benchmark. Smaller: yes. Less memory: yes. Faster: no. One command measures speed, size, peak memory and perplexity on your own CPU, where the answer may well differ. No PyTorch, no GPU.
$ pip install quantcost $ quantcost run # fp32 / int8 / int4, measured on this machine fp32 515 MB 83.7 tok/s baseline ppl 23.1 int8 131 MB 39.4 tok/s 0.47x ppl 25.0 q4 174 MB 57.4 tok/s 0.69x ppl 28.4 → 4× smaller, 26% less RAM, 2.1× slower. → Quantize to fit in RAM, not to go fast.
What quantization actually bought
SmolLM2-135M-Instruct, pre-quantized ONNX from the Hub, on an Apple M4 (arm64, macOS): 4 performance cores, 32 new tokens, greedy, 2 warmup / 5 measured runs, each precision in its own process. Perplexity over 8 × 512-token windows of a pinned WikiText-2 slice. No instability warning on this run.
This is one machine, and that is the point: on hardware with native int8 kernels the speed column flips. The leaderboard is where that spread becomes visible.
What quantization costs, across real machines
Every row below was measured by someone on their own hardware and submitted as a result card. Speedup is relative to that machine's own fp32 baseline, so the numbers compare quantization choices rather than hardware budgets.
Loading results…
$ pip install quantcost $ quantcost run $ quantcost submit # opens the PR for you
One model, taken apart
A different measurement from the leaderboard above, and worth keeping separate: this is Qwen2.5-0.5B-Instruct through the authoring path — locally exported and quantized ONNX, driven by Optimum and by the C++ harness, rather than the pre-quantized artifacts the leaderboard measures. Different model, different runtime, so the numbers are not comparable with the table above and are not meant to be.
These runs predate a methodology fix. Peak RAM was measured with every precision in one process, which reports the union of their footprints, and the file cache was not warmed before timing — which penalises whichever artifact is largest, here the 2.4 GB fp32 baseline. Treat the relative throughput here with suspicion and the leaderboard as the trustworthy source; the size and perplexity columns are unaffected.
INT4 is bigger than INT8 here — weight-only quantization compresses MatMul but not the ~544 MB FP32 token-embedding Gather of a 152k-token vocab.
Same bit width, different scheme: PyTorch's per-tensor dynamic INT8 is both slower and accuracy-destroying on ARM.
The C++ harness beats the Python ORT wrapper (69.9 vs 38.3 tok/s) on the same INT8 model — lower per-step overhead, identical weights.
Full benchmark table
Model: Qwen/Qwen2.5-0.5B-Instruct · 64 new tokens, greedy · warmup 2 / measured 5 · perplexity on WikiText-2.
| Backend | Precision | Device | Size (MB) | Latency (s) | Throughput (tok/s) | Peak RAM (MB) | Perplexity |
|---|---|---|---|---|---|---|---|
| pytorch | fp32 | mps | 942.32 | 2.96 ± 0.02 | 21.60 | 2495.7 | 19.00 |
| ort-cpu | fp32 | cpu | 2404.97 | 3.25 ± 0.09 | 19.71 | 2910.8 | 19.00 |
| ort-cpu best | int8 | cpu | 604.98 | 1.67 ± 0.04 | 38.27 | 2861.8 | 20.10 |
| ort-cpu | int4 | cpu | 770.93 | 3.21 ± 0.07 | 19.97 | 3194.2 | 24.58 |
| pytorch | int8 | cpu | n/a¹ | 4.02 ± 0.08 | 15.92 | 5396.5 | 58.42 |
¹ PyTorch dynamic INT8 is an in-memory transform with no standalone on-disk artifact. The PyTorch FP32 row measures the hub's bf16 safetensors (942 MB); the ONNX FP32 export is true fp32 (2405 MB) — the INT8/INT4 rows are the meaningful size comparison. Correctness anchor: PyTorch and ORT FP32 give identical perplexity (19.00), confirming the export is numerically faithful before any quantization.
From Hugging Face checkpoint to on-device tokens
This is the pipeline that produces quantized artifacts, behind the .[quantize] extra. Benchmarking does not need any of it — quantcost run consumes ONNX already published on the Hub, which is what keeps that install torch-free. Reach for these commands when you want to quantize a model yourself.
Load
Resolve model + device (CPU / MPS / CUDA) and generate text with a swappable loader.
edgellm generateExport
Optimum-based ONNX FP32 export, verified numerically faithful against PyTorch.
edgellm exportQuantize
INT8 per-channel dynamic and INT4 block-wise weight-only, with honest backend availability.
edgellm quantizeBenchmark
Latency, tokens/sec, peak RAM, on-disk size and perplexity across every backend.
edgellm benchmarkReport
Regenerate the results table and bar chart straight from the measurements JSON.
edgellm reportFour ways to run the quantized model
Python + ONNX Runtime
PyTorch and ORT runners behind one interface, so every precision is measured the same way.
✓ measuredC++17 harness
edgellm_infer does autoregressive greedy decoding with a KV cache through the ORT C++ API, self-configuring from config.json.
✓ 69.9 tok/s INT8Android (ORT Mobile)
Kotlin app with the same KV-cache greedy decode, showing on-device tok/s in the UI.
✓ scaffold builds in Android StudioSnapdragon NPU
Two paths: Qualcomm AI Hub device-farm profiling, and the ORT QNN execution provider.
⏳ needs AI Hub tokenWhat the numbers actually taught me
Including the results that did not go the textbook way — kept in, not hidden.
"Quantize it, it'll be faster" is half wrong
Smaller is reliable. Faster is not. On an Apple M4, running the pre-quantized ONNX everyone actually downloads, int8 generates tokens at 0.47× the fp32 rate — and at 0.47× again on Qwen2.5-0.5B. The same figure twice, across a 135M and a 500M model, looks like a property of the runtime rather than one odd benchmark.
The cause is not slow integer arithmetic. Give the same artifacts a 512-token prefill instead of one token at a time and the int8 penalty roughly halves — 0.43× → 0.79× on SmolLM2, 0.59× → 0.97× on Qwen. That is the signature of a fixed per-step cost: ONNX Runtime unpacks the weights back to float inside each matmul, and one token cannot amortise unpacking a whole weight matrix. fp32 decode is bandwidth-bound on reading weights; int8 decode is compute-bound on unpacking them, so it does strictly more work while being 4× smaller.
int8 and q4 trade places
q4 is the better choice for generation and much the worse for prompt processing — exactly inverted from int8:
| prefill (512 tok) | decode (1 tok/step) | |
|---|---|---|
| int8 | 0.79× / 0.97× | 0.43× / 0.59× |
| q4 | 0.31× / 0.36× | 0.68× / 1.02× |
SmolLM2 / Qwen. They run through different kernels — MatMulNBits for q4, dynamic-quantize + MatMulInteger for int8 — with opposite strengths. So the right format depends on whether your workload is prompt-heavy or generation-heavy, which is not a question a single "speedup" number can answer.
Efficiency cores cost int8 two-thirds of its speed
This M4 has 4 performance and 6 efficiency cores. Pinning ONNX Runtime to all ten physical cores — the obvious default, and the one this tool originally shipped — measured int8 at 24.1 tok/s. Pinning the 4 performance cores measured 40.0. Same model, same bytes, 66% of the throughput thrown away, and run-to-run spread doubled.
An ORT parallel region ends on a barrier, so one thread scheduled onto an efficiency core gates the whole region, and the OS places threads differently each run. fp32 barely noticed — it is bandwidth-bound — so the bad default penalised precisely the quantized formats under test and made the headline look worse than reality. Every card now records which rule chose its thread count.
A cold file cache cost 4× on the baseline
ONNX Runtime memory-maps weights and faults them in lazily. Measured straight after download, the fp32 baseline read 17 tok/s; warm, the same machine and build read 64. Every speedup is divided by that number, so a cold run does not mis-measure one row — it corrupts the whole comparison.
The benchmark now reads each artifact through once before timing, runs every precision in its own process so peak RAM is not the union of all of them, and says so out loud when run-to-run variance exceeds 15%.
Not all INT8 is equal
ONNX Runtime's per-channel INT8 holds perplexity at 20.10. PyTorch's default per-tensor dynamic INT8 — same 8 bits — collapses to 58.42, because on ARM/qnnpack it also ignores reduce_range.
That row stays in the benchmark table as a demonstration of the failure mode rather than a result to bury.
INT4 didn't beat INT8 on size
771 MB vs 605 MB. Qwen2.5-0.5B has a ~152k-token vocab, so its tied embedding / lm_head weight dominates the file. Weight-only INT4 compresses the MatMuls but not the FP32 embedding Gather, and the tied head adds a 4-bit copy.
Perplexity degrades further too (24.58). On a smaller-vocab model, INT4 wins.
The same weights are faster in C++
Identical INT8 ONNX model: 38.3 tok/s through the Python wrapper, 69.9 tok/s through the C++ harness, with prefill down from 257 ms to 66 ms. The gap is per-step Python overhead, not the model.
Authoring-path measurement on Qwen2.5-0.5B, taken before the thread-pinning and file-cache fixes above. The direction is solid — a C++ decode loop avoids per-token interpreter overhead — but treat the magnitude as unverified until it is re-measured.
The compiler is a strong baseline
A hand-written NEON SDOT INT8 GEMM hits 59.2 GOPS vs 3.63 for a scalar loop with vectorization disabled — 16.3×. But at -O3 the compiler auto-vectorizes that same naive loop to ~74 GOPS.
The real lesson: beating a modern compiler on simple INT8 dot products takes register and cache blocking, not just intrinsics.
Run it yourself
- Python 3.10–3.14. The default install is ONNX Runtime, NumPy and a tokenizer — no PyTorch, no GPU, about ten seconds.
- quantcost models -m <repo> lists which precisions a model actually publishes before anything downloads.
- Quantizing your own artifacts needs the heavier .[quantize] extra — PyTorch, Optimum and the C++ harness. Benchmarking does not.
- CI runs ruff and pytest across Python 3.10–3.13, and asserts the default install never pulls in torch.
# 1 · benchmark your own machine $ pip install quantcost $ quantcost run $ quantcost submit # opens the PR for you # 2 · try another model $ quantcost run -m onnx-community/Qwen2.5-0.5B-Instruct # 3 · author your own quantized artifacts (heavy extra) $ pip install -e ".[quantize]" $ edgellm export && edgellm quantize $ cmake -S cpp -B cpp/build -DCMAKE_BUILD_TYPE=Release && cmake --build cpp/build
Run it on your machine
pip install quantcost — about ten seconds, no PyTorch, no GPU. Then quantcost submit puts your hardware on the leaderboard. The quantization pipeline, C++ harness, SIMD kernel and Android app are all in the same MIT-licensed repository.