Keep the evidence.
Cut the noise.

Score a whole retrieval batch against the query in one Clef call, then fill a token budget with the chunks worth sending to your LLM. Chunks are kept verbatim or removed with a recorded reason, so citations stay auditable.

pip install clef-compactor source
12s · 1280x720 · no audiodocs/brag.mp4
how it works

One question per chunk. One call per 64 chunks.

Clef answers P(relevant) for every chunk against your query. The compactor ranks by score, drops what sits below the threshold, and fills the remaining budget with the best chunks. Every cut carries a reason your code can log or display.

pipeline

The gate, animated

Batched noul questions, deterministic ranking, greedy budget fill.

verbatim

Never rewritten

Kept chunks keep their exact text. Summaries would break citations; deletion keeps them honest.

auditable

Reasoned cuts

Each dropped chunk records irrelevant or budget_exhausted, plus its score and token count.

measured results

Real model, measured on real hardware

Open-weights clef-flash 9B, float16 on 2x Kaggle T4, 20 hand-labelled cases, 104 chunks. Reproducible from the public kernel. The replay baseline (simulated scorer) and a hosted-API run are documented in the repo.

metricopen-weights 9B on 2xT4replay baseline
chunk accuracy0.7120.990
kept precision0.7711.000
relevant recall0.7460.984
kept F10.7580.992
context tokens saved38.4%34.0%
scoring latency p501,274 ms<1 ms
cost per 1k calls$0.00 self-hosted$0.019
clef vs laya

Different points on the same trade-off

Published numbers from Cloudflare's "Introducing Clef" blog post. Laya runs locally and wins raw latency; Clef wins decision quality by a wide margin; clef-flash sits in between.

benchmarkclefclef-flashlaya
BFCL, case exact98.4798.7638.13
ToolRet, nDCG@1069.1966.4312.69
API-Bank accuracy91.9393.1111.41
median latency, ms209.338.85.8
p95 latency, ms238.6122.4222.5
context window65,53665,53632k reported
quick start

Three lines to a smaller context

from clef_compactor import ClefCompactor

compactor = ClefCompactor()  # reads CLEF_ACCOUNT_ID / CLEF_API_TOKEN
result = compactor.compact("What is the refund policy?", chunks, token_budget=1000)
context = "\n\n".join(result.kept_texts())

Async client, CLI, an OpenAI-compatible endpoint and LangChain / LlamaIndex adapters ship in the box. See the README for copy-paste examples of each.

limitations

What this tool is not