Keep the evidence.
Cut the noise.

Score a whole retrieval batch against the query in one Clef call, then fill a token budget with the chunks worth sending to your LLM. Chunks are kept verbatim or removed with a recorded reason, so citations stay auditable.

pip install clef-compactor source
12s · 1280x720 · no audiodocs/brag.mp4
how it works

One question per chunk. One call per 64 chunks.

Clef answers P(relevant) for every chunk against your query. The compactor ranks by score, drops what sits below the threshold, and fills the remaining budget with the best chunks. Every cut carries a reason your code can log or display.

pipeline

The gate, animated

Batched noul questions, deterministic ranking, greedy budget fill.

verbatim

Never rewritten

Kept chunks keep their exact text. Summaries would break citations; deletion keeps them honest.

auditable

Reasoned cuts

Each dropped chunk records irrelevant or budget_exhausted, plus its score and token count.

measured results

Replay baseline, committed and reproducible

Deterministic simulated scorer, 20 cases, 104 chunks, seed 20261001. This validates the pipeline; it does not measure Clef. Live numbers land via evals/run_eval.py --mode live.

metricvalue
chunk accuracy0.990
kept precision1.000
relevant recall0.984
kept F10.992
context tokens saved34.0%
cost per 1k calls$0.019
clef vs laya

Different points on the same trade-off

Published numbers from Cloudflare's "Introducing Clef" blog post. Laya runs locally and wins raw latency; Clef wins decision quality by a wide margin; clef-flash sits in between.

benchmarkclefclef-flashlaya
BFCL, case exact98.4798.7638.13
ToolRet, nDCG@1069.1966.4312.69
API-Bank accuracy91.9393.1111.41
median latency, ms209.338.85.8
p95 latency, ms238.6122.4222.5
context window65,53665,53632k reported
quick start

Three lines to a smaller context

from clef_compactor import ClefCompactor

compactor = ClefCompactor()  # reads CLEF_ACCOUNT_ID / CLEF_API_TOKEN
result = compactor.compact("What is the refund policy?", chunks, token_budget=1000)
context = "\n\n".join(result.kept_texts())

Async client, CLI, an OpenAI-compatible endpoint and LangChain / LlamaIndex adapters ship in the box. See the README for copy-paste examples of each.

limitations

What this tool is not