Score a whole retrieval batch against the query in one Clef call, then fill a token budget with the chunks worth sending to your LLM. Chunks are kept verbatim or removed with a recorded reason, so citations stay auditable.
Clef answers P(relevant) for every chunk against your query. The compactor ranks by score, drops what sits below the threshold, and fills the remaining budget with the best chunks. Every cut carries a reason your code can log or display.
Batched noul questions, deterministic ranking, greedy budget fill.
Kept chunks keep their exact text. Summaries would break citations; deletion keeps them honest.
Each dropped chunk records irrelevant or budget_exhausted, plus its score and token count.
Open-weights clef-flash 9B, float16 on 2x Kaggle T4, 20 hand-labelled cases, 104 chunks. Reproducible from the public kernel. The replay baseline (simulated scorer) and a hosted-API run are documented in the repo.
| metric | open-weights 9B on 2xT4 | replay baseline |
|---|---|---|
| chunk accuracy | 0.712 | 0.990 |
| kept precision | 0.771 | 1.000 |
| relevant recall | 0.746 | 0.984 |
| kept F1 | 0.758 | 0.992 |
| context tokens saved | 38.4% | 34.0% |
| scoring latency p50 | 1,274 ms | <1 ms |
| cost per 1k calls | $0.00 self-hosted | $0.019 |
Published numbers from Cloudflare's "Introducing Clef" blog post. Laya runs locally and wins raw latency; Clef wins decision quality by a wide margin; clef-flash sits in between.
| benchmark | clef | clef-flash | laya |
|---|---|---|---|
| BFCL, case exact | 98.47 | 98.76 | 38.13 |
| ToolRet, nDCG@10 | 69.19 | 66.43 | 12.69 |
| API-Bank accuracy | 91.93 | 93.11 | 11.41 |
| median latency, ms | 209.3 | 38.8 | 5.8 |
| p95 latency, ms | 238.6 | 122.4 | 222.5 |
| context window | 65,536 | 65,536 | 32k reported |
from clef_compactor import ClefCompactor compactor = ClefCompactor() # reads CLEF_ACCOUNT_ID / CLEF_API_TOKEN result = compactor.compact("What is the refund policy?", chunks, token_budget=1000) context = "\n\n".join(result.kept_texts())
Async client, CLI, an OpenAI-compatible endpoint and LangChain / LlamaIndex adapters ship in the box. See the README for copy-paste examples of each.