Score a whole retrieval batch against the query in one Clef call, then fill a token budget with the chunks worth sending to your LLM. Chunks are kept verbatim or removed with a recorded reason, so citations stay auditable.
Clef answers P(relevant) for every chunk against your query. The compactor ranks by score, drops what sits below the threshold, and fills the remaining budget with the best chunks. Every cut carries a reason your code can log or display.
Batched noul questions, deterministic ranking, greedy budget fill.
Kept chunks keep their exact text. Summaries would break citations; deletion keeps them honest.
Each dropped chunk records irrelevant or budget_exhausted, plus its score and token count.
Deterministic simulated scorer, 20 cases, 104 chunks, seed 20261001. This validates the pipeline; it does not measure Clef. Live numbers land via evals/run_eval.py --mode live.
| metric | value |
|---|---|
| chunk accuracy | 0.990 |
| kept precision | 1.000 |
| relevant recall | 0.984 |
| kept F1 | 0.992 |
| context tokens saved | 34.0% |
| cost per 1k calls | $0.019 |
Published numbers from Cloudflare's "Introducing Clef" blog post. Laya runs locally and wins raw latency; Clef wins decision quality by a wide margin; clef-flash sits in between.
| benchmark | clef | clef-flash | laya |
|---|---|---|---|
| BFCL, case exact | 98.47 | 98.76 | 38.13 |
| ToolRet, nDCG@10 | 69.19 | 66.43 | 12.69 |
| API-Bank accuracy | 91.93 | 93.11 | 11.41 |
| median latency, ms | 209.3 | 38.8 | 5.8 |
| p95 latency, ms | 238.6 | 122.4 | 222.5 |
| context window | 65,536 | 65,536 | 32k reported |
from clef_compactor import ClefCompactor compactor = ClefCompactor() # reads CLEF_ACCOUNT_ID / CLEF_API_TOKEN result = compactor.compact("What is the refund policy?", chunks, token_budget=1000) context = "\n\n".join(result.kept_texts())
Async client, CLI, an OpenAI-compatible endpoint and LangChain / LlamaIndex adapters ship in the box. See the README for copy-paste examples of each.