Metadata-Version: 2.4
Name: kvcalib
Version: 0.4.0
Summary: Distribution-free statistical guarantees for KV-cache eviction budgets, calibrated on your own workload.
Project-URL: Homepage, https://github.com/amir2628/KVCalib
Project-URL: Repository, https://github.com/amir2628/KVCalib
Author: The KVCalib Authors
License-Expression: Apache-2.0
License-File: LICENSE
Keywords: calibration,conformal-prediction,kv-cache,llm-inference,risk-control
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: License :: OSI Approved :: Apache Software License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Python: >=3.10
Requires-Dist: numpy>=1.24
Requires-Dist: pydantic>=2.0
Requires-Dist: typer>=0.12
Provides-Extra: dev
Requires-Dist: hypothesis>=6.100; extra == 'dev'
Requires-Dist: opentelemetry-sdk>=1.20; extra == 'dev'
Requires-Dist: pytest-cov>=4.0; extra == 'dev'
Requires-Dist: pytest>=7.0; extra == 'dev'
Provides-Extra: kvpress
Requires-Dist: kvpress>=0.5; extra == 'kvpress'
Requires-Dist: torch>=2.2; extra == 'kvpress'
Requires-Dist: transformers>=4.45; extra == 'kvpress'
Provides-Extra: otel
Requires-Dist: opentelemetry-api>=1.20; extra == 'otel'
Description-Content-Type: text/markdown

# KVCalib

KVCalib turns "how aggressively can I evict my KV cache" into a decision backed by a
**distribution-free statistical guarantee**, calibrated on your own workload, instead
of a benchmark number computed on someone else's traffic.

It is a thin calibration/monitoring layer on top of [`kvpress`](https://github.com/NVIDIA/kvpress)
(NVIDIA's KV-cache eviction backends and benchmark harness) -- not a new eviction
heuristic, not a fork of `kvpress`, and not a competitor to it. See
[`docs/architecture.md`](docs/architecture.md) for the full design rationale and
[`docs/understanding_the_guarantee.md`](docs/understanding_the_guarantee.md) for what
the guarantee does and doesn't promise, including how it differs from H2O's own
approximation-ratio guarantee.

**Status**: Phase 2 complete. Multi-backend: H2O, SnapKV, and StreamingLLM are all
validated (monotonicity + repeat-split coverage) on exact-match loss; `PyramidKV` is
implemented but documented as excluded on non-CUDA hardware (see
[`docs/monotonicity_reports/pyramidkv_exact_match.md`](docs/monotonicity_reports/pyramidkv_exact_match.md)).
Cross-backend comparison: SnapKV achieves the smallest certified budget at alpha=0.80
(see [Cross-backend comparison](#cross-backend-comparison)). Metric-degradation loss
is implemented but **parked**, not validated -- its own monotonicity check failed on
real data (see
[`docs/monotonicity_reports/h2o_metric_degradation.md`](docs/monotonicity_reports/h2o_metric_degradation.md)).
An ACI-based drift monitor (`kvcalib.monitor.aci`) is implemented and validated
against a real distribution shift (see [Negative controls](#negative-controls)).
Phase 3 complete: `MondrianRiskControl` adds group-conditional (Mondrian-style)
calibration, analysis/reporting only for now, validated on real data at both a
genuine budget tightening and a genuine calibration-data cost (see
[Group-conditional coverage](#group-conditional-coverage)). Phase 4 complete:
`kvcalib.telemetry.otel_export` adds opt-in OpenTelemetry span/event export (see
[Install](#install)); `kvcalib.monitor.streaming.StreamingCalibrator` adds
windowed recalibration triggered by the existing drift monitor, not a new online
calibration algorithm; score-threshold eviction knobs were considered and scoped
out (all four backends and `kvpress`'s shared scoring base are rank-based, with
no score-cutoff quantity to expose); packaging polish and a PyPI Trusted
Publishing release workflow round out the phase. See
[`KVCALIB_PROJECT_SPEC.md`](KVCALIB_PROJECT_SPEC.md) for full phase-by-phase
detail and `CHANGELOG.md` for the versioned history.

## Install

```bash
pip install "kvcalib[kvpress]"
```

Not yet published to PyPI -- until the first release goes out, install from a
checkout instead:

```bash
pip install -e ".[kvpress]"   # from a checkout, until the PyPI release is published
```

**Known upstream issue on Python 3.13**: `kvpress` 0.5.4 (latest as of this writing)
transitively depends on `google-fire`, which still imports the `pipes` stdlib module
removed in Python 3.13 -- `import kvpress` fails with `ModuleNotFoundError: No module
named 'pipes'` on 3.13, regardless of anything in this project. This is a `kvpress`/
`fire` compatibility gap, not a KVCalib issue -- `kvcalib`'s own package imports lazily
and does not hit it. Until it's fixed upstream, either use Python 3.10-3.12 for the
`[kvpress]` extra, or work around it locally with a one-line shim:
`echo 'from shlex import quote' > "$(python3 -c "import sysconfig; print(sysconfig.get_paths()['purelib'])")/pipes.py"`.

**Known upstream limitation: `PyramidKVBackend` cannot run generation on non-CUDA
hardware.** PyramidKV gives each transformer layer a different KV-cache length by
design; `transformers`' `eager`/`sdpa`/`flex_attention` paths build one causal mask
(sized for one assumed cache length) and reuse it for every layer, which crashes
whenever a prompt/budget pair produces non-uniform per-layer lengths (common, not an
edge case). Only `flash_attention_2`/`_3` are exempt, and both are hard-gated in
`transformers` to CUDA/HIP/Cambricon-MLU hardware -- unavailable on Apple Silicon
regardless of attention-kernel package. See
[`docs/monotonicity_reports/pyramidkv_exact_match.md`](docs/monotonicity_reports/pyramidkv_exact_match.md)
for the full investigation, including two remediation paths (PyTorch 2.13's MPS
FlexAttention, the `mps-flash-attn` package) that were concretely tried and ruled out
rather than assumed incompatible. `PyramidKVBackend` is implemented and its
budget-to-ratio translation is correct -- this is a `kvpress`/`transformers`
limitation on this hardware, not a bug in this project -- but it is excluded from
Phase 2's coverage validation and cross-backend comparison as a result.

Calibration and deployment require `kvpress` + `transformers` + `torch` and an actual
model -- there is no synthetic substitute for "run the model," since the whole point
is measuring eviction's real effect. The core calibration math
(`kvcalib.core`) has no such dependency and can be tested and used standalone (e.g.
against precomputed loss curves).

Optional: `pip install "kvcalib[otel]"` (or `.[otel]` from a checkout) adds
`opentelemetry-api` for `kvcalib.telemetry.otel_export`'s span/event export
(`export_calibration_span`, `export_generation_event`) -- this is opt-in and
caller-driven, not required for calibration or deployment.

## Quickstart

```python
from transformers import AutoModelForCausalLM, AutoTokenizer
from kvcalib import CalibrationPrompt, Calibrator
from kvcalib.backends import H2OBackend

tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen2.5-0.5B-Instruct")
model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-0.5B-Instruct", attn_implementation="eager")

calibration_set = [
    CalibrationPrompt(prompt="...", reference_answer="..."),
    # ...prompts representative of your own workload...
]

calibrator = Calibrator(alpha=0.1, backend=H2OBackend())
calibrator.fit(calibration_set, model, budget_grid=[8, 16, 32, 64, 512], tokenizer=tokenizer)
calibrator.save("calibration.json")
print(calibrator.guarantee)

# --- at deploy time ---
calibrator = Calibrator.load("calibration.json")
cache = calibrator.risk_controlled_cache()  # drop-in kvpress-compatible press

from kvpress import KVPressTextGenerationPipeline
pipe = KVPressTextGenerationPipeline(model=model, tokenizer=tokenizer)
print(pipe("Context: ...", question="...", press=cache, max_new_tokens=10))
```

Running `examples/single_backend_quickstart.py` end to end (a tiny, 14-prompt
calibration set of short factual questions) prints:

```
Expected exact_match accuracy loss from evicting to budget=8 tokens is <= 0.1,
assuming calibration and live traffic are exchangeable. Calibrated on n=14 prompts
on 2026-07-22 using the h2o backend. This guarantee holds only if calibration and
live traffic are exchangeable [...]
Generated: Madrid
```

See `examples/` for this quickstart, a cross-backend comparison, and a fast
(no-GPU) walkthrough of the negative-control refusal path.

## Coverage validation

This is the project's actual evidence, not a benchmark score borrowed from someone
else's traffic: real model + real `kvpress` H2O eviction, run on LongBench
`multifieldqa_en`, then repeat-split many times from that one real, measured loss
matrix (see `tests/coverage_validation/`, `pytest tests/coverage_validation -m gpu`).

**Setup**: Qwen2.5-0.5B-Instruct, H2O backend, exact-match-with-containment loss,
n=39 calibration prompts (all `multifieldqa_en` records whose full document fits
under the 4000-token budget grid ceiling -- see
[`docs/monotonicity_reports/h2o_exact_match.md`](docs/monotonicity_reports/h2o_exact_match.md)
for why), 70/30 calibration/test split, 200 repeat-splits per alpha.

| alpha target | result | mean realized loss (held-out) | n_calib | n_test |
|---|---|---|---|---|
| 0.01 | **refused** -- `InsufficientCalibrationDataError` (n_calib=27 < required 100) | -- | 27 | 12 |
| 0.05 | falls back to largest budget (4096) in all 200 splits | 0.556 | 27 | 12 |
| 0.10 | falls back to largest budget (4096) in all 200 splits | 0.556 | 27 | 12 |
| 0.20 | falls back to largest budget (4096) in all 200 splits | 0.556 | 27 | 12 |
| 0.60 | falls back to largest budget (4096) in all 200 splits | 0.556 | 27 | 12 |
| 0.80 | genuine, non-trivial budget selection -- 512, 1024, or 4096 depending on the split | 0.722 | 27 | 12 |

These are the real, reported numbers -- not illustrative ones. Two things worth being
upfront about, since both are legitimate outcomes this tool is specifically designed
to surface rather than hide:

- **alpha=0.01 is refused, not silently miscalibrated.** With only 27 calibration
  prompts in a 70/30 split of this 39-prompt sample, the derived minimum-n floor for
  alpha=0.01 is 100 (`kvcalib.core.validation.minimum_calibration_size`); the guard
  fires and the tool says so, rather than returning a technically-computed but
  meaningless budget.
- **alpha=0.05/0.10/0.20/0.60 all fall back to the largest budget in the grid, with
  an identical realized loss, in every one of the 200 repeat-splits.** Qwen2.5-0.5B-
  Instruct's own baseline exact-match accuracy on this task -- even with *no*
  eviction at all -- has a loss around 0.56 (see the monotonicity report). None of
  these four targets is achievable by any budget for this model on this task, and
  KVCalib correctly refuses to pretend otherwise: it returns the least-aggressive
  budget in the grid and says so in the guarantee text, rather than quietly
  certifying a number the model can't back up. (An earlier draft of this table
  described alpha=0.60 as "calibrating a non-trivial budget" -- checking directly
  which budget got selected across all 200 splits showed that was wrong: it selects
  4096 every time, identically to 0.05/0.10/0.20. Left here as a correction rather
  than silently fixed, since it's exactly the kind of claim this table exists to get
  right.) alpha=0.80 is included specifically to also show what happens once the
  target is above that floor: real, non-trivial budget selection (512, 1024, or 4096
  depending on the split), with realized loss tracking the target as it relaxes.

A more capable pilot model would very likely clear the smaller alpha targets too;
this table reports what a real 0.5B model on real hardware in this session actually
does, which is the entire point of calibrating on your own workload rather than
trusting someone else's benchmark number.

### Cross-backend comparison

The same repeat-split protocol (real generation, real eviction, 200 repeat-splits,
70/30 calibration/test) extended to SnapKV and StreamingLLM: at alpha in
{0.05, 0.10, 0.20}, all three backends fall back to the largest grid budget in every
split, identically -- this model's own baseline exact-match floor on this task, not a
backend difference. At alpha=0.80, the one target already known to clear that floor,
the backends separate:

| rank | backend | median selected budget | mean selected budget | fraction at max budget | mean realized loss |
|---|---|---|---|---|---|
| 1 | **SnapKV** | **256** | **349** | 1% | 0.743 |
| 2 | H2O | 1024 | 1782 | 28% | 0.722 |
| 3 | StreamingLLM | 4096 | 2606 | 51.5% | 0.684 |

SnapKV achieves the smallest certified budget at alpha=0.80 by a wide margin --
consistent with SnapKV's content-scored eviction outperforming both H2O's cumulative-
attention scoring and StreamingLLM's purely positional sink+recency scheme on this
single-document QA task (see the per-backend monotonicity reports for the underlying
per-budget curves). See
[`docs/coverage_validation/cross_backend_comparison.md`](docs/coverage_validation/cross_backend_comparison.md)
for the full per-backend repeat-split tables and discussion.

### Group-conditional coverage

`MondrianRiskControl` (Phase 3) calibrates a separate budget per group instead of
one pooled budget -- see
[`docs/group_conditional_risk_control.md`](docs/group_conditional_risk_control.md)
for what this does and doesn't change (analysis/reporting only: it does not change
what `risk_controlled_cache()` serves). Validated on the same real H2O +
`multifieldqa_en` matrix, grouped by context-length bucket ("short"/"long", split
at the median): at alpha=0.80, the pooled budget (1782 tokens) is the wrong number
for either group -- short-context prompts only need **1508** (15% less), long-context
prompts need **2171** (22% more). At alpha=0.05, this dataset shows the real cost:
`n_min=20` exceeds the entire 17-member "long" group before any split is even
drawn, a structural floor, not unlucky sampling, so that group correctly falls back
to the pooled guarantee instead of fabricating one. See
[`docs/coverage_validation/group_conditional_coverage.md`](docs/coverage_validation/group_conditional_coverage.md)
for the full repeat-split table and discussion.

## Negative controls

All three are run on real model output and detailed in
[`docs/negative_controls.md`](docs/negative_controls.md); the first two are mandatory
per Sec. 7.2 of the project brief, the third is Phase 2's drift monitor validated
against the same real distribution shift as the second:

| control | result |
|---|---|
| Broken monotonicity (real per-prompt losses, budget assignment shuffled per prompt) | `MonotonicityViolationError` raised; calibration refused |
| Broken exchangeability (calibrated on `multifieldqa_en` at alpha=0.80, tested on `hotpotqa`) | Calibrated budget **1024** (a genuine ~75% eviction, not a fallback); realized loss **0.867** on `hotpotqa` vs. target alpha=0.80 -- guarantee measurably broken, as expected |
| Drift monitor (`monitor.aci.ACIDriftMonitor`, same shift replayed live-traffic-style) | Silent across all 39 in-control rounds; raised `GuaranteeInvalidatedError` at round 28 of 60 on the shifted replay |

The exchangeability control specifically calibrates at an alpha that selects a
genuinely compressed budget (see `docs/negative_controls.md` for why alpha=0.10,
used in an earlier draft, was a weaker version of this control: it fell back to the
largest budget in the grid, so there was no real eviction decision in effect for the
distribution shift to actually invalidate).

## License

Apache 2.0.
