Metadata-Version: 2.4
Name: deltacert
Version: 1.2.0
Summary: Calibrated divergence certification for LLM serving systems
License-Expression: Apache-2.0
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Requires-Python: >=3.9
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: numpy
Requires-Dist: torch>=2.0
Requires-Dist: transformers>=4.38
Requires-Dist: accelerate>=0.30
Requires-Dist: cryptography>=42.0
Provides-Extra: bnb
Requires-Dist: bitsandbytes>=0.43; extra == "bnb"
Provides-Extra: peft
Requires-Dist: peft>=0.9; extra == "peft"
Provides-Extra: vllm
Requires-Dist: vllm>=0.4; extra == "vllm"
Provides-Extra: gptq
Requires-Dist: gptqmodel>=1.0; extra == "gptq"
Requires-Dist: optimum>=1.20; extra == "gptq"
Provides-Extra: validation
Requires-Dist: datasets>=2.20; extra == "validation"
Requires-Dist: lm-eval>=0.4; extra == "validation"
Requires-Dist: matplotlib; extra == "validation"
Requires-Dist: human-eval>=1.0; extra == "validation"
Dynamic: license-file

# DeltaCert

[![PyPI](https://img.shields.io/pypi/v/deltacert)](https://pypi.org/project/deltacert/) [![License](https://img.shields.io/pypi/l/deltacert)](LICENSE) [![Python](https://img.shields.io/pypi/pyversions/deltacert)](https://pypi.org/project/deltacert/)

**Certify any change to your LLM serving stack — quantization, engine upgrades, batch size, model updates — with a mathematical bound, before you deploy.** Unlock the cost savings (2x KV-cache batch capacity, 75% VRAM reduction) that fear of silent regressions currently blocks, and catch the two failure classes standard benchmarks and single-pass checks both miss: benchmark-blind long-generation forking, and feedback-driven collapse that passes every logit-level check and only shows up in the model's own output.

Built by [Threvo Labs](https://pypi.org/user/Shorya/).

[PyPI](https://pypi.org/project/deltacert/) · [SPEC.md](SPEC.md) · [LICENSE](LICENSE)

```bash
pip install deltacert
```

---

## The catch — two ways a change can lie to you, both caught

**Way 1: a benchmark says "fine" but generations fork anyway.**

| Config | Short eval said | d_COMM | safe_until_token | failure_after_token | Verdict |
|---|---|---|---|---|---|
| Llama-3.1-8B fp16 → nf4 | GSM8K 5-shot exact-match +1.0% (looks safe) | 0.00 | 31 | 14 | **unsafe** |
| Qwen2.5-7B fp16 → nf4 | GSM8K exact-match +0.0% (looks safe) | 0.00 | 17 | 18 | **unsafe** |

Same pattern, two model families, two labs, two tokenizers. GSM8K missed both. DeltaCert's trajectory certification caught both — generations fork from the fp16 reference within ~15-18 tokens on long-form coding tasks, well before a short-form benchmark would ever see it.

**Way 2 — the one that made us rewrite our own default mode: a config passes *every* teacher-forced check and is still destroyed.**

| Config | Single-position | Trajectory (30,643 positions) | Downstream reality | Verdict |
|---|---|---|---|---|
| Qwen2.5-7B fp8 KV-cache | safe (d=3.72, cosines ≥0.9998) | safe (d≥3.52 at every position) | GSM8K 0.88 → 0.00, 0/100 correct | **unsafe** |
| Llama-3.1-8B fp8 KV-cache, identical flag | safe | safe | GSM8K within noise | genuinely safe |

The identical engine flag is benign on Llama and catastrophic on Qwen. Both teacher-forced modes — single-position and full trajectory — certify the Qwen collapse safe, because the damage doesn't live in the logits; it lives in the autoregressive feedback loop, which teacher forcing structurally can't see. Catching this needed a third instrument: a free-running collector that runs the deployed decode policy on both engines and measures the actual output process, with a McNemar-exact-test guard so a benign fork's ordinary spontaneous-repetition rate can't be mistaken for caused collapse. It fires decisively on Qwen (79% excess degeneration, p≈10⁻¹⁰) and stays quiet on Llama (2.3%, p=0.63) — sensitivity on the real failure, specificity on the real clean case, same instrument, same thresholds.

Robustness-checked: excluding all 7 references with degenerate repetition, every trajectory statistic is identical (`cert_trajectory_clean7.json`); the fp8-KV single-position measurement was independently reproduced on a different host/stack to three decimal places.

Reproduce it yourself:

```bash
deltacert generate-cases --model meta-llama/Llama-3.1-8B-Instruct --output cases.jsonl
deltacert certify --model meta-llama/Llama-3.1-8B-Instruct --quantization int4 \
    --checks trajectory --trajectory-cases cases.jsonl
```

Every number above traces to a real certificate in `validation_results/`; every row in the table below reproduces with one script.

## The proof

Seven real flagship tests, each a full before/after comparison on a real model, real GPU, real downstream benchmark:

| Change | Business gain | d_COMM | Downstream effect | Verdict |
|---|---|---|---|---|
| Llama-3.1-8B batch=1 → batch=64 | 64 concurrent requests, same GPU | 6.21 | GSM8K -1.0 pt | ✅ Safe |
| vLLM 0.8.5 → vLLM 0.9.0 | take the upgrade same-week, not months later | 16.09 | GSM8K -1.0 pt | ✅ Safe |
| KV cache default → fp8 (vLLM native) | 2x concurrent capacity | 4.83 | GSM8K -1.0 pt | ✅ Safe |
| gpt-4o-mini pinned snapshot → current alias | same-day provider-drift check | 6.65 | canary acc +0.0 pt | ✅ Safe |
| Standard decode → speculative decode (k=5) | claimed ~2x throughput | 15.38 | GSM8K +0.0 pt, **measured 0.28x** (slower) | ✅ Safe on quality, not on speed |
| Llama-3.1-8B fp16 → nf4 (W4) | +60% VRAM reduction | 0.00 | forks at token 14 on long generations | ❌ Unsafe |
| Llama-3.1-8B fp16 → GPTQ int4 | +75% VRAM reduction | 1.16 | GSM8K -8.0 pts | ❌ Unsafe |

Five safe, two unsafe. A tool that only ever says "safe" isn't measuring anything — the two unsafe rows above are DeltaCert doing its job.

## How it works

DeltaCert compares output distributions before and after a change. `d_COMM` ("commutator distance," from the operator-algebraic commutator bound it's derived from) is computed from the cosine similarity `c` between two runs:

```
Δ = 4c√(1-c²)        (commutator magnitude)
d = -log(Δ/2)        (algebraic distance)
divergence_bound = 2·exp(-d)
```

`d` is an algebraic distance; `2e⁻ᵈ` is the certified bound on output divergence — deterministic, minutes to compute, checkable at every token position, no eval harness or labeled data required.

> **"Certified" throughout this document means:** measured against a calibrated threshold with a stated bound — not a guarantee of downstream quality.

- **Full derivation, clamp behavior, per-method calibration (with sample sizes disclosed), and the top-k logprobs caveat:** see `SPEC.md`.
- **This README asserts. The spec defends. Nothing here is a proof.**

## Getting started

```bash
pip install deltacert

deltacert certify --model meta-llama/Llama-3.1-8B-Instruct --quantization int8
```

That uses DeltaCert's shipped reference calibration (from the 7-test suite above). For a threshold tuned to **your own model and workload**, run the sweep yourself:

```bash
deltacert capture --model your-model --output baseline.npz
deltacert capture --model your-model --quantization int8 --output candidate.npz
deltacert calibrate --baseline baseline.npz --candidates candidate.npz \
    --names int8 --downstream-file your_evals.json
```

`deltacert certify` always tells you when it's using the shipped calibration instead of your own.

For KV-cache and any change whose damage can be feedback-driven, add a free-running check (separate subcommand — its certificate is McNemar/degeneration-based, not `d_COMM`-based):

```bash
deltacert free-running --model your-model --kv-cache-dtype fp8 --output cert_free_running.json
```

## Integrations

- **vLLM plugin** — official `vllm.general_plugins` entry point, already wired in this package. Opt-in only: complete no-op unless `DELTACERT_ENFORCE=1` is set, so `pip install deltacert` is safe in a shared image; when enabled, a serving engine refuses to start on an uncertified change.
- **CI/CD gate** — `python -m deltacert.integrations.cicd_hook --cert ./cert.json` exits 1 (blocks the pipeline) if not certified, 0 if certified. Works with GitHub Actions, GitLab CI, Jenkins, or any CI that checks exit codes.
- **HuggingFace auto-wiring** — `from deltacert.integrations.hf_integration import auto_certify` picks the right collectors for you from what's active in your config (quantization, LoRA, prefix cache) instead of calling `certify_system()` with raw parameters yourself.

## Verifying a certificate

Optional — certificates work unsigned; signing adds tamper-evidence for sharing certs across teams or with auditors.

```bash
deltacert keygen --private-key mykey.pem --public-key mykey.pub
deltacert sign --cert cert.json --key-file mykey.pem
deltacert verify --cert cert.json --key-file mykey.pub
```

`verify` exits 0 if the signature is valid, 1 if the certificate was modified after signing or signed by a different key. Every certificate also carries a `validation_status` field (`flagship_validated` vs `implemented_pending_validation`) — signed as part of the payload, so a signature can never make an unvalidated collector's result look more trustworthy than it is.

Keep your private key secret — never commit it, never share it. Only the public key is meant to be distributed.

All 13 reference certificates in `validation_results/` are signed with Threvo's key (`deltacert-public.pem`, committed in this repo). Don't take our numbers on faith — check them yourself:

```bash
deltacert verify --cert validation_results/weight_quant/cert_nf4.json --key-file deltacert-public.pem
```

## What it certifies

DeltaCert ships all 21 collectors described in the design — the code is real and implemented, and doesn't get deleted just because a given check hasn't been run in a full end-to-end validation yet. 8 have real flagship validation results behind them so far (the tables above, including the free-running collector that catches feedback-driven failures teacher-forced checks miss); the rest are working code with the same math, not yet run through that process.

| # | Check | CLI-drivable | Status |
|---|---|---|---|
| 1 | `weight_quant` | ✅ | ✅ validated |
| 2 | `kv_cache_quant` | ✅ | ✅ validated |
| 3 | `batch_divergence` | ✅ (needs vLLM) | ✅ validated |
| 4 | `spec_decoding` | ✅ (needs vLLM) | ✅ validated |
| 5 | `engine_swap` | ✅ | ✅ validated |
| 6 | `provider_drift` | ✅ | ✅ validated |
| 7 | `trajectory` | ✅ | ✅ validated |
| 8 | `free_running` | ✅ (needs vLLM) | ✅ validated |
| 9 | `activation_quant` | ✅ | 🔬 implemented — validation run pending |
| 10 | `prefix_cache` | ✅ | 🔬 implemented — validation run pending |
| 11 | `lora` | ✅ | 🔬 implemented — validation run pending |
| 12 | `model_swap` | ✅ | 🔬 implemented — validation run pending |
| 13 | `prompt_swap` | ✅ | 🔬 implemented — validation run pending |
| 14 | `sparse_attention` | Python API only | 🔬 implemented — validation run pending |
| 15 | `moe_token_dropping` | Python API only | 🔬 implemented — validation run pending |
| 16 | `neuron_skipping` | Python API only | 🔬 implemented — validation run pending |
| 17 | `allreduce_tp` | Python API only | 🔬 implemented — validation run pending |
| 18 | `alltoall_ep` | Python API only | 🔬 implemented — validation run pending |
| 19 | `pipeline_parallel` | Python API only | 🔬 implemented — validation run pending |
| 20 | `kv_transfer` | Python API only | 🔬 implemented — validation run pending |
| 21 | `gradient_compress` | Python API only | 🔬 implemented — validation run pending |

"Python API only" means the check needs code you supply (a custom `compress_fn`, attention mask, etc.) — the CLI can't conjure that for you; see `import deltacert as dc; dc.certify_system(...)`.

## Limitations (stated up front, not discovered by you later)

- Validated on two open model families (Llama-3.1-8B-Instruct, Qwen2.5-7B-Instruct) plus one hosted API (gpt-4o-mini); serving-time flagships beyond weight/KV-cache quantization are validated on Llama only.
- Per-method calibration (e.g. the bnb/GPTQ thresholds) derives from n=6 configs per model — not a settled constant. Run `deltacert calibrate` on your own model/workload rather than trusting the shipped default for anything production-critical.
- The provider_drift result above is a same-day proxy (pinned snapshot vs. current alias), not the real weekly-cadence drift measurement, which needs two runs across real time.
- `d_comm` is a reliable within-method damage indicator but is not directly comparable across different compression methods — see `SPEC.md` for the bnb-vs-GPTQ false-negative this caused and how it's handled.
- The feedback-driven failure class (fp8 KV-cache on Qwen) currently has one confirmed member after a pre-registered four-candidate hunt; whether other configurations populate it is still open.

## Roadmap

- More models, more downstream tasks — firming up per-method calibration beyond n=5
- Full validation pass on the remaining 13 collectors
- Real weekly-cadence provider_drift run (beyond the same-day proxy)
- Cross-backend certification: extend `capture` to TensorRT-LLM / SGLang (comparison logic is already backend-agnostic)

## Citation

If you use DeltaCert, cite it as:

```bibtex
@software{deltacert2026,
  title  = {DeltaCert: Calibrated Divergence Certification for LLM Serving Systems},
  author = {Shorya},
  year   = {2026},
  url    = {https://pypi.org/project/deltacert/}
}
```

## License

Apache-2.0. See [LICENSE](LICENSE).

## Contact

- Issues and questions: open a GitHub issue
- Full validation data: `validation_results/` in this repo
