Metadata-Version: 2.5
Name: auditkit
Version: 1.1.0
Summary: Evaluate any model on any dataset and any task — standalone, no platform required
Project-URL: Homepage, https://github.com/Lexsi-Labs/AuditKIT
Project-URL: Repository, https://github.com/Lexsi-Labs/AuditKIT
Author: Lexsi Labs
License: # Lexsi Labs Source Available License (LSAL), Version 1.2 (AuditKit)
        
        ## Preamble
        
        This Source Available License governs use of the software known as **AuditKit**, together with its documentation, CLI, and example pipelines (collectively, the "Licensed Work"), developed and owned by **Lexsi Labs (Lithasa Technologies Pvt. Ltd.)** ("Licensor").
        
        This is **not** an open-source license as defined by the [Open Source Initiative (OSI)](https://opensource.org/). It grants **academic research and teaching** MIT-like permissions: free use, modification, and redistribution with the notice intact. Organizations must acknowledge their use or obtain permission (Section 1A). It also bars commercial exploitation.
        
        ---
        
        ## 1. Grant of Rights
        
        Subject to the terms of this License, permission is hereby granted, free of charge, to any person obtaining a copy of the Licensed Work, to use, copy, modify, merge, publish, and redistribute the Licensed Work and derivative works thereof, **for Noncommercial Purposes only**, provided that the above copyright notice and this License are included in full in all copies or substantial portions of the Licensed Work.
        
        For academic research and teaching, this grant is MIT-like: you may use, modify, and redistribute the Licensed Work, provided the copyright notice and this License travel with every copy.
        
        **Noncommercial Purposes** means:
        
        * personal use for research, experimentation, private study, or hobby projects;
        * academic and scholarly research, teaching, and publication (including use in papers, theses, benchmarks, and reproducibility artifacts).
        
        Research and teaching by academics and their groups, including work done at a university or public research body, falls under this Section. Section 1A governs use by an organization as such, including a charitable organization, educational institution, public research organization, or government body acting for its own operations.
        
        ## 1A. Use by Organizations
        
        Before any organization, including a charitable organization, educational institution, public research organization, or government body, uses the Licensed Work for **internal evaluation, red-teaming, benchmarking, auditing, or any use in connection with developing, evaluating, or safeguarding its own models, products, or services**, whether or not it charges a fee or provides the Licensed Work to third parties, it must either:
        
        * obtain Licensor's **prior written permission**; or
        * provide Licensor a written **acknowledgement of use**, identifying the organization, the intended use, and an undertaking to attribute the Licensed Work, which Licensor may accept in place of negotiated permission.
        
        Both go to **support@lexsi.ai**.
        
        Licensor grants permission at its discretion. The written permission sets out the terms, which may include sharing with Licensor the evaluation results the Licensed Work produces, in AuditKit's standard report format, excluding model weights, training data, prompt text, and model generations. Licensor uses shared results for research and for improving the Licensed Work, and handles them under the terms stated in the written permission.
        
        Permission under this Section does not authorize any use described in Section 2.
        
        ## 2. Commercial Restriction
        
        Without a separate **commercial license** from Lexsi Labs, you may **not** Sell the Licensed Work. "**Sell**" means practicing any right granted to you under this License to provide to third parties, for a fee or other consideration, a product or service whose value derives, entirely or substantially, from the functionality of the Licensed Work, including:
        
        * offering the Licensed Work, or a derivative of it, as a commercial product, paid service, SaaS, hosted, or API offering;
        * embedding the Licensed Work in proprietary or revenue-generating software;
        * paid consulting or support whose substance is the Licensed Work.
        
        If you redistribute a modified version, you must mark it as modified and must not present it as the original. You may also not **re-license, rebrand, or redistribute** the Licensed Work under different terms, nor use **Lexsi Labs**, **AuditKit**, or related trademarks, logos, or branding except to identify unmodified, licensed copies.
        
        ## 3. Patents
        
        NO EXPRESS OR IMPLIED LICENSES TO ANY PARTY'S PATENT RIGHTS ARE GRANTED BY THIS LICENSE. Any patent rights of Licensor relating to the Licensed Work are reserved and may be licensed only under a separate written agreement with Licensor.
        
        ## 3A. Third-Party Components
        
        The Licensed Work evaluates and audits third-party base models and draws on third-party benchmarks and datasets. Each carries its own license, and this License grants no rights under any of them. You must comply with those terms in addition to this License, including any restrictions a base-model or dataset license places on your use of it.
        
        ## 4. Ownership
        
        All rights, title, and interest in and to the Licensed Work remain with **Lithasa Technologies Pvt. Ltd.** Except as expressly stated in Sections 1 and 1A, nothing in this License transfers ownership or any other rights to the Licensee.
        
        ## 5. Contributions
        
        If you submit modifications, pull requests, or patches ("Contributions") to Lexsi Labs, you grant Lexsi Labs a perpetual, worldwide, royalty-free right to use, modify, distribute, and license your Contributions under any terms, including commercial ones, and you represent that you have the right to make such Contributions.
        
        ## 6. Warranty Disclaimer and Responsibility for Use
        
        THE LICENSED WORK IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE, AND NONINFRINGEMENT. IN NO EVENT SHALL THE LICENSOR OR CONTRIBUTORS BE LIABLE FOR ANY CLAIM, DAMAGES, OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT, OR OTHERWISE, ARISING FROM, OUT OF, OR IN CONNECTION WITH THE LICENSED WORK OR THE USE OR OTHER DEALINGS IN THE LICENSED WORK.
        
        **You alone are responsible for your use of the Licensed Work.** This includes every evaluation, audit, output, or system you produce with it, and your compliance with applicable law. This responsibility applies to academic, personal, organizational, and commercial use alike. Licensor has no duty to monitor your use and bears no liability for misuse of the Licensed Work by you or by any third party who obtained it through you.
        
        The Licensed Work is a research and auditing tool. Any measurements or findings it reports hold only under the conditions you configure. They do not certify that any model is safe to deploy; that decision rests with you.
        
        If you use the Licensed Work under Section 1A or under a commercial license, you will, to the extent permitted by applicable law, indemnify and hold harmless Licensor and its contributors against any claim, loss, or expense, including legal fees, arising from your use or misuse of the Licensed Work or from any breach of this License by you. This indemnity does not apply to personal or academic use under Section 1; the two preceding paragraphs govern that use.
        
        ## 7. Termination
        
        This License terminates without notice if you breach any of its terms. On termination, you must stop using the Licensed Work and destroy every copy in your possession. Licenses of parties who received the Licensed Work from you remain in force provided they remain in compliance.
        
        Sections 3, 3A, 4, 5, 6, and 8 survive termination.
        
        ## 8. Governing Law and General Terms
        
        This License shall be governed by and construed in accordance with the **laws of India**, without regard to its conflict of law principles. If a court holds any provision of this License unenforceable, the remaining provisions stay in effect. Licensor may publish revised versions of this License; a copy of the Licensed Work remains under the version it was received under unless you accept a later version.
        
        ## 9. Contact
        
        For acknowledgements and permission requests under Section 1A, and for commercial use, partnership, or redistribution rights under Section 2, contact:
        **support@lexsi.ai** · **https://lexsi.ai**
        
        ## 10. Notice
        
        **AuditKit** © 2026 **Lithasa Technologies Pvt. Ltd.**
        Licensed under the **Lexsi Labs Source Available License (LSAL) v1.2**.
        **Academic research and teaching are free on MIT-like terms. Use by any organization requires written acknowledgement or permission (Section 1A). Commercial use requires a commercial license (Section 2).**
License-File: LICENSE.md
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Topic :: Software Development :: Testing
Requires-Python: >=3.10
Provides-Extra: all
Requires-Dist: accelerate>=1.0; extra == 'all'
Requires-Dist: anthropic; extra == 'all'
Requires-Dist: bert-score; extra == 'all'
Requires-Dist: datasets; extra == 'all'
Requires-Dist: flask; extra == 'all'
Requires-Dist: gitpython; extra == 'all'
Requires-Dist: lexsi-sdk; extra == 'all'
Requires-Dist: litellm; extra == 'all'
Requires-Dist: lm-eval; extra == 'all'
Requires-Dist: mlcroissant; extra == 'all'
Requires-Dist: mlflow; extra == 'all'
Requires-Dist: openai; extra == 'all'
Requires-Dist: peft; extra == 'all'
Requires-Dist: pillow; extra == 'all'
Requires-Dist: pytest; extra == 'all'
Requires-Dist: pytest-asyncio; extra == 'all'
Requires-Dist: pytest-mock; extra == 'all'
Requires-Dist: requests; extra == 'all'
Requires-Dist: torch>=2.4; extra == 'all'
Requires-Dist: torchmetrics; extra == 'all'
Requires-Dist: torchvision; extra == 'all'
Requires-Dist: transformers<6,>=5.15; extra == 'all'
Requires-Dist: vllm>=0.30; extra == 'all'
Requires-Dist: werkzeug; extra == 'all'
Provides-Extra: anthropic
Requires-Dist: anthropic; extra == 'anthropic'
Provides-Extra: bert-score
Requires-Dist: bert-score; extra == 'bert-score'
Requires-Dist: torch>=2.4; extra == 'bert-score'
Provides-Extra: dev
Requires-Dist: pytest; extra == 'dev'
Requires-Dist: pytest-asyncio; extra == 'dev'
Requires-Dist: pytest-mock; extra == 'dev'
Provides-Extra: interop
Requires-Dist: datasets; extra == 'interop'
Requires-Dist: gitpython; extra == 'interop'
Requires-Dist: mlcroissant; extra == 'interop'
Provides-Extra: lexsi-sdk
Requires-Dist: lexsi-sdk; extra == 'lexsi-sdk'
Provides-Extra: litellm
Requires-Dist: litellm; extra == 'litellm'
Provides-Extra: lmeval
Requires-Dist: accelerate>=1.0; extra == 'lmeval'
Requires-Dist: lm-eval; extra == 'lmeval'
Provides-Extra: mlflow
Requires-Dist: mlflow; extra == 'mlflow'
Provides-Extra: openai
Requires-Dist: openai; extra == 'openai'
Provides-Extra: relay
Requires-Dist: flask; extra == 'relay'
Requires-Dist: requests; extra == 'relay'
Requires-Dist: werkzeug; extra == 'relay'
Provides-Extra: requests
Requires-Dist: requests; extra == 'requests'
Provides-Extra: sglang
Requires-Dist: sglang<0.6,>=0.5.18; extra == 'sglang'
Provides-Extra: transformers
Requires-Dist: peft; extra == 'transformers'
Requires-Dist: torch>=2.4; extra == 'transformers'
Requires-Dist: transformers<6,>=5.15; extra == 'transformers'
Provides-Extra: vision
Requires-Dist: pillow; extra == 'vision'
Requires-Dist: torchmetrics; extra == 'vision'
Requires-Dist: torchvision; extra == 'vision'
Provides-Extra: vllm
Requires-Dist: vllm>=0.30; extra == 'vllm'
Description-Content-Type: text/markdown

<p align="center">
  <a href="https://github.com/Lexsi-Labs/AuditKIT">
    <picture>
      <source media="(prefers-color-scheme: dark)" srcset="docs/assets/auditkit-logo-white.png">
      <img src="docs/assets/auditkit-logo-black.png" alt="AuditKit" width="480">
    </picture>
  </a>
</p>

<p align="center">
  <b>Evaluate any model on any dataset and any task.</b><br>
  One library for benchmark, judge, code, red-team, security, and performance<br>
  evaluation — zero required deps.
</p>

<p align="center">
  <a href="https://pypi.org/project/auditkit/"><img src="https://img.shields.io/badge/pypi-v1.2.0-0a8868" alt="PyPI v1.2.0"></a>
  <a href="https://www.python.org/"><img src="https://img.shields.io/badge/python-3.10%2B-blue" alt="Python 3.10+"></a>
  <a href="https://github.com/Lexsi-Labs/AuditKIT/blob/main/LICENSE.md"><img src="https://img.shields.io/badge/license-LSAL--1.2-lightgrey" alt="License: LSAL-1.2 (source-available, noncommercial)"></a>
  <a href="https://auditkit.lexsi.ai/"><img src="https://img.shields.io/badge/docs-auditkit.lexsi.ai-4c6ef5" alt="Documentation"></a>
  <a href="https://github.com/Lexsi-Labs/AuditKIT/actions"><img src="https://img.shields.io/badge/tests-2.8k%20passing-brightgreen" alt="Tests: 2846 passing, 141 skipped"></a>
</p>

<p align="center">
  <a href="https://auditkit.lexsi.ai/">Documentation</a> ·
  <a href="#quickstart">Quickstart</a> ·
  <a href="#examples">Examples</a> ·
  <a href="docs/community/contributing.md">Contributing</a>
</p>

---

## TL;DR

| | |
|---|---|
| **What** | A unified Python library that evaluates any AI model on any dataset across any technique. |
| **Architecture** | 5-stage spine: Dataset → Adapter → Model → Metrics → RunResult. `ak.evaluate()` orchestrates it all. |
| **Techniques** | Benchmark, LLM-as-judge (GEval), code checks, RAG, hallucination, embedding similarity, toxicity/bias, pairwise/preference, red-teaming, security, performance, agent/tool use |
| **Models** | 10 backends: OpenAI, Anthropic, HuggingFace, Lexsi, vLLM, LiteLLM, API, Groq, OpenRouter, Agent (HTTP endpoint) — all resolved via `AutoModel.resolve()`. Any `list[str] → list[str]` callable also works. |
| **Zero deps** | Core runs on stdlib. Backends and heavy metrics are optional extras (`pip install auditkit[openai]`). |
| **Fingerprints** | Every run gets a stable sha256 — results are cacheable, comparable, and reproducible by construction. |
| **Status** | ~3,000 tests, v1.2.0, LSAL-1.2 license (source-available, noncommercial). 11 metric families, 6 CLI subcommands, YAML config, MKDocs site. |

---

## Why AuditKIT

Evaluating AI models is fragmented. Academic benchmarks (MMLU, GSM8K) use one tool. LLM-as-judge evaluations use another. Red-teaming and performance profiling each have their own frameworks. There is no single library that does all of them with a consistent API, zero required dependencies, and a provenance-first data model.

AuditKIT is that library. It provides a unified evaluation spine that supports every technique, so you can compare results across benchmarks, judge evaluations, red-team probes, and performance profiles — all from a single `ak.evaluate()` call.

## Key features

- **Metric families.** Benchmark (exact match, F1, BLEU, ROUGE, ChrF), LLM-as-judge (GEval, rubric items), code/deterministic (contains, regex, JSON validation), RAG (lexical groundedness, context overlap/coverage), embedding similarity, hallucination detection, toxicity/bias, pairwise/preference (win rate, Elo, preference accuracy), security (DEFCON grade), performance (latency, throughput).
- **Zero required deps.** Core runs on the Python standard library alone. Heavy backends (BERTScore, vLLM, LiteLLM, OpenAI, Anthropic, HuggingFace) are optional extras.
- **Any model backend.** `echo` for testing, `openai:`, `anthropic:`, `hf:`, `lexsi:`, `vllm:`, `litellm:` (which also reaches Ollama, e.g. `litellm:ollama/llama3.1`), `api:`, `groq:`, `openrouter:`, `agent:` (a deployed agent over HTTP), or any callable. Auto-resolved via `AutoModel.resolve()`.
- **Many datasets, one model.** `evaluate_many()` runs one model across several datasets in a single call, returning one `RunResult` per dataset.
- **Red teaming.** Built-in adversarial probes (prompt injection, jailbreak, encoding, over-refusal) and detectors (keyword, refusal, injection success, system prompt leak) via `RedTeamRunner`.
- **Agent and RAG evaluation.** Score tool calls, parallel/dependency behavior and retrieval on any model. The `auditkit.agent_eval` layer evaluates a whole task episode: a verified outcome (final state, artifact, answer, or a custom predicate) as the headline, with tool, trace and judge scores as diagnostics. It reads recorded runs (AgentTune, OpenAI messages, a `Sample` trace) or calls a deployed `agent:` endpoint, and works with any framework that emits OpenAI messages or serves over HTTP.
- **Experiment tracking.** Named experiments with `ExperimentDB`, MLflow logging, cross-run comparison with bootstrap significance tests.
- **Model comparison.** `compare_models()` runs the same dataset against multiple models and produces side-by-side results with pairwise significance.
- **CLI with YAML config.** `auditkit eval`, `init`, `list`, `redteam`, `compare`, `agent` subcommands. Define evaluations in YAML files with `prompts:`, model config, tags, and split strategies.
- **Provenance by default.** Every run writes a `RunResult` with a stable fingerprint (sha256 over model+seed+tasks+config). Results are cacheable and comparable by fingerprint.

## How it works

```mermaid
flowchart LR
    A["Dataset<br/>Samples · Scenario · CSV"] --> B["Adapter<br/>generation · chat · instruction"]
    B --> C["Model<br/>echo · openai · hf · vllm"]
    C --> D{"Metrics"}
    D -->|per sample| E["Score + Prediction"]
    E --> F["Aggregate<br/>Stats · Headline"]
    F --> G["RunResult<br/>fingerprint · cache"]
```

## Install

```bash
pip install auditkit                     # core (zero deps)
pip install "auditkit[openai]"           # OpenAI backend
pip install "auditkit[all]"              # all backends + heavy metrics
```

Requires Python 3.10+ on Linux, macOS, or Windows.

<details>
<summary>More install options</summary>

```bash
pip install "auditkit[anthropic]"            # Anthropic backend
pip install "auditkit[litellm]"              # LiteLLM (Ollama, etc.)
pip install "auditkit[vllm]"                 # vLLM backend
pip install "auditkit[transformers]"         # HuggingFace + hallucination + embedding similarity
pip install "auditkit[bert-score]"           # BERTScore
pip install "auditkit[mlflow]"               # MLflow experiment tracking

# From source
pip install "auditkit @ git+https://github.com/Lexsi-Labs/AuditKIT.git"
```

</details>

## Quickstart

```python
import auditkit as ak

samples = [
    ak.Sample(input="What is 2+2?", target="4"),
    ak.Sample(input="What is 3+3?", target="6"),
]

result = ak.evaluate(samples, model=lambda prompts: prompts)
print(result.summary())
```

**From the CLI:**

```bash
auditkit eval --model hf:gpt2 <<< "What is 2+2?"
```

**With a YAML config:**

```yaml
model: openai:gpt-4o-mini
temperature: 0.0
prompts:
  - "What is the capital of France?"
output: results.json
```

```bash
auditkit eval --config auditkit.yaml
```

### Cohere / Aya models

Aya Expanse, Tiny Aya, Aya Vision and North Micro Vision run on the `hf:` backend
(`pip install "auditkit[transformers,vision]"`, transformers 5.15 or newer).
Score a model before and after fine-tuning:

```python
import auditkit as ak

samples = [ak.Sample(input="Translate to French: good morning", target="bonjour")]
metrics = ["exact_match", "f1_score"]

base = ak.evaluate(samples, "hf:CohereLabs/aya-expanse-8b", metrics)
tuned = ak.evaluate(samples, "hf:<you>/aya-expanse-8b-tuned", metrics)
print(ak.compare(base, tuned).summary())
```

For the vision models (`CohereLabs/aya-vision-8b`, `CohereLabs/aya-vision-32b`,
`CohereLabs/North-Micro-Vision-Instruct`), attach images to the sample; on the `hf:`
backend they go through the model's processor:

```python
from PIL import Image

sample = ak.Sample(input="What does this chart show?", target="sales by month",
                   images=[Image.open("chart.png")])
result = ak.evaluate([sample], "hf:CohereLabs/aya-vision-8b", ["f1_score"])
```

**Images travel on `hf:` and on `api:` in chat mode**, whichever host `api:` points
at: a self-hosted vLLM or SGLang server, or a hosted provider. `api:` raw-prompt mode raises
rather than drop them. The offline `vllm:` backend doesn't read `Sample.images` yet and drops
them **without an error**: for a vision eval on vLLM, `vllm serve` the model and use `api:`.

Aya Vision **does** work text-only on `hf:` — measured, vision tower unused, 5/5 on a
five-language probe, and identical to Aya Expanse 8B on the same probe. For text
work prefer Expanse: same quality or better, and it runs natively on vLLM and SGLang,
where Aya Vision needs each one's Transformers backend.

The Cohere models run on `hf:`, `vllm:` and SGLang (all nine measured on each), but not all natively. Aya Vision has no native server class (vLLM
dropped it after 0.24.0, SGLang never had it), so both servers run it through their
Transformers backends: automatically on vLLM, with `compat.py --patch-sglang` on SGLang. North Micro Vision is native on vLLM; on SGLang
it needs transformers >= 5.15 in the SGLang env plus `--patch-sglang`. Tiny Aya and Aya Expanse
are native on both. The full
model-by-runtime table is in [docs/model_backends.md](docs/model_backends.md).

Tiny Aya (`CohereLabs/tiny-aya-global` and siblings `fire`, `water`, `earth`),
Aya Expanse and Aya Vision are gated
on the Hub: accept the licence on the model page, then set `HF_TOKEN` (or pass `token=` to
`auditkit.model.hf_gen.HFGenModel`). North Micro Vision is not gated.

## What it evaluates

| Task | Technique | Example metrics |
|------|-----------|-----------------|
| Benchmark text | MCQ, generation | `ExactMatch`, `Bleu`, `F1Score` |
| LLM output quality | LLM-as-judge | `GEval`, `RubricItem` |
| RAG pipelines | RAG | `LexicalGroundedness`, `ContextOverlap` |
| Safety | Red-team | `KeywordDetector`, `DefconGrade`, probes |
| Performance | Latency/throughput | `LatencyStats`, `Throughput` |

## Output

Every `evaluate()` returns a `RunResult`:

| Field | Contents |
|-------|----------|
| `headline` | Per-metric averages |
| `predictions` | Per-sample scores with input, output, expected |
| `stats` | Full statistics (mean, std, min, max, count) |
| `fingerprint` | Stable sha256 for caching and comparison |
| `config` | The `RunConfig` used (seed, temperature, etc.) |
| `errors` | Any errors encountered during the run |

## Examples

Colab-ready notebooks with real models and real datasets — see
[`examples/README.md`](examples/README.md) for the full, current list.

| Example | File |
|---------|------|
| Full pipeline: adapter, 5 metrics, LLM judge, LLM annotator | [`01_full_evaluation_pipeline.ipynb`](examples/01_full_evaluation_pipeline.ipynb) |
| Generation across two real HF model families, compared | [`02_generation_across_hf_families.ipynb`](examples/02_generation_across_hf_families.ipynb) |
| Custom annotators (regex, LLM-backed, fully custom) | [`03_custom_annotators.ipynb`](examples/03_custom_annotators.ipynb) |
| Metrics deep dive: built-in, custom, LLM-as-judge, RAG | [`04_metrics_deep_dive.ipynb`](examples/04_metrics_deep_dive.ipynb) |
| Every data type/task kind: generative, MCQ, RAG, precomputed, chat | [`05_data_types.ipynb`](examples/05_data_types.ipynb) |
| Model comparison deep dive: `compare_models()` + `RunComparison` | [`06_model_comparison.ipynb`](examples/06_model_comparison.ipynb) |
| Annotators across 4 real model families, then compared together | [`07_annotators_across_models.ipynb`](examples/07_annotators_across_models.ipynb) |
| `CompareResult` deep dive: what `compare_models()`'s native per-model support still can't express (different scorers per model), full method surface | [`08_compare_result_deep_dive.ipynb`](examples/08_compare_result_deep_dive.ipynb) |
| `GuardJudge` implementation check: all 5 profiles, real and offline | [`09_guard_judge_implementation_check.ipynb`](examples/09_guard_judge_implementation_check.ipynb) |
| `GuardJudge` with real, flagship guard models | [`10_guard_judge_legit_models.ipynb`](examples/10_guard_judge_legit_models.ipynb) |
| `EncoderJudge`: the base mechanism plus its two prebuilt subclasses | [`11_encoder_judge_prebuilts.ipynb`](examples/11_encoder_judge_prebuilts.ipynb) |
| Performance metrics on a real, large-scale evaluation | [`12_performance_metrics_demo.ipynb`](examples/12_performance_metrics_demo.ipynb) |
| SGLang: environment checks, server mode via `api:`, tool calls and parallel tool calls | [`13_sglang_compatibility.ipynb`](examples/13_sglang_compatibility.ipynb) |
| Agent and RAG evals: tool calls, parallel tool calls, deployed agent endpoints, AgentTune data, retrieval and judged RAG metrics | [`14_agent_and_rag_evals.ipynb`](examples/14_agent_and_rag_evals.ipynb) |
| End-to-end agent evaluation, fully offline (`.py` script): recorded episodes, a verified state outcome, and a fabricated answer that does not pass a state case | [`agent_eval_offline.py`](examples/agent_eval_offline.py) |

## Applications

Real-world scenarios answered end to end, not feature tours — see
[`examples/applications/README.md`](examples/applications/README.md). Only `01` uses the
`lm-evaluation-harness` integration (`ak.run_lmeval()`, needs
`auditkit[lmeval]`), as an authoritative cross-check alongside the native
evaluation; none of the others do.

| Application | File |
|-------------|------|
| How much does pruning severity (20%/40%/60%) degrade a model? Real BoolQ, real annotator, ship/no-ship verdicts, cross-checked via `ak.run_lmeval()` | [`01_application_pruned_llama_boolq.ipynb`](examples/applications/01_application_pruned_llama_boolq.ipynb) |
| Healthcare: clinical QA correctness vs. grounding (PubMedQA) | [`02_application_healthcare_pubmedqa.ipynb`](examples/applications/02_application_healthcare_pubmedqa.ipynb) |
| Finance: QA grounded in real SEC 10-K filings | [`03_application_finance_10k_qa.ipynb`](examples/applications/03_application_finance_10k_qa.ipynb) |
| E-commerce: review-sentiment triage at scale | [`04_application_ecommerce_review_triage.ipynb`](examples/applications/04_application_ecommerce_review_triage.ipynb) |
| Education: auto-graded tutoring, correctness vs. explanation | [`05_application_education_arc_tutor.ipynb`](examples/applications/05_application_education_arc_tutor.ipynb) |
| Enterprise search: internal knowledge assistant (real retrieval + RAG grounding) | [`06_application_enterprise_search_rag.ipynb`](examples/applications/06_application_enterprise_search_rag.ipynb) |
| LLM-as-judge via a real BERT NLI classifier, not a generative model | [`07_application_bert_nli_judge.ipynb`](examples/applications/07_application_bert_nli_judge.ipynb) |

## Repository map

| Directory | Description | README |
|-----------|-------------|--------|
| [`src/auditkit/`](src/auditkit/) | Core library — spine, runner, metrics, models, CLI | [README](src/auditkit/README.md) |
| [`src/auditkit/metrics/`](src/auditkit/metrics/) | Metric families | [README](src/auditkit/metrics/README.md) |
| [`src/auditkit/model/`](src/auditkit/model/) | Model backends (echo, openai, hf, vllm, etc.) | [README](src/auditkit/model/README.md) |
| [`src/auditkit/redteam/`](src/auditkit/redteam/) | Red team probes and detectors | [README](src/auditkit/redteam/README.md) |
| [`src/auditkit/scenarios/`](src/auditkit/scenarios/) | Built-in benchmark datasets | [README](src/auditkit/scenarios/README.md) |
| [`tests/`](tests/) | Test suite (~3,000 tests) | [README](tests/README.md) |
| [`examples/`](examples/) | Colab-ready example notebooks | [README](examples/README.md) |
| [`examples/applications/`](examples/applications/) | Real-world application notebooks | [README](examples/applications/README.md) |
| [`docs/`](docs/) | MkDocs source for [auditkit.lexsi.ai](https://auditkit.lexsi.ai/) | [index](docs/index.md) |

---

## Developed by Lexsi Labs

<p align="center">
  Created by the team at <strong>Lexsi Labs</strong>, AuditKIT provides a unified evaluation spine for AI models across every technique and modality.
</p>

Shreeyans Arora · Utsav Avaiya · Zera Lyngkhoi · Ram Mohan Rao Kadiyala · Hem Gosalia · Vinay Kumar Sankarapu · Pratinav Seth

---

## Citation

If you use AuditKIT in research, please cite it:

```bibtex
@software{auditkit2026,
  title  = {AuditKIT: one library to evaluate any model on any dataset and any task},
  author = {Arora, Shreeyans and Avaiya, Utsav and Lyngkhoi, Zera and
            Kadiyala, Ram Mohan Rao and Gosalia, Hem and Sankarapu, Vinay Kumar and
            Seth, Pratinav},
  year   = {2026},
  version = {1.2.0},
  url    = {https://github.com/Lexsi-Labs/AuditKIT}
}
```

`CITATION.cff` carries the same author list for GitHub's citation widget.

---

## License

This project is released under the [Lexsi Labs Source Available License (LSAL) v1.2](LICENSE.md) — free for academic research and teaching on MIT-like terms; use by any organization requires written acknowledgement or permission (Section 1A); a separate commercial license is required to sell it or embed it in a paid product.

---

## Join Community / Contribute

- Issues and discussions are welcomed on the [GitHub issue tracker](https://github.com/Lexsi-Labs/AuditKIT/issues).
- See [Contributing](docs/community/contributing.md) for the development setup, code standards and the pull request process.
