# Verdict

> Regression testing for AI agent products: run your real code with side effects contained, grade outcomes, and get an honest verdict on every change.

Verdict is an open-source eval harness for AI agents in Python apps. Every page below is plain Markdown. The whole site in one file: https://mitej23.github.io/llms-full.txt

## Get started

- [Introduction](https://mitej23.github.io/docs.md): Verdict is regression testing for AI agent products, run against your app's real code.
- [Installation](https://mitej23.github.io/docs/installation.md): Install Verdict and the optional extras for SQLAlchemy, a frozen clock, OpenTelemetry and the web UI.
- [Quickstart](https://mitej23.github.io/docs/quickstart.md): Run the refunds example, open the web UI, and catch a seeded regression with verdict compare.

## Core concepts

- [Overview](https://mitej23.github.io/docs/concepts.md): The parts of Verdict in one page, each linking to its own concept page.
- [Cases and worlds](https://mitej23.github.io/docs/concepts/cases-and-worlds.md): A case is a folder with the input, the starting world and the checks for one situation.
- [Boundary fakes](https://mitej23.github.io/docs/concepts/boundary-fakes.md): Every external service is replaced by a fake that records what your app asked it to do.
- [In-process isolation and the network guard](https://mitej23.github.io/docs/concepts/isolation-and-network-guard.md): How Verdict contains a trial inside your app's own Python process, and why a forgotten boundary fails loudly.
- [State adapters](https://mitej23.github.io/docs/concepts/state-adapters.md): A state adapter gives every trial its own database, built from the case's world.
- [Checks](https://mitej23.github.io/docs/concepts/checks.md): Checks grade what a trial did, its effects, final state, output and error, and say why they failed.
- [Trials and pass^k](https://mitej23.github.io/docs/concepts/trials-and-pass-k.md): Each case runs k times because agents are noisy; pass^k means every one of the k trials passed.
- [Code versions and fingerprints](https://mitej23.github.io/docs/concepts/code-versions.md): Every run records a fingerprint of the code it tested, so results always name their code.
- [Compare and the gate](https://mitej23.github.io/docs/concepts/compare-and-the-gate.md): How Verdict decides whether a change helped beyond noise, what it flags, and when the CI gate fails.
- [Grader self-test and regrade](https://mitej23.github.io/docs/concepts/self-test-and-regrade.md): Prove each case's checks are right before any run, and re-score old trials after you change them.
- [Traces and observations](https://mitej23.github.io/docs/concepts/traces-and-observations.md): Verdict keeps each trial's OpenTelemetry spans, turns them into typed steps, and points failed checks at the step behind them.
- [Generating cases](https://mitej23.github.io/docs/concepts/generating-cases.md): verdict generate writes many realistic cases from a spec without letting a model decide what is correct.

## Guides

- [Add Verdict to your app](https://mitej23.github.io/docs/guides/add-verdict-to-your-app.md): Point Verdict at your app's entry point, database and external services, then write a first case.
- [Write checks](https://mitej23.github.io/docs/guides/write-checks.md): Write checks that pin down what your app must do, using effects and state before output.
- [Multi-turn cases](https://mitej23.github.io/docs/guides/multi-turn-cases.md): Test a chat agent over a whole conversation, with a scripted customer or one played by a model.
- [Use the SQLAlchemy adapter](https://mitej23.github.io/docs/guides/sqlalchemy-adapter.md): Give every trial its own SQLite database from your SQLAlchemy models and the case's world.
- [Freeze the clock](https://mitej23.github.io/docs/guides/freeze-the-clock.md): Run a case at a fixed time so dates like "10 days ago" or "next Monday" are reproducible.
- [Instrument with OpenTelemetry](https://mitej23.github.io/docs/guides/opentelemetry.md): Collect each trial's spans for the trace view, step blame and token cost.
- [Write a generator spec](https://mitej23.github.io/docs/guides/write-a-generator-spec.md): Turn your product rules into a generator spec, using the refunds example, and generate gated cases.
- [Compare a candidate](https://mitej23.github.io/docs/guides/compare-a-candidate.md): Run the same cases against another version of your code, from a folder or a git ref, and compare.
- [Gate CI on compare](https://mitej23.github.io/docs/guides/gate-ci.md): Fail a CI job when a change regresses a case, worsens one beyond noise, or raises cost too much.
- [Keep costs down](https://mitej23.github.io/docs/guides/keep-costs-down.md): Spend model calls only where they change a decision, and price your models so cost is visible.

## Walkthroughs

- [Build an eval for your agent, end to end](https://mitej23.github.io/docs/tutorial.md): From an app with no evals to a gated CI comparison, step by step, using the refunds agent as the app.
- [Web UI](https://mitej23.github.io/docs/web-ui.md): A tour of the local web app for one project, from cases and runs to a single trial's trace and a code diff.

## Project

- [Contributing](https://mitej23.github.io/docs/project/contributing.md): Set up the repository, run the tests, and follow the rules that keep Verdict trustworthy.
- [Release policy](https://mitej23.github.io/docs/project/release-policy.md): What is stable in Verdict 0.x, how file formats are versioned, and what a release must include.
- [AI agents](https://mitej23.github.io/docs/project/ai-agents.md): Read these docs from an AI agent or coding assistant with llms.txt, llms-full.txt and per-page Markdown.
- [License](https://mitej23.github.io/docs/project/license.md): Verdict is licensed under the Apache License, Version 2.0.

## Reference

- [Reference](https://mitej23.github.io/docs/reference.md): Complete reference for the verdict CLI, every file format, the check kinds and the Python API.
- [CLI](https://mitej23.github.io/docs/reference/cli.md): Every verdict command, flag, default and exit code.
- [verdict.toml](https://mitej23.github.io/docs/reference/verdict-toml.md): Every key in the project config, with its type and default.
- [case.toml](https://mitej23.github.io/docs/reference/case-toml.md): Every key in a case's case.toml, the file that defines one case's input and checks.
- [world.json](https://mitej23.github.io/docs/reference/world-json.md): The starting database rows and scripted service responses a case's trials begin from.
- [oracle.json and known_bad.json](https://mitej23.github.io/docs/reference/oracle-and-known-bad.md): Known-good and known-bad results that prove a case's checks are right.
- [holdout.json](https://mitej23.github.io/docs/reference/holdout-json.md): The list of cases compare reports separately as the holdout split.
- [generator.toml](https://mitej23.github.io/docs/reference/generator-toml.md): Every key in a case generator spec, plus the files verdict generate writes.
- [Generator hooks](https://mitej23.github.io/docs/reference/generator-hooks.md): The functions and schema a generator spec's scenarios.py defines, with their arguments and return values.
- [Check kinds](https://mitej23.github.io/docs/reference/checks.md): Every check kind, its fields, how it matches, and the label it generates.
- [Python API](https://mitej23.github.io/docs/reference/python-api.md): The functions and classes you use from Python, in custom fakes, Python checks, generator hooks and scripts.
- [Store schema](https://mitej23.github.io/docs/reference/store-schema.md): The SQLite tables Verdict saves runs, per-case summaries, trials and spans in.

## Changelog

- [Changelog](https://mitej23.github.io/docs/changelog.md): What changed in each Verdict release.
