# Introduction

> Verdict is regression testing for AI agent products, run against your app's real code.

Source: https://mitej23.github.io/docs



Verdict runs your app's real code path against a recorded world, contains every side effect at the edges, grades what your app did, and tells you whether a code change made things better or worse beyond noise.

It is for teams whose agents act in business systems (calendars, CRMs, payments, tickets), where a wrong action costs money or trust. You point Verdict at the function production calls, describe each situation as a folder of plain files, and run it like a test suite.

```mermaid
flowchart LR
  C[&#x22;Case<br/>case.toml + world.json&#x22;] --> E[&#x22;Your app's entry point<br/>the real code path&#x22;]
  E --> B[&#x22;Boundary fakes<br/>record every call&#x22;]
  E --> S[&#x22;State adapter<br/>fresh database per trial&#x22;]
  B --> K[&#x22;Checks<br/>effects, state, output&#x22;]
  S --> K
  E --> K
  K --> V[&#x22;Verdict<br/>per-case compare and gate&#x22;]
```

1. A **case** gives the input, the starting data (the **world**) and the **checks**.
2. Verdict calls your **entry point** with that input, several times (**trials**).
3. Each trial gets its own database, **fakes** for every external service, a blocked network and, optionally, a frozen clock.
4. **Checks** grade what happened: the calls made to services, the final database rows, the output.
5. **Compare** two runs (a baseline and a candidate) case by case and get a verdict and a CI exit code.

## Why Verdict [#why-verdict]

* **Your code, our edges.** Nothing in your app is mocked except its boundaries. Verdict never edits your app.
* **Outcomes over wording.** "Refunded $24.00 once, marked the order refunded, emailed the customer" matters more than how the reply reads.
* **Honest statistics.** The case is the unit. Verdict measures the noise floor, tests each case separately with a correction for multiple comparisons, and says "no difference beyond noise" when that's the truth.
* **Fail loud.** An outbound call that no fake handles raises an error instead of reaching a real service.

## Where to start [#where-to-start]

<Cards>
  <Card title="Installation" href="/docs/installation">
    Install the package and the extras your app needs.
  </Card>

  <Card title="Quickstart" href="/docs/quickstart">
    Run the refunds example, catch a seeded regression, and gate on it.
  </Card>

  <Card title="Core concepts" href="/docs/concepts">
    Cases, fakes, isolation, checks, trials and the statistics behind a verdict.
  </Card>

  <Card title="Tutorial" href="/docs/tutorial">
    Build an eval for your own agent, end to end.
  </Card>
</Cards>

## Status [#status]

Verdict is at 0.1.0 and early. The file formats are versioned (`schema = 1`), and the Python API may still change. It runs Python apps in-process; a process mode for other languages is on the roadmap. See the [changelog](/docs/changelog) and the [release policy](/docs/project/release-policy).


---

# Installation

> Install Verdict and the optional extras for SQLAlchemy, a frozen clock, OpenTelemetry and the web UI.

Source: https://mitej23.github.io/docs/installation



Verdict is a Python package, `verdict-evals`, with no required dependencies. It needs Python 3.11 or later.

<Callout type="warn" title="Not on PyPI yet">
  Verdict 0.1.0 has not been published to PyPI. Until it is, install it from the Git repository. The package name and extras below are the ones it will be published under.
</Callout>

## Install the package [#install-the-package]

Add it to your app's environment, with the extras your app needs.

<Tabs items="['uv', 'pip', 'From a checkout']">
  <Tab value="uv">
    ```bash
    uv add "verdict-evals[sqlalchemy,clock,otel] @ git+https://github.com/mitej23/verdict"
    ```
  </Tab>

  <Tab value="pip">
    ```bash
    pip install "verdict-evals[sqlalchemy,clock,otel] @ git+https://github.com/mitej23/verdict"
    ```
  </Tab>

  <Tab value="From a checkout">
    ```bash
    git clone https://github.com/mitej23/verdict
    cd verdict
    uv sync --extra dev        # every extra, plus pytest
    uv run verdict --version
    ```
  </Tab>
</Tabs>

Check that the CLI is on your path:

```bash
verdict --version
# verdict 0.1.0
```

## Extras [#extras]

Verdict's core uses only the standard library. Integrations are optional extras.

| Extra         | Installs                                                   | You need it when                                                                                                                                                                    |
| ------------- | ---------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `sqlalchemy`  | `sqlalchemy>=2.0`                                          | Your app uses SQLAlchemy and you set `[state] adapter = "sqlalchemy"`. See [State adapters](/docs/concepts/state-adapters).                                                         |
| `clock`       | `time-machine>=2.14`                                       | A case sets `clock` to freeze time. See [Freeze the clock](/docs/guides/freeze-the-clock).                                                                                          |
| `otel`        | `opentelemetry-sdk>=1.24`                                  | You want each trial's spans, the trace view, and token cost. See [Instrument with OpenTelemetry](/docs/guides/opentelemetry).                                                       |
| `otel-export` | `otel` plus `opentelemetry-exporter-otlp-proto-http>=1.24` | You also send traces to Logfire, Langfuse or another OTLP backend with `[trace.export]`. See [Send traces to your backend](/docs/guides/opentelemetry#send-traces-to-your-backend). |
| `web`         | `fastapi>=0.110`, `uvicorn>=0.29`                          | You run `verdict serve`. See [Web UI](/docs/web-ui).                                                                                                                                |
| `all`         | every extra above                                          | You want everything, e.g. to try the bundled example.                                                                                                                               |
| `dev`         | all of the above, `pytest`, `httpx`                        | You work on Verdict itself.                                                                                                                                                         |

To try it before wiring up your own app, copy the bundled example and run it:

```bash
pip install "verdict-evals[all] @ git+https://github.com/mitej23/verdict"
verdict init --example refunds demo && cd demo
verdict doctor && verdict run -k 2 && verdict serve --open
```

Verdict tells you which extra is missing when a feature needs one (`verdict doctor` lists them all at once), for example:

```text
[state] adapter = "sqlalchemy" needs the sqlalchemy extra: pip install 'verdict-evals[sqlalchemy]'
```

## The web UI client [#the-web-ui-client]

`verdict serve` serves a React app. Released packages include it built. From a checkout of the repository, build it once; this needs Node 20 or later.

```bash
cd web
npm install
npm run build      # writes into src/verdict/web/dist
```

Without the build, `verdict serve` answers every page with a 503 and these instructions.

## Case generation [#case-generation]

`verdict generate` writes cases through Claude Code's headless mode, so it needs the `claude` CLI on your path and a Claude Code login. `--estimate` and `--dry-run` don't. See [Generating cases](/docs/concepts/generating-cases).

## Next [#next]

Run the bundled example in the [Quickstart](/docs/quickstart).


---

# Quickstart

> Run the refunds example, open the web UI, and catch a seeded regression with verdict compare.

Source: https://mitej23.github.io/docs/quickstart



This page runs the example app that ships with Verdict, compares it with a candidate that has a bug, and shows the gate failing. It takes about five minutes and makes no model calls.

## The example [#the-example]

`examples/refunds` is a small support agent that refunds orders. Its decisions are plain code, so it needs no API keys, but it's wired like a real agent: a SQLAlchemy database, a `PaymentsClient` class and a `send_email` function that call real URLs in production.

<Files>
  <Folder name="examples/refunds">
    <File name="verdict.toml" />

    <Folder name="refunds_app">
      <File name="agent.py" />

      <File name="db.py" />

      <File name="services.py" />
    </Folder>

    <Folder name="candidate">
      <Folder name="refunds_app" />
    </Folder>

    <Folder name="cases">
      <Folder name="refund-within-window" />

      <Folder name="refund-outside-window" />

      <Folder name="large-order-escalates" />

      <Folder name="not-delivered-yet" />

      <Folder name="payments-down" />

      <File name="holdout.json" />
    </Folder>

    <Folder name="generator" />
  </Folder>
</Files>

`candidate/` is a copy of the app with one line changed: the refund window is 7 days instead of 30. It's the regression you'll catch.

<Steps>
  <Step>
    ### Get the code [#get-the-code]

    ```bash
    git clone https://github.com/mitej23/verdict
    cd verdict
    uv sync --extra dev
    cd examples/refunds
    ```
  </Step>

  <Step>
    ### Check the graders [#check-the-graders]

    Before spending anything on runs, check that each case's checks pass a known-good result and fail a known-bad one.

    ```bash
    uv run verdict self-test
    ```

    ```text
    —    large-order-escalates: no oracle.json or known_bad.json
    —    not-delivered-yet: no oracle.json or known_bad.json
    ok   payments-down: known-good 1.0, known-bad 0.25
    —    refund-outside-window: no oracle.json or known_bad.json
    ok   refund-within-window: known-good 1.0, known-bad 0.25
    ```
  </Step>

  <Step>
    ### Run the baseline [#run-the-baseline]

    Run every case three times and label the run.

    ```bash
    uv run verdict run -k 3 --label baseline
    ```

    ```text
    → large-order-escalates  (3 trials)
      ✓ ✓ ✓   3/3 passed
    → not-delivered-yet  (3 trials)
      ✓ ✓ ✓   3/3 passed
    → payments-down  (3 trials)
      ✓ ✓ ✓   3/3 passed
    → refund-outside-window  (3 trials)
      ✓ ✓ ✓   3/3 passed
    → refund-within-window  (3 trials)
      ✓ ✓ ✓   3/3 passed

    Run 20261008-191805-nogit-ba44 saved to .../examples/refunds/.verdict/verdict.db
    ```

    Each trial got a fresh in-memory database built from the case's `world.json`, a recording fake for payments and email, and a clock frozen at 2026-10-08 10:00 UTC. Nothing left the process.
  </Step>

  <Step>
    ### Run the candidate [#run-the-candidate]

    `--app-root` runs the same cases against another checkout of the code.

    ```bash
    uv run verdict run -k 3 --app-root candidate --label candidate
    ```

    ```text
    → refund-within-window  (3 trials)
      ✕ ✕ ✕   0/3 passed
    ```

    The other four cases still pass.
  </Step>

  <Step>
    ### Compare [#compare]

    List the runs to get their ids, then compare baseline with candidate.

    ```bash
    uv run verdict runs
    ```

    ```text
    20261008-191805-nogit-8dd8           finished      12/15   code 4a5ffdf5d401  candidate
    20261008-191805-nogit-ba44           finished      15/15   code 8cfd65598bb3  baseline
    ```

    ```bash
    uv run verdict compare 20261008-191805-nogit-ba44 20261008-191805-nogit-8dd8 --gate
    echo $?
    ```

    ```text
    Baseline  20261008-191805-nogit-ba44
    Candidate 20261008-191805-nogit-8dd8

      all       5 cases   pass 100% → 80%   change -20% (95% CI -50% to +14%)   all trials pass 5 → 4
      dev       4 cases   pass 100% → 75%   change -25% (95% CI -60% to +15%)   all trials pass 4 → 3
      holdout   1 cases   pass 100% → 100%   change +0% (95% CI -46% to +47%)   all trials pass 1 → 1
      cost     $0.0000 → $0.0000 per trial
      time     0.0s → 0.0s per trial

      Cases that changed:
        regressed refund-within-window                     3/3 → 0/3  (dev, p=0.1, adjusted 0.5)

    Verdict: fails the gate
      - 1 case(s) regressed: refund-within-window
    1
    ```

    Read it like this:

    * The overall interval crosses zero, so the average alone wouldn't block the change.
    * `refund-within-window` passed every baseline trial and failed every candidate trial, so it's flagged **regressed**, and a regression fails the gate.
    * `--gate` turns that into exit code 1 for CI.
    * `payments-down` is listed in `holdout.json`, so it's reported on its own line.
  </Step>

  <Step>
    ### Open the web UI [#open-the-web-ui]

    Build the client once (see [Installation](/docs/installation#the-web-ui-client)), then serve this project.

    ```bash
    uv run verdict serve
    # Verdict for refunds: http://127.0.0.1:8765
    ```

    Open a failing candidate trial. The **Outcome** tab shows the failed checks and the step behind them. The **Data changes** tab shows the order's status stayed `delivered`.
  </Step>
</Steps>

## Same code on both sides [#same-code-on-both-sides]

Comparing a run with itself shows what an A/A comparison looks like. Verdict recognises the matching code fingerprint and reports the noise floor instead of judging.

```bash
uv run verdict compare 20261008-191805-nogit-ba44 20261008-191805-nogit-ba44
```

```text
Verdict: A/A: same code on both sides, so this difference is the noise floor
```

## Next [#next]

* [Core concepts](/docs/concepts) explains each part you just used.
* [Build an eval for your agent](/docs/tutorial) does the same for your app.


---

# Overview

> The parts of Verdict in one page, each linking to its own concept page.

Source: https://mitej23.github.io/docs/concepts



A **run** takes a set of cases and runs your app's entry point several times per case, each time inside a sandbox built from the case's world. What the app did is graded by the case's checks and saved to a local SQLite store. **Compare** turns two sets of runs into a verdict.

```mermaid
flowchart TB
  T[&#x22;verdict.toml&#x22;] --> R[&#x22;runner&#x22;]
  C[&#x22;cases/&lt;id&gt;/&#x22;] --> R
  R -->|per trial| SB[&#x22;sandbox from the world&#x22;]
  SB --> SA[&#x22;state adapter: fresh database&#x22;]
  SB --> BF[&#x22;boundary: recording fakes&#x22;]
  SB --> NG[&#x22;network guard&#x22;]
  SB --> CL[&#x22;clock: frozen&#x22;]
  R --> CH[&#x22;checks → reward&#x22;]
  CH --> ST[(&#x22;store: SQLite&#x22;)]
  ST --> CMP[&#x22;compare → verdict, exit code&#x22;]
```

## Cases and worlds [#cases-and-worlds]

A case is a folder: `case.toml` holds the input and the checks, `world.json` holds the starting database rows and scripted service responses. [Read more](/docs/concepts/cases-and-worlds).

## Boundary fakes [#boundary-fakes]

Every external service your app calls is declared as a `[[boundary]]`. Verdict replaces it with a fake that records each call as an **effect** and answers from the world. [Read more](/docs/concepts/boundary-fakes).

## In-process isolation and the network guard [#in-process-isolation-and-the-network-guard]

Verdict imports your app into its own process and swaps things at the edges. Any host that no fake covers and `[network].allow` doesn't list fails the trial. [Read more](/docs/concepts/isolation-and-network-guard).

## State adapters [#state-adapters]

The SQLAlchemy adapter gives each trial its own in-memory database, built from the world, and swaps your app's session factory to use it. [Read more](/docs/concepts/state-adapters).

## Checks [#checks]

Checks grade effects, final state, output and errors. Every check returns pass or fail and a detail that says why. [Read more](/docs/concepts/checks).

## Trials and pass^k [#trials-and-passk]

Agents are noisy, so each case runs `k` times. A case "passes every trial" (pass^k) only when all `k` trials pass. [Read more](/docs/concepts/trials-and-pass-k).

## Code versions and fingerprints [#code-versions-and-fingerprints]

Every run records a fingerprint of the code it tested, so two runs of the same code are recognised and a candidate can run from another folder or a git ref. [Read more](/docs/concepts/code-versions).

## Compare and the gate [#compare-and-the-gate]

Compare pairs runs case by case: a bootstrap interval for the overall change, Holm-adjusted Fisher tests per case, regressed and watch flags, a holdout split, A/A detection and a CI gate. [Read more](/docs/concepts/compare-and-the-gate).

## Grader self-test and regrade [#grader-self-test-and-regrade]

`oracle.json` and `known_bad.json` prove each case's checks are right before any run. `verdict regrade` re-scores stored trials after you change checks. [Read more](/docs/concepts/self-test-and-regrade).

## Traces and observations [#traces-and-observations]

With the `otel` extra, every trial keeps its OpenTelemetry spans. Verdict turns them into typed steps and points each failed check at the step behind it. [Read more](/docs/concepts/traces-and-observations).

## Generating cases [#generating-cases]

`verdict generate` writes many cases from a spec. Your code decides what is correct; Claude Code writes the content; gates and an independent reviewer decide what becomes a case. [Read more](/docs/concepts/generating-cases).


---

# Cases and worlds

> A case is a folder with the input, the starting world and the checks for one situation.

Source: https://mitej23.github.io/docs/concepts/cases-and-worlds



A case is one situation your app must handle, written as a folder of plain files. A world is the state that situation starts from.

<Files>
  <Folder name="cases">
    <Folder name="refund-within-window">
      <File name="case.toml" />

      <File name="world.json" />

      <File name="oracle.json" />

      <File name="known_bad.json" />

      <File name="checks.py" />
    </Folder>

    <File name="holdout.json" />
  </Folder>
</Files>

| File             | Required | What it holds                                                          |
| ---------------- | -------- | ---------------------------------------------------------------------- |
| `case.toml`      | yes      | id, title, input, checks, and optionally tags, entry, trials and clock |
| `world.json`     | no       | starting database rows and scripted service responses                  |
| `oracle.json`    | no       | a known-good result the checks must pass                               |
| `known_bad.json` | no       | a known-bad result the checks must fail                                |
| `checks.py`      | no       | Python checks the case refers to                                       |

## The case [#the-case]

```toml title="cases/refund-within-window/case.toml"
schema = 1
id = "refund-within-window"
title = "Broken mug, delivered 10 days ago: refund it"
tags = ["refund", "happy-path"]
clock = "2026-10-08T10:00:00+00:00"

[input]
order_id = "A100"
message = "My mug arrived broken, can I get my money back?"

[[check]]
kind = "effect"
service = "payments"
action = "refund"
count = 1
args = { order_id = "A100", amount = 24.0 }

[[check]]
kind = "state"
table = "orders"
where = { id = "A100" }
equals = { status = "refunded" }
```

`[input]` is passed to your entry point as a dict. The checks describe the outcome, not the wording.

## The world [#the-world]

```json title="cases/refund-within-window/world.json"
{"schema": 1,
 "tables": {"orders": [{"id": "A100", "customer_email": "maya@example.com", "status": "delivered",
                        "total": 24.0, "delivered_at": "2026-09-28T12:00:00"}]},
 "services": {"payments": {"refund": {"returns": {"refund_id": "R-1", "status": "ok"}}}}}
```

* `tables` are the rows each trial's database starts with.
* `services` script what each faked method returns: a value, a sequence of values, or an error.

Every trial gets its own deep copy of the world, so trials never see each other's changes.

## Rules [#rules]

* **Ids are stable and unique.** Lowercase letters, digits, `.`, `_` and `-`. The id defaults to the folder name.
* **At least one check.** A case without checks doesn't load.
* **Folders starting with `_` or `.` are skipped.** The generator uses `_rejected/` and `_staging/` this way.
* **Formats are versioned.** Every file carries `schema = 1`. A different schema fails loudly.

## Holdout cases [#holdout-cases]

List case ids in `cases/holdout.json` to report them separately in compare. Tune on the rest; judge on holdout.

```json title="cases/holdout.json"
{"holdout": ["payments-down"]}
```

## Reference [#reference]

* [case.toml](/docs/reference/case-toml)
* [world.json](/docs/reference/world-json)
* [oracle.json and known_bad.json](/docs/reference/oracle-and-known-bad)
* [holdout.json](/docs/reference/holdout-json)


---

# Boundary fakes

> Every external service is replaced by a fake that records what your app asked it to do.

Source: https://mitej23.github.io/docs/concepts/boundary-fakes



A boundary is an external service your app calls: payments, email, a calendar API. Verdict replaces it with a fake for the length of a run. The fake records every call as an **effect** and answers from the case's world.

```toml title="verdict.toml"
[[boundary]]
service = "payments"
patch = "refunds_app.services:PaymentsClient"   # a class

[[boundary]]
service = "email"
patch = "refunds_app.services:send_email"       # a function
```

## Effects [#effects]

Every faked call becomes one effect, in the order it happened:

```json
{"seq": 1, "service": "payments", "action": "refund",
 "args": {"order_id": "A100", "amount": 24.0}, "ok": true,
 "result": {"refund_id": "R-1", "status": "ok"}}
```

* `action` is the method or function name.
* `args` are the call's arguments by parameter name.
* `ok` is false when the call raised.

Checks such as `effect` and `no_effect` read this list. With OpenTelemetry installed, each effect also carries the `span_id` of the step that made it.

## Class targets [#class-targets]

When `patch` names a class, every public method defined on that class is replaced. Instances created anywhere, before or after the swap, are covered. Methods whose names start with `_`, static methods and class methods are left alone.

## Function targets [#function-targets]

When `patch` names a function, Verdict swaps it in every loaded module under your app root that refers to it, however it was imported. Apps write `from services import send_email`, which copies the function into the importing module; patching only `services` would miss that copy. Modules imported later are swapped before each case.

## Scripted responses [#scripted-responses]

By default a faked call returns what the world scripts for it:

```json title="world.json"
{"services": {"payments": {
  "refund": {"returns": {"refund_id": "R-1", "status": "ok"}},
  "lookup": {"sequence": [{"found": true}, {"found": false}]},
  "charge": {"raises": "card declined"}
}}}
```

| Script     | Behaviour                                                                |
| ---------- | ------------------------------------------------------------------------ |
| `returns`  | returns the value on every call                                          |
| `sequence` | one value per call, repeating the last once the list runs out            |
| `raises`   | raises `BoundaryError(message)`; the effect is recorded with `ok: false` |
| nothing    | returns `None`                                                           |

## Custom fakes [#custom-fakes]

When a service needs logic (a limit, a lookup in the world), write a fake class. Name it with `fake`:

```toml title="verdict.toml"
[[boundary]]
service = "payments"
patch = "refunds_app.services:PaymentsClient"
fake = "my_evals.fakes:Payments"
```

```python title="my_evals/fakes.py"
from verdict.boundary import record, scripted

class Payments:
    def refund(self, order_id, amount):
        if amount > 500:
            record("payments", "refund", {"order_id": order_id, "amount": amount}, ok=False, result="limit")
            raise RuntimeError("refund limit")
        result = scripted("payments", "refund", default={"status": "ok"})
        record("payments", "refund", {"order_id": order_id, "amount": amount}, ok=True, result=result)
        return result
```

* Methods take the same arguments as the real ones.
* A custom fake records its own effects with `record()`.
* One instance is created per trial and service.
* Methods the fake doesn't define fall back to the recording stub.

See the [Python API](/docs/reference/python-api#verdictboundary) for `record` and `scripted`.

## Outside a trial [#outside-a-trial]

A fake called outside a Verdict trial raises `BoundaryError`. Everything installed is undone when the run ends.


---

# In-process isolation and the network guard

> How Verdict contains a trial inside your app's own Python process, and why a forgotten boundary fails loudly.

Source: https://mitej23.github.io/docs/concepts/isolation-and-network-guard



Verdict runs your app in-process: it imports your code into its own Python process and swaps things at the edges. Nothing in your app is edited.

## What gets swapped [#what-gets-swapped]

When a run starts, Verdict:

1. Puts your app root first on `sys.path` and imports your entry points.
2. Installs the [state adapter](/docs/concepts/state-adapters), which swaps your session factory.
3. Installs a fake for every [`[[boundary]]`](/docs/concepts/boundary-fakes).
4. Installs the network guard.
5. Starts OpenTelemetry collection if the `otel` extra is installed.

For each case it freezes the clock if the case sets `clock`, then runs the trials. When the run ends, everything is undone in reverse order.

## The trial context [#the-trial-context]

Trials of one case run concurrently: async entry points as asyncio tasks, sync ones in threads. The current trial lives in a `ContextVar`, which fakes and the state adapter read, so concurrent trials never share effects or databases.

## The network guard [#the-network-guard]

While a trial is running, resolving a host name that isn't allowed raises `NetworkBlocked`:

```text
NetworkBlocked: The app tried to reach 'payments.example.com', which no boundary fakes and [network].allow doesn't list
```

The app's error is recorded on the trial and the checks still run, so a boundary you forgot shows up as a failed trial instead of a real refund.

Allow the hosts your app may reach, usually only model APIs:

```toml title="verdict.toml"
[network]
allow = ["api.openai.com", "api.anthropic.com"]
```

A host matches when it equals an entry or ends with `.` plus the entry, so `"openai.com"` also allows `api.openai.com`. The guard works at DNS resolution (`socket.getaddrinfo`) and is only active inside a trial.

## Errors are data [#errors-are-data]

An exception from your app doesn't stop the run. Verdict stores it as the trial's `error` (with a short traceback), still dumps the final state, and grades what happened before the error. Use a `no_error` check when an error itself is a failure.

## Limits [#limits]

* **Python only.** The app must expose an importable entry callable and, for state, an importable session factory.
* **Wiring matters.** Isolation depends on the app reaching services through the classes and functions you declare. A client created from a third-party SDK that isn't declared as a boundary is caught by the network guard, not faked.
* **One process per run.** The swaps are process-wide, which is why the web UI starts each run as a separate `verdict run` process.

A process mode, with the app in a container behind an egress proxy, is planned for apps in other languages. The file formats are designed so it can sit underneath them unchanged.


---

# State adapters

> A state adapter gives every trial its own database, built from the case's world.

Source: https://mitej23.github.io/docs/concepts/state-adapters



A state adapter gives each trial a fresh copy of your app's state, built from `world.json`, and dumps the final state for checks. Your production database is never opened.

```toml title="verdict.toml"
[state]
adapter = "sqlalchemy"
metadata = "refunds_app.db:Base.metadata"
session = "refunds_app.db:SessionLocal"
```

## The SQLAlchemy adapter [#the-sqlalchemy-adapter]

`sqlalchemy` is the only adapter today. Install the `sqlalchemy` extra. For each trial it:

1. Creates an in-memory SQLite database and runs `metadata.create_all`.
2. Inserts the rows from `world.tables`, converting ISO strings for date, datetime and time columns.
3. Hands your app sessions bound to that database, through a stand-in for your session factory.
4. After the entry point returns, dumps every table in the metadata as `{"tables": {name: [rows]}}`.

The stand-in replaces the `session` target everywhere it was imported by name, the same way function boundaries are swapped. Calling it, or reading an attribute such as `SessionLocal.kw`, goes to a sessionmaker bound to the current trial's database.

## Postgres types [#postgres-types]

SQLite lacks some Postgres types, so they are mapped before `create_all`:

| Postgres type                                           | Becomes      |
| ------------------------------------------------------- | ------------ |
| `UUID`, `INET`, `CIDR`, `MACADDR`, `CITEXT`, `TSVECTOR` | `String(64)` |
| `ARRAY`, `JSONB`, `HSTORE`                              | `JSON`       |

Server defaults that call a function, such as `gen_random_uuid()` or `now()`, are dropped (`CURRENT_TIMESTAMP` is kept).

## Final state [#final-state]

Checks of kind `state` read the dump. Values are made JSON-safe: dates become ISO strings, decimals become floats, UUIDs become strings.

```toml
[[check]]
kind = "state"
table = "orders"
where = { id = "A100" }
equals = { status = "refunded" }
```

## Errors [#errors]

* A table in `world.json` that the metadata doesn't define raises an error.
* Your app opening a session outside a trial raises `StateError`.

## Known limit [#known-limit]

Apps that rely on Postgres-specific behaviour can fail on SQLite. One example: SQLAlchemy bulk inserts with UUID primary keys hit an "insertmanyvalues sentinel" error. A Postgres mode is planned for those apps.

See [Use the SQLAlchemy adapter](/docs/guides/sqlalchemy-adapter).


---

# Checks

> Checks grade what a trial did, its effects, final state, output and error, and say why they failed.

Source: https://mitej23.github.io/docs/concepts/checks



A check is one assertion about what a trial did. A case has one or more `[[check]]` tables; a trial passes only when every check passes.

```toml title="case.toml"
[[check]]
kind = "no_effect"
service = "payments"
action = "refund"

[[check]]
kind = "effect"
service = "email"
args = { to = "support@shop.example" }
count = 1

[[check]]
kind = "state"
table = "orders"
where = { id = "A103" }
equals = { status = "escalated" }
```

## Kinds [#kinds]

| Kind           | Passes when                                                                          |
| -------------- | ------------------------------------------------------------------------------------ |
| `effect`       | matching calls were made (exactly `count`, or at least one)                          |
| `no_effect`    | no matching call was made                                                            |
| `effect_count` | the number of effects, optionally on one service, equals `count`                     |
| `state`        | a row matching `where` exists and has `equals`; with `absent = true`, no row matches |
| `output`       | the output, or a field of it, contains, matches or equals a value                    |
| `no_error`     | the entry point didn't raise                                                         |
| `python`       | your function returns true                                                           |

Every field is in the [check kinds reference](/docs/reference/checks).

## Subset matching [#subset-matching]

`args`, `where` and `equals` match on the keys you write. Other keys are ignored.

* Dicts match when every expected key matches.
* Lists match element by element and must be the same length.
* Numbers match within 1e-9.

`args = { order_id = "A100" }` matches a refund of any amount on order A100.

## Details say why [#details-say-why]

Every check returns a name, pass or fail, and a detail. The detail is what you read when a trial fails:

```text
payments.refund(order_id=A100, amount=24.0) ×1   0 matching (want 1); calls to it: no matching calls
orders[id=A100] status=refunded                  got {'status': 'delivered'}
```

Set `name` on a check to replace its generated label.

## Reward [#reward]

Each trial gets a reward: the fraction of its checks that passed. `passed` is true only when all of them did. Compare works on `passed`, never on the reward.

## A check that breaks [#a-check-that-breaks]

A check that raises (a bad regex, a missing Python function) fails with `check raised …` in its detail. It never crashes the run.

See [Write checks](/docs/guides/write-checks).


---

# Trials and pass^k

> Each case runs k times because agents are noisy; pass^k means every one of the k trials passed.

Source: https://mitej23.github.io/docs/concepts/trials-and-pass-k



A trial is one call to your entry point for one case. A run gives every case `k` trials, because the same agent on the same input can pass once and fail the next time.

## Why more than one [#why-more-than-one]

On the first production app Verdict was built for, one case went from 0/5 to 3/5 with identical code. A single trial can't tell a fix from luck. Several trials per case give each case a pass rate, and compare works on those rates.

## How many [#how-many]

The number of trials for a case is, in order of precedence:

1. `-k` / `--trials` on `verdict run`;
2. `trials` in the case's `case.toml`;
3. `[run].trials` in `verdict.toml` (default 2).

```bash
verdict run -k 3
```

Trials of one case run concurrently, up to `[run].concurrency` at a time (default 4). Cases run one after another.

## pass^k [#passk]

A case **passes every trial**, or pass^k, when all `k` of its trials pass. This is stricter than pass\@k (at least one of `k` passed), and it's what a user of a product experiences: the agent has to get it right every time.

Verdict reports both views:

* `verdict run` prints each case's marks, e.g. `✓ ✕ ✓   2/3 passed`.
* The store keeps each case's `passed_trials`, `k` and `pass_k` (1 when every trial passed).
* `verdict compare` prints "all trials pass 5 → 4": how many cases passed every trial on each side.

Verdict never mixes single-shot and best-of-N results. Reporting the best of several attempts as if it were one inflates results.

## The case is the unit [#the-case-is-the-unit]

Trials are pooled within a case, never across cases. Ten trials of an easy case don't outweigh one trial of a hard one. See [Compare and the gate](/docs/concepts/compare-and-the-gate).

## Cost [#cost]

Trials multiply cost. Start low during iteration and run the full suite once before merging. See [Keep costs down](/docs/guides/keep-costs-down).


---

# Code versions and fingerprints

> Every run records a fingerprint of the code it tested, so results always name their code.

Source: https://mitej23.github.io/docs/concepts/code-versions



Every run records which code it tested. A result always names the code that produced it, and two runs of identical code are recognised as an A/A comparison.

## The fingerprint [#the-fingerprint]

The fingerprint is a hash of every `.py` file under `[app].paths`, as it is on disk, including uncommitted edits. File paths are part of the hash, so a rename changes it. `__pycache__` and `.verdict` are skipped.

```toml title="verdict.toml"
[app]
root = "."
paths = ["refunds_app"]     # what the fingerprint covers; empty means the whole app root
```

`verdict runs` shows the first 12 hex characters:

```text
20261008-191805-nogit-8dd8           finished      12/15   code 4a5ffdf5d401  candidate
20261008-191805-nogit-ba44           finished      15/15   code 8cfd65598bb3  baseline
```

Keep `paths` to your app's code. Prompts or config stored outside `.py` files are not part of the fingerprint.

## What a run records [#what-a-run-records]

Alongside the fingerprint, each run stores:

* the Verdict version, the config path and the app root;
* git information: short SHA, branch, and whether the paths are dirty;
* when dirty, the uncommitted diff of those paths, new files included (up to 200,000 characters);
* the planned number of trials per case.

The run id is built from the time and the commit: `YYYYMMDD-HHMMSS-<sha>-<random>`, with `nogit` outside a repository.

## Candidates [#candidates]

A candidate is another version of your code run against the same cases. There are two ways to point a run at one:

```bash
# a folder: another checkout, a copy, a worktree you made
verdict run --app-root ../my-app-feature --label candidate

# a git ref: Verdict checks it out in its own worktree
verdict run --code my-branch --label candidate
```

`--code` creates a detached git worktree under `.verdict/worktrees/<sha>` (next to the store) and runs the app from the same relative path inside it. Your working copy is untouched. An existing worktree for that SHA is reused.

Both options keep the cases, checks and store from your `verdict.toml`: only the code changes.

## A/A [#aa]

When both sides of a comparison have one and the same fingerprint, compare reports the noise floor instead of judging, and `--gate` never fails. Run your baseline twice to see how much your cases move on their own.

## In the web UI [#in-the-web-ui]

The **Code** page lists `main` (the configured app root, with its uncommitted changes), every candidate folder a run has used, and every git worktree of the repository, each diffed against `main`. See [Web UI](/docs/web-ui#code).

See [Compare a candidate](/docs/guides/compare-a-candidate).


---

# Compare and the gate

> How Verdict decides whether a change helped beyond noise, what it flags, and when the CI gate fails.

Source: https://mitej23.github.io/docs/concepts/compare-and-the-gate



`verdict compare BASE CAND` compares two sets of runs case by case and says whether the candidate is better, worse or the same beyond noise. With `--gate` it exits 1 on a regression.

```bash
verdict compare <baseline-run> <candidate-run> --gate
```

Each side is one or more run ids, comma-separated. Trials of the same case from several runs are pooled: two baseline runs just give more baseline trials.

## The method [#the-method]

The case is the statistical unit. Trials are pooled within a case, never across cases.

1. **Per case:** the candidate pass rate minus the baseline pass rate.
2. **Overall change:** the mean of those per-case differences.
3. **95% interval:** a two-stage bootstrap (4,000 resamples). Resample cases, then draw each case's pass rate from Beta(passes + ½, fails + ½).
4. **Per case test:** Fisher's exact test on the case's 2×2 table, Holm-adjusted for the number of cases compared. A case is **beyond noise** when its adjusted p is below 0.05.
5. **Flags** for each case (below).
6. **A/A detection:** when both sides have the same [code fingerprint](/docs/concepts/code-versions).
7. **Splits:** holdout cases are summarised separately from dev cases.

Only cases present on both sides are compared. Cases on one side only are counted as "not compared".

### Why not a plain average [#why-not-a-plain-average]

* **Resampling trials collapses intervals.** A case that went 4/4 would never fail in a trial bootstrap, which makes small-`k` intervals falsely narrow. The Beta draw keeps that uncertainty.
* **An average dilutes targeted changes.** On a 7-case suite, a fix confined to 2 cases left the overall interval at −8% to +56% ("no difference"). The per-case Holm-adjusted test caught it: 0/10 to 5/5, adjusted p = 0.002.

## Flags [#flags]

| Flag               | When                                                                                                             |
| ------------------ | ---------------------------------------------------------------------------------------------------------------- |
| `regressed`        | the case passed every baseline trial and fails 2 or more candidate trials, or its pass rate fell by half or more |
| `watch`            | the case passed every baseline trial and fails exactly one candidate trial: rerun it with more trials            |
| `improved`         | the pass rate rose by half or more, or went from never passing to always passing                                 |
| `better` / `worse` | the pass rate rose or fell by less                                                                               |
| `same`             | no change                                                                                                        |

## The gate [#the-gate]

With `--gate`, compare exits 1 when any of these holds:

* a case **regressed**;
* a case is **worse beyond noise** (adjusted p \< 0.05 and a negative change);
* the overall 95% interval is **entirely below zero**;
* cost per trial rose by more than `--max-cost-increase` (a fraction: `0.2` is 20%).

It also exits 1 when the two sides share no cases. It never fails an A/A comparison.

## Verdicts [#verdicts]

The last line is one of:

| Verdict                                                               | Meaning                                     |
| --------------------------------------------------------------------- | ------------------------------------------- |
| `A/A: same code on both sides, so this difference is the noise floor` | identical fingerprints; nothing is judged   |
| `fails the gate`                                                      | one or more gate reasons, listed below it   |
| `better beyond noise`                                                 | the overall interval is entirely above zero |
| `N case(s) better beyond noise, the rest unchanged within noise`      | some cases improved significantly           |
| `no difference beyond noise`                                          | nothing stands out                          |
| `nothing to compare: the two sides share no cases`                    | no overlap                                  |

"No difference beyond noise" is a common, honest answer. It means you need more cases or more trials to tell, not that the change did nothing.

## Holdout [#holdout]

Cases listed in `cases/holdout.json` are reported on their own line. Iterate against dev cases, then judge the change on holdout cases you haven't tuned against. `verdict run --split dev` and `--split holdout` run one side only; see [holdout.json](/docs/reference/holdout-json#run-one-split).

## Cost and time [#cost-and-time]

When trials have a cost (from [OpenTelemetry spans and pricing](/docs/concepts/traces-and-observations)), compare prints mean cost per trial on each side and the relative change. It prints mean seconds per trial too.

## Sensitivity [#sensitivity]

With about 28 cases and 5 trials each, only overall changes of roughly 10 points or more stand out. A per-case test needs about 10 baseline trials against 5 candidate trials to call a 0% to 100% change beyond noise after Holm correction. Use `watch` flags to decide which cases deserve more trials.

## Machine-readable output [#machine-readable-output]

`--json` prints the full comparison: verdict, gate reasons, per-split summaries, cost and time, and one row per case with its flag, p and adjusted p.

See [Gate CI on compare](/docs/guides/gate-ci) and the [CLI reference](/docs/reference/cli#verdict-compare).


---

# Grader self-test and regrade

> Prove each case's checks are right before any run, and re-score old trials after you change them.

Source: https://mitej23.github.io/docs/concepts/self-test-and-regrade



A wrong grader makes every result wrong: it fakes progress or fakes a ceiling. The self-test checks every grader against results you already know the answer for, before any model is called.

## Known-good and known-bad results [#known-good-and-known-bad-results]

Put two optional files next to a case:

* `oracle.json`: what a correct app does. **Every** check must pass on it.
* `known_bad.json`: what a plausibly broken app does. **At least one** check must fail on it.

```json title="cases/payments-down/known_bad.json"
{"output": {"decision": "refunded", "reply": "Done: your refund of $80.00 is on its way."},
 "effects": [{"seq": 1, "service": "payments", "action": "refund", "args": {"order_id": "A104", "amount": 80.0}, "ok": false, "result": "503"},
             {"seq": 2, "service": "email", "action": "send_email", "args": {"to": "sam@example.com"}, "ok": true, "result": null}],
 "state": {"tables": {"orders": [{"id": "A104", "status": "refunded"}]}}}
```

Each file is a trial result. Only `output`, `effects`, `state` and `error` are needed.

## Run the self-test [#run-the-self-test]

```bash
verdict self-test
```

```text
—    large-order-escalates: no oracle.json or known_bad.json
ok   payments-down: known-good 1.0, known-bad 0.25
ok   refund-within-window: known-good 1.0, known-bad 0.25
```

The numbers are the reward each result got. A failure names the problem:

```text
FAIL refund-within-window: known-good 0.75, known-bad 0.25
       checks fail the known-good result: output.reply contains 'refund' (Done: your money is on its way.)
```

`verdict self-test` exits 1 when any case fails.

## Before every run [#before-every-run]

`verdict run` self-tests the selected cases first and refuses to start if one fails:

```text
Self-test failed for refund-within-window: run `verdict self-test`. A wrong grader makes every result wrong (--skip-self-test to run anyway).
```

Pass `--skip-self-test` to run anyway.

## Regrade [#regrade]

When you fix a check, stored trials still carry their old grades. `verdict regrade` re-scores them with the current checks. It reads each trial's stored result, so it makes no model calls and doesn't rerun your app.

```bash
verdict regrade                    # every finished, failed or cancelled run
verdict regrade --runs RUN_A,RUN_B
```

```text
Re-graded 2 run(s); 0 trial(s) changed result.
```

It updates each trial's grade, each run's per-case summary and the run totals. Trials of cases that no longer exist are left as they were.

## In the web UI [#in-the-web-ui]

A case's **Checks** tab shows the self-test result and a table of every check against both results: ✓ or ✕ for the known-good and the known-bad, with the reason each failed. Below it, the two results side by side (output and calls). A wrong grader shows up there as a ✕ in the known-good column, or a known-bad column of ✓. When a grader is wrong, the case page says so at the top. **Runs** has a "Re-grade all runs" button.


---

# Traces and observations

> Verdict keeps each trial's OpenTelemetry spans, turns them into typed steps, and points failed checks at the step behind them.

Source: https://mitej23.github.io/docs/concepts/traces-and-observations



With the `otel` extra installed, Verdict collects the OpenTelemetry spans your app emits during each trial. It stores them with the trial, prices model calls, and turns them into a trace of typed steps.

## Collection [#collection]

When a run starts, Verdict adds a span processor to the global tracer provider (it creates an SDK `TracerProvider` if your app hasn't set one). Each trial runs inside a `verdict.trial` span with the attributes `verdict.case` and `verdict.trial`, and every span in that trace is kept with the trial.

Frameworks that emit OpenTelemetry GenAI spans show up without app changes, for example pydantic-ai (via Logfire), the OpenAI Agents SDK, and LangGraph via OpenInference.

Spans stay local unless you add [`[trace.export]`](/docs/guides/opentelemetry#send-traces-to-your-backend), which also sends them to an OTLP backend, tagged as eval traffic.

Each trial also records its usage: model calls, tool calls, tokens in and out, the models used and the estimated cost. The run page sums them (mean reward, tokens, models, errored trials, average time per case).

## Kinds [#kinds]

Each span gets a kind from its attributes:

| Kind    | Inferred from                                                                                                |
| ------- | ------------------------------------------------------------------------------------------------------------ |
| `llm`   | `gen_ai.operation.name` is `chat`, `text_completion` or `generate_content`, or `gen_ai.request.model` is set |
| `tool`  | `gen_ai.operation.name` is `execute_tool`, or `gen_ai.tool.name` is set                                      |
| `agent` | `gen_ai.operation.name` is `invoke_agent`, or `gen_ai.agent.name` / `agent_name` is set                      |
| `log`   | Logfire log spans                                                                                            |
| `span`  | anything else                                                                                                |

## Cost [#cost]

For every `llm` span, Verdict reads `gen_ai.usage.input_tokens`, `gen_ai.usage.output_tokens` and cached tokens, finds the model in `gen_ai.response.model` or `gen_ai.request.model`, and prices it from [pricing.toml](/docs/guides/keep-costs-down#pricing). A trial with no model calls costs 0. A trial that calls a model with no price has an unknown cost, never a wrong one.

## Observations [#observations]

The trace view turns spans into steps, in the style of Langfuse:

* **agent:** the agent's name, its system prompt, the first user message and its final result;
* **generation:** one model call, with model, usage, cost, the new input and the output. Its row reads as the decision, e.g. `→ refund(order_id=A1, amount=24.0)` or `says "…"`;
* **tool:** the tool's name, arguments and response, marked as a warning when the response has an error status;
* **event:** a log line;
* **span:** anything else.

Verdict reads two attribute styles: the OpenTelemetry GenAI semantic conventions (`gen_ai.input.messages`, `gen_ai.output.messages`, `gen_ai.tool.call.arguments`, `gen_ai.tool.call.result`, `gen_ai.agent.name`, `gen_ai.usage.*`) and Logfire / pydantic-ai attributes. Wrapper spans such as pydantic-ai's "running 1 tool" are folded away.

Name agents for display in `verdict.toml`:

```toml title="verdict.toml"
[trace]
agent_names = { refund_agent = "Refund agent" }
```

## Blame [#blame]

Each recorded [effect](/docs/concepts/boundary-fakes#effects) carries the id of the span that was current when it was made. For each failed check, Verdict picks the step most likely behind it:

* **effect checks** (`effect`, `no_effect`, `effect_count`): the step where the offending or nearest call was made;
* **every other check:** the outermost agent, which owns the output and the final state.

The web UI shows this as "Where it went wrong" on a failing trial, linking to the step in the trace.

## Without OpenTelemetry [#without-opentelemetry]

Without the extra, runs work the same; trials just have no spans, no trace and no cost.

See [Instrument with OpenTelemetry](/docs/guides/opentelemetry).


---

# Generating cases

> verdict generate writes many realistic cases from a spec without letting a model decide what is correct.

Source: https://mitej23.github.io/docs/concepts/generating-cases



`verdict generate` writes cases from a generator spec. Your code decides what is correct for each scenario; Claude Code only writes the content around it; gates and an independent reviewer decide what becomes a case.

```bash
verdict generate --estimate -n 25     # what it would cost; no calls
verdict generate --dry-run -n 10      # the whole pipeline with a free stand-in writer
verdict generate -n 25 --budget 5     # real generation through Claude Code
```

## The idea [#the-idea]

Writing cases by hand doesn't scale, and letting a model write them, checks included, means the model decides what "correct" is. Verdict splits the job:

* **Code decides correctness.** Your spec's `plan(values, rng)` turns a scenario into concrete facts (ids, amounts, dates) and the case's checks. It's your product rules written as code.
* **Claude writes content.** The customer's message, the replies, the world rows: whatever your spec asks for, as JSON matching a schema. It never writes checks.
* **Gates filter.** A case is written only after it passes every gate.

That is what keeps volume from costing correctness.

```mermaid
flowchart TB
  S[&#x22;sample scenarios<br/>fill coverage gaps&#x22;] --> P[&#x22;plan()<br/>facts + checks&#x22;]
  P --> W[&#x22;Claude writes content&#x22;]
  W --> B[&#x22;build()&#x22;]
  B --> G[&#x22;gates&#x22;]
  G -->|pass| R[&#x22;independent reviewer&#x22;]
  R -->|pass| C[&#x22;cases/&lt;id&gt;/&#x22;]
  G -->|fail| F[&#x22;retry once with feedback&#x22;]
  R -->|fail| F
  F -->|fails again| X[&#x22;cases/_rejected/&lt;id&gt;.json&#x22;]
```

## The spec [#the-spec]

A spec is three files, by default in `generator/` next to `verdict.toml`:

<Files>
  <Folder name="generator">
    <File name="generator.toml" />

    <File name="rules.md" />

    <File name="scenarios.py" />
  </Folder>
</Files>

* `generator.toml`: the scenario dimensions with weights, models and budget. See [generator.toml](/docs/reference/generator-toml).
* `rules.md`: the product's rules in plain words, shown to the writer and the reviewer.
* `scenarios.py`: the hooks. `plan()` is required. See [Generator hooks](/docs/reference/generator-hooks).

## Sampling [#sampling]

A scenario is one value from each dimension, e.g. `situation = delivered_recently`, `order_size = small`, `payments = down`, `tone = angry`.

* Values are drawn by weight, and a value is boosted when it has fewer cases than its weight asks for, counting the generated cases already in the folder. A batch fills coverage gaps while keeping the configured mix: a 9:1 split stays near 9:1.
* The optional `normalise()` hook blanks out dimensions that can't matter for a combination, so duplicates collapse.
* The same `--seed` draws the same scenarios (default 7).

Case ids are `<prefix>-<value of id_from>-<hash>`, e.g. `gen-delivered-recently-ca132a`.

## The gates [#the-gates]

Cheapest first. A case must pass all of them:

1. **Builds.** The content turns into a case with `title`, `input`, `world`, `oracle` and `known_bad`.
2. **No placeholders.** Nothing like `[NAME]`, `{{var}}`, `<USER>`, "lorem ipsum", TODO or TBD.
3. **Facts used.** Every fact `plan()` fixed appears in the input or world, so the checks test something the case set up.
4. **Not a duplicate.** The input isn't a near-copy of an existing case's input (similarity at or above `duplicate_threshold`, default 0.9).
5. **The spec's `validate()`.** Your own extra rules.
6. **Loads.** The case folder loads like any other case.
7. **Grader self-test.** The checks pass the generated known-good result and fail the known-bad one. See [self-test](/docs/concepts/self-test-and-regrade).
8. **Independent reviewer.** A second Claude call, with the rules and the scenario, must say there is a single correct outcome and that it's the scenario's, that the content matches the scenario, and give realism of at least 3 out of 5.

## Retry and rejection [#retry-and-rejection]

A scenario that fails a gate is retried once, with the issues fed back to the writer. If it fails again, it goes to `cases/_rejected/<id>.json` with every attempt's issues, for auditing. Folders and files starting with `_` are never loaded as cases, so rejected scenarios never run.

## Output [#output]

A case that passes is written straight to the cases folder and runs with the rest:

<Files>
  <Folder name="cases/gen-delivered-recently-ca132a">
    <File name="case.toml" />

    <File name="world.json" />

    <File name="oracle.json" />

    <File name="known_bad.json" />

    <File name="generation.json" />
  </Folder>
</Files>

* `case.toml` carries the tags `generated` plus the plan's tags, and a `[generator]` table with the scenario's values and every dimension's values (`_dimensions`), which the web UI uses for coverage.
* `generation.json` records the scenario, Claude's raw content, every attempt with the gates' issues and the reviewer's answer, and the cost. Nothing reads it at run time.

## Through Claude Code [#through-claude-code]

Calls go through Claude Code's headless mode (`claude -p`), so generation uses your Claude Code login rather than an API key. Each call runs in an empty temporary directory with no tools, no MCP servers, no user or project settings and no saved session, with Verdict's own system prompt. `ANTHROPIC_API_KEY` is removed from the call's environment so a key in your shell can't switch it to pay-per-token billing.

## Cost [#cost]

Costs are reported as API-equivalent prices. Through a Claude Code subscription they are not billed per token.

* A 3-case batch on the refunds example cost $0.068 API-equivalent over 6 calls in 18 seconds.
* `verdict generate --estimate -n 25` on the refunds example prints about $0.42 with prompt caching ($0.017 a case) and $0.70 without.

`[budget] usd` (or `--budget`) stops new cases from starting once spend passes it; scenarios not started are reported as skipped.

See [Write a generator spec](/docs/guides/write-a-generator-spec).


---

# Add Verdict to your app

> Point Verdict at your app's entry point, database and external services, then write a first case.

Source: https://mitej23.github.io/docs/guides/add-verdict-to-your-app



This guide connects an existing Python app to Verdict through three seams: the entry point it calls, the database it uses, and the services it reaches. You don't change the app.

## 1. Scaffold [#1-scaffold]

From your project root:

```bash
verdict init --target refunds_app.agent:handle
```

```text
Wrote /path/to/app/verdict.toml, /path/to/app/cases/example.
Next: verdict doctor
```

`verdict init DIR` scaffolds into another directory instead. It refuses to overwrite a `verdict.toml`. Without `--target`, edit `[entry.default] target` in `verdict.toml` yourself. Once the seams below are filled in, `verdict doctor` imports the app the way a run does and lists anything missing.

## 2. Find the entry point [#2-find-the-entry-point]

The entry point is the function production calls: a route handler's body, a queue consumer, an agent's `respond()`. It takes one dict and returns anything; sync or async both work.

```python title="refunds_app/agent.py"
def handle(request: dict) -> dict:
    ...
    return {"decision": "refunded", "reply": "Done: your refund of $24.00 is on its way."}
```

```toml title="verdict.toml"
[app]
root = "."
paths = ["refunds_app"]

[entry.default]
target = "refunds_app.agent:handle"
```

`target` is `module:attribute`, importable from `[app].root`. If your handler takes a request object instead of a dict, add a thin function that builds one from the dict and calls the real handler, and point `target` at that.

Several entry points are fine. Name each `[entry.<name>]` and pick one per case with `entry = "<name>"`.

## 3. Declare the database [#3-declare-the-database]

If your app uses SQLAlchemy, point Verdict at its metadata and the session factory the app calls:

```toml title="verdict.toml"
[state]
adapter = "sqlalchemy"
metadata = "refunds_app.db:Base.metadata"
session = "refunds_app.db:SessionLocal"
```

See [Use the SQLAlchemy adapter](/docs/guides/sqlalchemy-adapter).

## 4. Declare every external service [#4-declare-every-external-service]

List each service your app calls, by the class or function it calls it through:

```toml title="verdict.toml"
[[boundary]]
service = "payments"
patch = "refunds_app.services:PaymentsClient"

[[boundary]]
service = "email"
patch = "refunds_app.services:send_email"

[network]
allow = []                 # add model API hosts if the agent calls a real model
```

You don't have to find every call site. The network guard does it for you: run once, and every host you missed fails a trial with `NetworkBlocked` and the host name.

## 5. Write the first case [#5-write-the-first-case]

Start from a real situation, ideally a production incident. Describe what must happen, not what the reply says.

```toml title="cases/refund-within-window/case.toml"
schema = 1
id = "refund-within-window"
title = "Broken mug, delivered 10 days ago: refund it"
clock = "2026-10-08T10:00:00+00:00"

[input]
order_id = "A100"
message = "My mug arrived broken, can I get my money back?"

[[check]]
kind = "effect"
service = "payments"
action = "refund"
count = 1
args = { order_id = "A100", amount = 24.0 }
```

```json title="cases/refund-within-window/world.json"
{"schema": 1,
 "tables": {"orders": [{"id": "A100", "customer_email": "maya@example.com", "status": "delivered",
                        "total": 24.0, "delivered_at": "2026-09-28T12:00:00"}]},
 "services": {"payments": {"refund": {"returns": {"refund_id": "R-1", "status": "ok"}}}}}
```

Delete `cases/example` once you have a real case.

## 6. Run [#6-run]

```bash
verdict run -k 2
```

If a trial fails with an error, read it: a `NetworkBlocked` names a missing boundary, a `StateError` means the session factory isn't the one the app uses, and `KeyError` on `entry` means a case names an entry `verdict.toml` doesn't define.

## Next [#next]

* [Write checks](/docs/guides/write-checks) that pin down the outcome.
* Add `oracle.json` and `known_bad.json` and run `verdict self-test`. See [Grader self-test](/docs/concepts/self-test-and-regrade).
* Follow the [tutorial](/docs/tutorial) for the whole loop.


---

# Write checks

> Write checks that pin down what your app must do, using effects and state before output.

Source: https://mitej23.github.io/docs/guides/write-checks



Good checks describe the outcome a user or the business cares about: what changed in the world, which messages went out. This guide shows how to write them, strongest first.

## Start from effects [#start-from-effects]

Effects are the calls your app made through a boundary. They are the outcome of most agent actions.

```toml
# exactly one refund, of the right amount, on the right order
[[check]]
kind = "effect"
service = "payments"
action = "refund"
count = 1
args = { order_id = "A100", amount = 24.0 }

# no refund at all
[[check]]
kind = "no_effect"
service = "payments"
action = "refund"

# no calls to any service
[[check]]
kind = "effect_count"
count = 0
```

* Without `count`, `effect` passes on one or more matching calls. Set `count = 1` when a duplicate would be a bug, such as a double refund.
* `args` is a subset match: list only the arguments that matter.
* Add `ok = true` to count only calls that succeeded, or `ok = false` for failed ones.

## Then final state [#then-final-state]

```toml
[[check]]
kind = "state"
table = "orders"
where = { id = "A104" }
equals = { status = "delivered" }     # the failed refund left it unchanged

[[check]]
kind = "state"
table = "refunds"
where = { order_id = "A104" }
absent = true                          # no row was written
```

## Output last, and loosely [#output-last-and-loosely]

Checks on wording break when the wording improves. Prefer a structured field, and use text checks to forbid a specific failure:

```toml
[[check]]
kind = "output"
path = "decision"
equals = "not_delivered"

[[check]]
kind = "output"
path = "reply"
not_matches = "refund (of|is) .* on its way|refunded"
```

* `path` is a dotted path into the output; list indexes are numbers (`messages.0.text`).
* `contains` and `not_contains` are case-insensitive substrings.
* `matches` and `not_matches` are case-insensitive regular expressions, searched anywhere in the text.
* `equals` compares the value with subset matching.

## Errors [#errors]

```toml
[[check]]
kind = "no_error"
```

Use it when the app must handle a failure itself, as in `payments-down`, where the provider fails and the app must still answer.

## Python checks [#python-checks]

When no built-in kind fits, write a function. Put it in `checks.py` in the case folder:

```python title="cases/refund-within-window/checks.py"
def reply_mentions_amount(result, case):
    reply = (result.get("output") or {}).get("reply", "")
    amount = case.world["tables"]["orders"][0]["total"]
    return f"{amount:.2f}" in reply, reply[:80]
```

```toml title="case.toml"
[[check]]
kind = "python"
target = "checks:reply_mentions_amount"
name = "reply states the refunded amount"
```

* `result` is the trial result: `output`, `error`, `effects`, `state`.
* Return `(passed, detail)`, or a bool.
* `target` can also be any importable `module:function` to share checks between cases.

## Prove the checks [#prove-the-checks]

Write an `oracle.json` that a correct app produces and a `known_bad.json` that shows the failure the case is about, then:

```bash
verdict self-test
```

If the known-bad result passes, your checks don't catch the failure. If the known-good result fails, they're too strict. See [Grader self-test](/docs/concepts/self-test-and-regrade).

## After changing checks [#after-changing-checks]

Re-score stored trials without rerunning anything:

```bash
verdict regrade
```

Full field list: [Check kinds](/docs/reference/checks).


---

# Multi-turn cases

> Test a chat agent over a whole conversation, with a scripted customer or one played by a model.

Source: https://mitej23.github.io/docs/guides/multi-turn-cases



Most agents don't finish in one message: they ask for an order number, confirm a time, or hand over to a person. A multi-turn case runs a conversation inside one trial, so fakes, the database and the clock carry across turns, and the checks grade what the app did by the end.

## 1. Add a turns entry [#1-add-a-turns-entry]

Point an entry at the function your chat front end calls once per customer message, with `mode = "turns"`:

```toml title="verdict.toml"
[entry.chat]
target = "refunds_app.chat:handle_turn"
mode = "turns"
```

```python title="refunds_app/chat.py"
def handle_turn(session: dict, message: str) -> dict:
    ...   # return the reply; anything with a reply/message/text field shows as the agent's turn
```

`session` is the same dict on every turn of a trial (`id`, `case`, `trial`, `input`). If your app keeps conversations in its database, that works too: the state adapter's database lives for the whole trial.

## 2. Write a scripted user [#2-write-a-scripted-user]

A scripted user is free and gives the same conversation every time:

```toml title="cases/chat-asks-for-the-order/case.toml"
entry = "chat"

[user]
turns = [
  "Hi, my mug arrived broken. Can I get my money back?",
  { when = "order number", say = "Sure, it's A100." },
]

[[check]]
kind = "output"
scope = "conversation"
contains = "order number"
name = "asked for the order number"

[[check]]
kind = "effect"
service = "payments"
action = "refund"
count = 1
```

A `{ when, say }` turn is sent only if the app's last reply matches `when` (a case-insensitive regex), so the script follows the app instead of talking past it. The conversation ends when the script runs out.

## 3. Or let a model play the customer [#3-or-let-a-model-play-the-customer]

A simulated user gets a scenario, not lines. The fields follow τ-bench: what the customer wants, what they know and don't, and how to behave.

```toml
[user]
reason_for_call = "Your mug arrived broken and you want your money back."
known_info = "Your order number is A100."
unknown_info = "When exactly it was delivered."
instructions = "Don't give the order number until you're asked for it."
persona = "Polite, a little impatient."
max_turns = 6
```

By default Claude Code plays it (`claude -p`, model `haiku`, through your subscription). It sees only the app's replies, never its tool calls, and ends the conversation with `###STOP###`, `###TRANSFER###` or `###OUT-OF-SCOPE###`. To use your own model, set `simulator = "my_evals.users:simulate"`, a function `(user, transcript) -> str`, on the case or as `[users] simulator` in `verdict.toml`.

A simulated user adds noise of its own: pin its model, run more trials, and compare an A/A pair to see how much.

## 4. Read the conversation [#4-read-the-conversation]

```bash
verdict trial latest chat-asks-for-the-order
```

```text
Conversation (scripted user, ended: script_end)
  user   Hi, my mug arrived broken. Can I get my money back?
  agent  Sorry to hear that. What's your order number?
  user   Sure, it's A100.
  agent  Done: your refund of $24.00 is on its way.
```

A conversation that reaches `max_turns` without the user ending it, or where the app raises, fails with the check `the user ended the conversation`.


---

# Use the SQLAlchemy adapter

> Give every trial its own SQLite database from your SQLAlchemy models and the case's world.

Source: https://mitej23.github.io/docs/guides/sqlalchemy-adapter



The SQLAlchemy adapter builds a fresh in-memory database for every trial from your app's metadata and the case's `world.tables`, and swaps your session factory so the app uses it.

## Install [#install]

```bash
pip install "verdict-evals[sqlalchemy]"
```

## Configure [#configure]

You need two targets from your app: the `MetaData` that defines the tables, and the session factory the app calls.

```python title="refunds_app/db.py"
from sqlalchemy import Column, DateTime, Float, String, create_engine
from sqlalchemy.orm import declarative_base, sessionmaker

Base = declarative_base()

class Order(Base):
    __tablename__ = "orders"
    id = Column(String, primary_key=True)
    customer_email = Column(String, nullable=False)
    status = Column(String, nullable=False)
    total = Column(Float, nullable=False)
    delivered_at = Column(DateTime, nullable=True)

engine = create_engine("sqlite:///refunds-production.db")
SessionLocal = sessionmaker(bind=engine)
```

```toml title="verdict.toml"
[state]
adapter = "sqlalchemy"
metadata = "refunds_app.db:Base.metadata"
session = "refunds_app.db:SessionLocal"
```

The app keeps calling `SessionLocal()` as it always does. Inside a trial, that returns a session on the trial's own database. The production engine is never connected to.

## Seed the world [#seed-the-world]

Rows go under `tables`, keyed by table name:

```json title="world.json"
{"schema": 1,
 "tables": {"orders": [
   {"id": "A101", "customer_email": "jon@example.com", "status": "delivered",
    "total": 60.0, "delivered_at": "2026-08-24T09:00:00"}
 ]}}
```

* ISO strings are converted for `Date`, `DateTime` and `Time` columns.
* Every table name must exist in the metadata.
* Tables you don't list start empty.

## Check the final state [#check-the-final-state]

After the entry point returns, every table is dumped. Grade it with `state` checks:

```toml
[[check]]
kind = "state"
table = "orders"
where = { id = "A101" }
equals = { status = "delivered" }
```

The trial's **Data changes** tab in the web UI shows each added, removed or changed row against the world.

## The session factory must be the one the app calls [#the-session-factory-must-be-the-one-the-app-calls]

Verdict swaps the factory wherever it was imported by name. If your app builds sessions some other way (a `Session(engine)` call inline, a dependency-injection container that captured the engine at import), point `session` at whatever the app actually calls per request. A session opened outside a trial raises `StateError`.

## Postgres apps [#postgres-apps]

Postgres-only column types (`UUID`, `JSONB`, `ARRAY`, `INET` and others) are mapped to portable ones, and function server defaults like `gen_random_uuid()` are dropped. Apps that depend on Postgres behaviour beyond that can fail on SQLite; a Postgres mode is planned.

See [State adapters](/docs/concepts/state-adapters).


---

# Freeze the clock

> Run a case at a fixed time so dates like "10 days ago" or "next Monday" are reproducible.

Source: https://mitej23.github.io/docs/guides/freeze-the-clock



Set `clock` on a case to run every trial of it with time frozen at that moment. "Delivered 10 days ago" stays 10 days ago next month.

## Install [#install]

```bash
pip install "verdict-evals[clock]"
```

The extra installs `time-machine`. A case that sets `clock` without it fails the run with:

```text
case refund-within-window sets clock; install the clock extra: pip install 'verdict-evals[clock]'
```

## Set the time [#set-the-time]

```toml title="case.toml"
clock = "2026-10-08T10:00:00+00:00"
```

The value is an ISO 8601 datetime. Include the offset: a case about time zones depends on it.

When the case starts, `datetime.now()`, `time.time()` and friends jump to that instant, then time moves on normally: a trial that takes 20 seconds ends 20 seconds after the clock. Every trial of the case starts from the same instant, and the clock is released before the next case. Letting it tick keeps traces honest (each step's real duration, in order) and matches how the app behaves in production; "10 days ago" stays 10 days ago to the second.

## Write the world relative to it [#write-the-world-relative-to-it]

The refunds example freezes every case at 2026-10-08 10:00 UTC and sets delivery dates around it:

| Case                    | `delivered_at`      | Days before the clock | Expected                      |
| ----------------------- | ------------------- | --------------------- | ----------------------------- |
| `refund-within-window`  | 2026-09-28T12:00:00 | 10                    | refund                        |
| `refund-outside-window` | 2026-08-24T09:00:00 | 45                    | no refund, explain the window |

## Time and model calls [#time-and-model-calls]

`time-machine` patches Python's time functions for the whole process while the case runs. Anything in your app that reads the time sees the case's time, including token expiry checks and request signing in client libraries. If a model client misbehaves with a clock set in the past or future, set `clock` only on the cases that need it.

## Trial results [#trial-results]

Each trial's result records the `clock` it ran with.


---

# Instrument with OpenTelemetry

> Collect each trial's spans for the trace view, step blame and token cost.

Source: https://mitej23.github.io/docs/guides/opentelemetry



Install the `otel` extra and Verdict keeps the OpenTelemetry spans your app emits during each trial. You get a trace per trial, the step behind each failed check, and token cost.

## Install [#install]

```bash
pip install "verdict-evals[otel]"
```

That's all Verdict needs. When a run starts, it adds a span processor to the global tracer provider, creating an SDK `TracerProvider` if your app hasn't set one.

## If your framework already emits GenAI spans [#if-your-framework-already-emits-genai-spans]

Frameworks with OpenTelemetry GenAI instrumentation need no changes: pydantic-ai through Logfire, the OpenAI Agents SDK, and LangGraph through OpenInference. Make sure their instrumentation is enabled in the code path your entry point runs.

## If it doesn't [#if-it-doesnt]

Emit spans yourself with the GenAI semantic conventions. The refunds example does this so its trace has an agent and tool steps:

```python title="refunds_app/agent.py"
from opentelemetry import trace

tracer = trace.get_tracer("refunds_app")

def handle(request: dict) -> dict:
    with tracer.start_as_current_span(
        "invoke_agent refund_agent",
        attributes={"gen_ai.operation.name": "invoke_agent", "gen_ai.agent.name": "refund_agent"},
    ):
        return _handle(request)

# around each tool call
with tracer.start_as_current_span(
    "execute_tool refund",
    attributes={
        "gen_ai.operation.name": "execute_tool",
        "gen_ai.tool.name": "refund",
        "gen_ai.tool.call.arguments": '{"order_id": "A100", "amount": 24.0}',
    },
):
    PaymentsClient().refund(order.id, order.total)
```

For a model call, set `gen_ai.operation.name = "chat"`, the model in `gen_ai.request.model` or `gen_ai.response.model`, token counts in `gen_ai.usage.input_tokens` and `gen_ai.usage.output_tokens`, and the messages in `gen_ai.input.messages` and `gen_ai.output.messages`.

## What you get [#what-you-get]

* **The trace.** The trial's **Trace** tab shows agents, model calls (each labelled with its decision), tools and events, with timing.
* **Blame.** Effects recorded inside a span carry its id, so a failed effect check points at the tool step that made the call.
* **Cost.** Model-call spans are priced per trial; compare shows cost per trial and `--max-cost-increase` can gate on it.

## Name your agents [#name-your-agents]

```toml title="verdict.toml"
[trace]
agent_names = { refund_agent = "Refund agent" }
```

Without a name, `refund_agent` is shown as "Refund agent" anyway; use `agent_names` when the span name isn't readable.

## Send traces to your backend [#send-traces-to-your-backend]

Spans are always kept locally with each trial. To also see eval traces where you see production traces (Logfire, Langfuse, Honeycomb, an OpenTelemetry collector), add an OTLP/HTTP endpoint:

```bash
pip install "verdict-evals[otel-export]"
```

```toml title="verdict.toml"
[trace.export]
endpoint = "https://logfire-api.pydantic.dev/v1/traces"
service = "refunds-evals"
headers = { Authorization = "env:LOGFIRE_EVALS_TOKEN" }
```

* Exported spans carry `deployment.environment = "eval"`, so filter them out of production dashboards, or send them to a separate project.
* `"env:NAME"` reads a header from the environment; a missing variable stops the run with a clear error. Keep tokens out of the file.
* The endpoint's host is allowed through the [network guard](/docs/concepts/isolation-and-network-guard) automatically; every other host stays blocked.
* Spans still queued are flushed when the run ends.

Full keys: [`[trace.export]`](/docs/reference/verdict-toml#traceexport).

## Price your models [#price-your-models]

Verdict ships prices for a few models. Add yours in a project `pricing.toml`, see [Keep costs down](/docs/guides/keep-costs-down#pricing). A model without a price makes the trial's cost unknown rather than wrong.

See [Traces and observations](/docs/concepts/traces-and-observations).


---

# Write a generator spec

> Turn your product rules into a generator spec, using the refunds example, and generate gated cases.

Source: https://mitej23.github.io/docs/guides/write-a-generator-spec



A generator spec tells `verdict generate` which scenarios exist and what is correct in each. This guide walks through the refunds example's spec in `examples/refunds/generator/`.

<Steps>
  <Step>
    ### Write the rules in words [#write-the-rules-in-words]

    `rules.md` is your product policy in plain language. The writer and the reviewer both read it.

    ```markdown title="generator/rules.md"
    # Refund policy

    1. **Unknown order:** say so and ask for the order number. Nothing else happens.
    2. **Not delivered yet** (status is not `delivered`): there is nothing to refund. Say it's on
       its way. No refund, no email.
    3. **The 30-day window:** a refund is possible only within 30 days of delivery, at the
       clock's time. Outside it, explain the window. No refund, no email, the order stays
       `delivered`.
    4. **Large orders:** over $200.00, the agent never refunds automatically. It marks the order
       `escalated`, emails support@shop.example with the order and amount, and tells the
       customer a person will review it within a day.
    5. **Refund:** otherwise it refunds the full order total through the payments provider, marks
       the order `refunded`, emails the customer a confirmation, and says the refund is on its way.
    6. **Payments failure:** if the provider fails, the refund did not happen. Never say it did.

    The customer's tone never changes the outcome.
    ```
  </Step>

  <Step>
    ### Name the dimensions [#name-the-dimensions]

    Each dimension is something that varies between situations, with a relative weight per value.

    ```toml title="generator/generator.toml"
    schema = 1
    module = "scenarios.py"
    rules = "rules.md"
    example = "../cases/refund-within-window"
    prefix = "gen"
    id_from = "situation"              # gen-<situation>-<hash>

    [model]
    write = "sonnet"
    review = "sonnet"

    [budget]
    usd = 2.0
    per_call = 0.40
    concurrency = 4

    [dimensions.situation]
    delivered_recently = 4             # within the 30-day window
    delivered_long_ago = 2             # outside it
    near_the_deadline = 1              # 29 days: still inside
    still_in_transit = 2

    [dimensions.order_size]
    small = 4                          # at or under the $200 automatic-refund limit
    large = 1                          # over it: a person decides

    [dimensions.payments]
    ok = 4
    down = 1

    [dimensions.tone]
    polite = 3
    angry = 2
    terse = 2
    ```

    Include the boundaries that matter (29 days, $200) as their own values: those are where agents break.
  </Step>

  <Step>
    ### Decide correctness in plan() [#decide-correctness-in-plan]

    `plan(values, rng)` is the rules in code. It fixes the facts and writes the checks. Claude never writes checks.

    ```python title="generator/scenarios.py"
    def plan(v, rng):
        outcome = _outcome(v)                      # not_delivered | outside_window | escalated | refunded | payment_failed
        order_id = f"G{rng.randint(100, 999)}"
        # ... pick a customer, an item and a total over or under $200, a delivery date ...
        facts = {"order_id": order_id, "customer_email": f"{who}@example.com", "item": item, "total": total,
                 "status": "shipped" if days is None else "delivered", "delivered_at": ...}

        decision = {"kind": "output", "path": "decision", "equals": outcome}
        if outcome == "refunded":
            checks = [{"kind": "effect", "service": "payments", "action": "refund",
                       "args": {"order_id": order_id, "amount": total}, "count": 1},
                      {"kind": "effect", "service": "email", "action": "send_email",
                       "args": {"to": facts["customer_email"]}, "count": 1},
                      {"kind": "state", "table": "orders", "where": {"id": order_id}, "equals": {"status": "refunded"}},
                      decision]
        # ... one branch per outcome ...
        return {"facts": facts, "checks": checks, "outcome": words, "known_bad": bad,
                "clock": "2026-10-08T10:00:00+00:00", "tags": [outcome.replace("_", "-"), v["tone"]],
                "decision": outcome}
    ```

    * `facts` are values the case must use; the facts-used gate rejects content that drops one.
    * `outcome` and `known_bad` are sentences for the writer and the reviewer.
    * Extra keys (`decision` here) are passed on to your other hooks.
  </Step>

  <Step>
    ### Collapse combinations that can't matter [#collapse-combinations-that-cant-matter]

    `normalise()` blanks out dimensions that don't change the outcome, so "in transit, large order, payments down" and "in transit, small order, payments ok" become one scenario.

    ```python title="generator/scenarios.py"
    NA = "n/a"

    def normalise(v):
        v = dict(v)
        if v["situation"] == "still_in_transit":
            v["order_size"] = v["payments"] = NA
        elif v["situation"] == "delivered_long_ago":
            v["order_size"] = v["payments"] = NA     # the window is checked before the amount
        elif v["order_size"] == "large":
            v["payments"] = NA                       # a person decides; no refund is attempted
        return v
    ```
  </Step>

  <Step>
    ### Ask Claude for less [#ask-claude-for-less]

    By default Claude writes a whole case. The refunds spec asks only for the parts that need language, with `CONTENT_SCHEMA`, and assembles the rest in `build()`:

    ```python title="generator/scenarios.py"
    CONTENT_SCHEMA = {
        "type": "object",
        "additionalProperties": False,
        "required": ["title", "message", "good_reply", "bad_reply", "known_bad_failure"],
        "properties": {
            "title": {"type": "string", "maxLength": 90},
            "message": {"type": "string", "description": "the customer's refund request, in the scenario's tone, mentioning the item"},
            "good_reply": {"type": "string", "description": "the agent's correct reply under the policy"},
            "bad_reply": {"type": "string", "description": "what a broken agent would reply"},
            "known_bad_failure": {"type": "string", "maxLength": 160},
        },
    }

    def build(v, plan, content):
        f, decision = plan["facts"], plan["decision"]
        return {
            "title": content["title"],
            "input": {"order_id": f["order_id"], "message": content["message"]},
            "world": {"tables": {"orders": [_order(f)]}, "services": {...}},
            "oracle": {"output": {"decision": decision, "reply": content["good_reply"]}, "effects": ..., "state": ...},
            "known_bad": {"output": {"decision": BAD[decision], "reply": content["bad_reply"]}, "effects": ..., "state": ...},
        }
    ```

    Smaller content is cheaper and leaves less for the model to get wrong. The world, the effects and the final state come from code.
  </Step>

  <Step>
    ### Add your own gate [#add-your-own-gate]

    `validate()` returns a list of problems; any problem fails the attempt and is fed back to the writer.

    ```python title="generator/scenarios.py"
    def validate(v, plan, case):
        issues = []
        if plan["facts"]["item"] not in case["input"]["message"].lower():
            issues.append(f"the customer's message should mention the {plan['facts']['item']}")
        return issues
    ```
  </Step>

  <Step>
    ### Add a free stand-in [#add-a-free-stand-in]

    `fake(values, plan, rng)` returns content in the shape of `CONTENT_SCHEMA` without calling Claude. It makes `--dry-run` work, and lets `--estimate` size the output.

    ```bash
    verdict generate --dry-run -n 3
    ```

    ```text
    [1/3] case     gen-delivered-recently-ca132a  $0.000
    [2/3] case     gen-delivered-long-ago-c4a8f0  $0.000
    [3/3] case     gen-near-the-deadline-dbbf7b  $0.000

    3 new case(s), 0 rejected, 0 skipped (budget); 0 needed a second attempt. $0.000 over 0 calls, 0.0 min.
    Run them: verdict run --config verdict.toml -k 2
    ```

    Dry-run cases go through every gate except the reviewer and are written to `cases/`. Delete them before a real batch if you don't want them.
  </Step>

  <Step>
    ### Estimate, then generate [#estimate-then-generate]

    ```bash
    verdict generate --estimate -n 25
    ```

    ```text
    25 cases, assuming 30% need a second attempt (API-equivalent prices):
      write   sonnet     32 calls  $0.17 with prompt caching, $0.34 without
      review  sonnet     30 calls  $0.25 with prompt caching, $0.36 without
      total            $0.42 ($0.017 a case); $0.70 uncached
      tokens per write call: 1069 system + 303 prompt -> ~291 out
    Through a Claude Code subscription these calls are not billed per token.
    ```

    Then generate for real. This needs the `claude` CLI and a Claude Code login.

    ```bash
    verdict generate -n 25 --budget 1
    ```

    A 3-case batch on this spec cost $0.068 API-equivalent over 6 calls in 18 seconds.
  </Step>

  <Step>
    ### Review what came out [#review-what-came-out]

    ```bash
    ls cases/_rejected/          # scenarios that failed twice, with every attempt's issues
    verdict self-test            # generated cases are self-tested like any other
    verdict run -k 2
    ```

    Read a few generated cases before trusting a batch. `generation.json` in each case records Claude's content, the gates' issues and the reviewer's answer.
  </Step>
</Steps>

## Run it again to fill gaps [#run-it-again-to-fill-gaps]

Each batch counts the generated cases already in the folder and favours under-represented values, so repeated runs of `-n 10` converge on the weights you set. Change `--seed` to draw different scenarios.

See [Generating cases](/docs/concepts/generating-cases), [generator.toml](/docs/reference/generator-toml) and [Generator hooks](/docs/reference/generator-hooks).


---

# Compare a candidate

> Run the same cases against another version of your code, from a folder or a git ref, and compare.

Source: https://mitej23.github.io/docs/guides/compare-a-candidate



A candidate is another version of your app's code: a branch, a commit, a folder with an experimental prompt. You run the same cases against it and compare with a baseline.

## Pick how to point at the candidate [#pick-how-to-point-at-the-candidate]

<Tabs items="['A git ref (--code)', 'A folder (--app-root)']">
  <Tab value="A git ref (--code)">
    ```bash
    verdict run --code my-branch -k 3 --label "my-branch"
    ```

    Verdict resolves the ref to a commit, creates a detached worktree at `.verdict/worktrees/<sha>`, and runs the app from the same relative path inside it. Your working copy is untouched, and the worktree is reused next time. Needs git.
  </Tab>

  <Tab value="A folder (--app-root)">
    ```bash
    verdict run --app-root ../my-app-feature -k 3 --label "feature"
    ```

    Runs the app from that directory instead of `[app].root`. Use it for a worktree you made yourself, a copy, or a folder of variants, like `examples/refunds/candidate`.
  </Tab>
</Tabs>

Either way, the cases, checks, store and boundaries come from your `verdict.toml`. Only the code changes. The `paths` you configured must exist in the candidate.

## Run the baseline the same way [#run-the-baseline-the-same-way]

```bash
verdict run -k 3 --label baseline
```

Use the same `-k` and the same cases on both sides. Compare only pairs cases present in both runs.

## Compare [#compare]

```bash
verdict runs
verdict compare BASELINE_RUN CANDIDATE_RUN
```

Pool several runs per side with commas. Two baseline runs of the same code give more baseline trials:

```bash
verdict compare RUN_A,RUN_B RUN_C
```

Compare a subset of cases:

```bash
verdict compare BASELINE_RUN CANDIDATE_RUN --cases refund-within-window,payments-down
```

## Iterate cheaply [#iterate-cheaply]

1. While iterating, run only the cases your change targets, at a low `k`:

   ```bash
   verdict run --code my-branch --cases refund-within-window -k 2
   ```

2. When a case comes back `watch`, rerun it with more trials (`-k 10`) before deciding.

3. Run the full suite once before merging, and judge on holdout cases.

## Know your noise floor [#know-your-noise-floor]

Run the baseline twice and compare the runs. Same fingerprint, so Verdict reports the noise floor, the size of difference you'd see with no change at all:

```bash
verdict run -k 3 --label "baseline again"
verdict compare BASELINE_RUN BASELINE_AGAIN_RUN
# Verdict: A/A: same code on both sides, so this difference is the noise floor
```

## In the web UI [#in-the-web-ui]

**Code** lists `main`, every candidate folder a run has used, and every git worktree, each with its diff against `main` and the runs that tested it. From a version's page you can compare its run with a baseline. See [Web UI](/docs/web-ui#code).

See [Code versions and fingerprints](/docs/concepts/code-versions) and [Compare and the gate](/docs/concepts/compare-and-the-gate).


---

# Gate CI on compare

> Fail a CI job when a change regresses a case, worsens one beyond noise, or raises cost too much.

Source: https://mitej23.github.io/docs/guides/gate-ci



`verdict compare --gate` exits 1 when the candidate regresses, so any CI system can block a merge on it.

```bash
verdict compare "$BASELINE" "$CANDIDATE" --gate --max-cost-increase 0.2
```

## What fails the gate [#what-fails-the-gate]

* a case **regressed**: it passed every baseline trial and fails two or more candidate trials, or its pass rate halved;
* a case is **worse beyond noise**: Holm-adjusted p \< 0.05 with a negative change;
* the overall 95% interval is **entirely below zero**;
* **cost per trial** rose by more than `--max-cost-increase` (a fraction: `0.2` means 20%);
* the two sides share **no cases**.

An A/A comparison (same code fingerprint on both sides) never fails the gate. See [Compare and the gate](/docs/concepts/compare-and-the-gate).

## Exit codes [#exit-codes]

| Command                  | Exit 0                                            | Exit 1                                                             |
| ------------------------ | ------------------------------------------------- | ------------------------------------------------------------------ |
| `verdict compare --gate` | the gate passed, or both sides ran the same code  | the gate failed; an unknown run id                                 |
| `verdict compare`        | always, once it has compared                      | an unknown run id                                                  |
| `verdict run`            | the run finished, whether trials passed or failed | a grader failed its self-test; a config or case error              |
| `verdict self-test`      | every grader is right                             | any case's checks fail its known-good or pass its known-bad result |

`verdict run` doesn't fail on failing trials. Failing trials are results; the gate is where you decide.

## A CI job [#a-ci-job]

Verdict doesn't ship a CI integration yet. A job needs to run both sides and compare. Shown here as a GitHub Actions job:

```yaml title=".github/workflows/verdict.yml"
name: verdict
on: pull_request

jobs:
  verdict:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
        with:
          fetch-depth: 0
      - uses: astral-sh/setup-uv@v5
      - run: uv sync
      - name: Baseline (the base branch, in its own worktree)
        run: uv run verdict run --code origin/${{ github.base_ref }} -k 3 --label baseline
      - name: Candidate (this checkout)
        run: uv run verdict run -k 3 --label candidate
      - name: Gate
        run: |
          base=$(uv run verdict runs | awk '$NF=="baseline"{print $1; exit}')
          cand=$(uv run verdict runs | awk '$NF=="candidate"{print $1; exit}')
          uv run verdict compare "$base" "$cand" --gate --max-cost-increase 0.2
```

* `verdict runs` lists newest first, with the label last, so the job picks the two runs it just made.
* `--code` needs the base branch's history: `fetch-depth: 0`.
* If your agent calls a real model, put the key in a secret and list the model host in `[network].allow`.
* For the JSON verdict, use `verdict compare ... --json`.

## Make the gate trustworthy [#make-the-gate-trustworthy]

* **Run the self-test first.** `verdict run` refuses to start when a grader is wrong; don't pass `--skip-self-test` in CI.
* **Use enough trials.** At `k = 1`, one unlucky failure looks like a regression. Use at least 2; `watch` flags tell you which cases need more.
* **Keep a holdout.** Cases in `holdout.json` you haven't tuned against give the honest answer.

## Cost in CI [#cost-in-ci]

Cost needs [OpenTelemetry spans](/docs/guides/opentelemetry) with token usage and a price for each model. Without them, cost per trial is 0 or unknown, and `--max-cost-increase` has nothing to compare.

See [Keep costs down](/docs/guides/keep-costs-down).


---

# Keep costs down

> Spend model calls only where they change a decision, and price your models so cost is visible.

Source: https://mitej23.github.io/docs/guides/keep-costs-down



Agent evals cost money: every trial is a full run of your agent. Spend calls where they change a decision.

## Cheapest proof first [#cheapest-proof-first]

1. **Prove code fixes for free.** If a fix is in code (a date calculation, a tool's arguments), replay recorded tool calls through the changed function in a unit test before running any model.

2. **Fix the grader for free.** Change checks, then `verdict self-test` and `verdict regrade`. Neither calls a model or reruns your app.

3. **Run targeted cases at a low k.** Only the cases your change is about, at `k = 2`:

   ```bash
   verdict run --cases refund-within-window,payments-down -k 2
   ```

4. **Run the full suite once**, before merging.

Never run the whole suite at a high `k` while iterating.

## Choose trials per case [#choose-trials-per-case]

* `[run].trials` in `verdict.toml` sets the default (2 if unset).
* `trials` in a `case.toml` raises it for a noisy case.
* `-k` overrides both for one run.
* A `watch` flag in compare means one new failure: rerun that case with `-k 10`, not the suite.

## Watch the cost [#watch-the-cost]

With [OpenTelemetry](/docs/guides/opentelemetry) and prices for your models, every trial has a cost:

* `verdict compare` prints mean cost per trial on each side and the change;
* `--max-cost-increase 0.2` fails the gate when cost per trial rises more than 20%;
* the web UI shows cost per trial on the Cases page and per run.

## Pricing [#pricing]

Verdict ships a `pricing.toml` with a few models. Prices are USD per million tokens, keyed by model name after any `provider/` prefix (`openai/gpt-5.1` is looked up as `gpt-5.1`).

Add or override models in your project:

```toml title="verdict.toml"
[pricing]
path = "pricing.toml"
```

```toml title="pricing.toml"
[models."gpt-5.1"]
input = 1.25            # USD per million input tokens
cached_input = 0.125    # per million cached input tokens; defaults to input
output = 10.0           # per million output tokens

[models."my-fine-tune"]
input = 3.0
output = 12.0
```

Your file is merged over the bundled one, model by model. Check prices against your provider's price page; the bundled ones are defaults, not read from a bill.

Cost per model call is `(input − cached) × input + cached × cached_input + output × output`, divided by a million. A model with no price makes the trial's cost unknown (shown as `None`/—) rather than wrong.

## Generating cases [#generating-cases]

`verdict generate` runs through Claude Code, so calls are not billed per token on a subscription; it still reports API-equivalent cost.

```bash
verdict generate --estimate -n 25     # no calls
verdict generate -n 25 --budget 1     # stop starting new cases past $1
```

See [Generating cases](/docs/concepts/generating-cases#cost).


---

# Build an eval for your agent, end to end

> From an app with no evals to a gated CI comparison, step by step, using the refunds agent as the app.

Source: https://mitej23.github.io/docs/tutorial



This tutorial takes an agent with no evals to a suite that catches regressions in CI. It rebuilds the refunds example's setup from scratch, so every file matches something you can run. Replace the refunds code with your own as you go.

You'll end with a `verdict.toml`, a handful of cases with proven graders, a baseline and its noise floor, and a gated comparison.

<Steps>
  <Step>
    ### Step 0: What you need [#step-0-what-you-need]

    * A Python 3.11+ app with an agent that acts on something: writes rows, calls services, sends messages.

    * Verdict installed with the extras your app needs (see [Installation](/docs/installation)):

      ```bash
      uv add "verdict-evals[sqlalchemy,clock,otel] @ git+https://github.com/mitej23/verdict"
      ```

    * One real situation your agent must get right, ideally a past incident.

    The app here is a support agent. `handle({"order_id", "message"})` looks up the order, then refunds it, escalates it, or explains why not:

    ```python title="refunds_app/agent.py"
    REFUND_WINDOW = timedelta(days=30)
    AUTO_REFUND_LIMIT = 200.0

    def handle(request: dict) -> dict:
        with SessionLocal() as db:
            order = db.get(Order, request["order_id"])
            ...
            PaymentsClient().refund(order.id, order.total)
            order.status = "refunded"
            db.commit()
            send_email(order.customer_email, "Your refund", ...)
            return {"decision": "refunded", "reply": f"Done: your refund of ${order.total:.2f} is on its way."}
    ```
  </Step>

  <Step>
    ### Step 1: Scaffold [#step-1-scaffold]

    ```bash
    verdict init
    ```

    This writes `verdict.toml` and `cases/example/`. Everything below edits those.
  </Step>

  <Step>
    ### Step 2: Point at the entry point [#step-2-point-at-the-entry-point]

    Find the function production calls and name it. Set `paths` to the code you want fingerprinted.

    ```toml title="verdict.toml"
    schema = 1

    [app]
    root = "."
    paths = ["refunds_app"]

    [entry.default]
    target = "refunds_app.agent:handle"
    ```
  </Step>

  <Step>
    ### Step 3: Give each trial its own database [#step-3-give-each-trial-its-own-database]

    ```toml title="verdict.toml"
    [state]
    adapter = "sqlalchemy"
    metadata = "refunds_app.db:Base.metadata"
    session = "refunds_app.db:SessionLocal"
    ```

    From now on, `SessionLocal()` inside a trial opens an in-memory SQLite database built from the case's world. Your real database is never touched.
  </Step>

  <Step>
    ### Step 4: Contain the services [#step-4-contain-the-services]

    Run once with no boundaries and an empty allow-list, and let the network guard find your services:

    ```toml title="verdict.toml"
    [network]
    allow = []
    ```

    ```bash
    verdict run -k 1
    ```

    Every outbound call now raises `NetworkBlocked` inside your app, naming the host: `The app tried to reach 'payments.example.com'…`. Where the app doesn't catch it, it becomes the trial's error. Where it does (the refunds agent catches payment failures and answers politely), the trial's output shows the fallback instead, which is worth knowing too. Declare each service through the class or function the app calls it with:

    ```toml title="verdict.toml"
    [[boundary]]
    service = "payments"
    patch = "refunds_app.services:PaymentsClient"

    [[boundary]]
    service = "email"
    patch = "refunds_app.services:send_email"
    ```

    If your agent calls a real model, add its host, e.g. `allow = ["api.openai.com"]`.
  </Step>

  <Step>
    ### Step 5: Write the first case [#step-5-write-the-first-case]

    Turn the incident into a folder. Fix the time, set up the data, and check the outcome.

    ```toml title="cases/refund-within-window/case.toml"
    schema = 1
    id = "refund-within-window"
    title = "Broken mug, delivered 10 days ago: refund it"
    tags = ["refund", "happy-path"]
    clock = "2026-10-08T10:00:00+00:00"

    [input]
    order_id = "A100"
    message = "My mug arrived broken, can I get my money back?"

    [[check]]
    kind = "effect"
    service = "payments"
    action = "refund"
    count = 1
    args = { order_id = "A100", amount = 24.0 }

    [[check]]
    kind = "state"
    table = "orders"
    where = { id = "A100" }
    equals = { status = "refunded" }

    [[check]]
    kind = "effect"
    service = "email"
    action = "send_email"
    count = 1
    args = { to = "maya@example.com" }

    [[check]]
    kind = "output"
    path = "reply"
    contains = "refund"
    ```

    ```json title="cases/refund-within-window/world.json"
    {"schema": 1,
     "tables": {"orders": [{"id": "A100", "customer_email": "maya@example.com", "status": "delivered", "total": 24.0, "delivered_at": "2026-09-28T12:00:00"}]},
     "services": {"payments": {"refund": {"returns": {"refund_id": "R-1", "status": "ok"}}}}}
    ```

    Delete `cases/example`, then run it:

    ```bash
    verdict run -k 2
    ```
  </Step>

  <Step>
    ### Step 6: Prove the grader [#step-6-prove-the-grader]

    Write what a correct agent does and what a broken one does, as results:

    ```json title="cases/refund-within-window/oracle.json"
    {"output": {"decision": "refunded", "reply": "Done: your refund of $24.00 is on its way."},
     "effects": [
       {"seq": 1, "service": "payments", "action": "refund", "args": {"order_id": "A100", "amount": 24.0}, "ok": true, "result": {"refund_id": "R-1"}},
       {"seq": 2, "service": "email", "action": "send_email", "args": {"to": "maya@example.com", "subject": "Your refund", "body": "..."}, "ok": true, "result": null}],
     "state": {"tables": {"orders": [{"id": "A100", "status": "refunded"}]}}}
    ```

    ```json title="cases/refund-within-window/known_bad.json"
    {"output": {"decision": "outside_window", "reply": "Refunds are available within 7 days of delivery."},
     "effects": [],
     "state": {"tables": {"orders": [{"id": "A100", "status": "delivered"}]}}}
    ```

    ```bash
    verdict self-test
    # ok   refund-within-window: known-good 1.0, known-bad 0.25
    ```

    The checks pass the good result and catch the bad one. From now on `verdict run` refuses to start if they stop doing so.
  </Step>

  <Step>
    ### Step 7: Cover the other outcomes [#step-7-cover-the-other-outcomes]

    One case per rule, at least. The refunds example has five: within the window, outside it, not delivered, a large order that escalates, and a payments failure. The failure case scripts the provider to fail:

    ```json title="cases/payments-down/world.json"
    {"schema": 1,
     "tables": {"orders": [{"id": "A104", "customer_email": "sam@example.com", "status": "delivered", "total": 80.0, "delivered_at": "2026-10-05T10:00:00"}]},
     "services": {"payments": {"refund": {"raises": "503 provider unavailable"}}}}
    ```

    and checks the agent doesn't claim a refund happened:

    ```toml title="cases/payments-down/case.toml"
    [[check]]
    kind = "output"
    path = "reply"
    not_matches = "refund (of|is) .* on its way|refunded"
    ```

    For volume, write a [generator spec](/docs/guides/write-a-generator-spec) and let `verdict generate` fill in the combinations.
  </Step>

  <Step>
    ### Step 8: Set a holdout [#step-8-set-a-holdout]

    Keep some cases out of your iteration loop:

    ```json title="cases/holdout.json"
    {"holdout": ["payments-down"]}
    ```
  </Step>

  <Step>
    ### Step 9: Run a baseline and measure the noise [#step-9-run-a-baseline-and-measure-the-noise]

    ```bash
    verdict run -k 3 --label baseline
    verdict run -k 3 --label "baseline again"
    verdict runs
    verdict compare BASELINE_RUN BASELINE_AGAIN_RUN
    ```

    Same code on both sides, so compare reports the noise floor. A deterministic app like this one doesn't move; an LLM agent will. Note how much your cases move with no change: that's the size of difference you can't trust.
  </Step>

  <Step>
    ### Step 10: Make a change and run it [#step-10-make-a-change-and-run-it]

    Change the code on a branch. Here, the refund window shrinks to 7 days:

    ```python title="refunds_app/agent.py"
    REFUND_WINDOW = timedelta(days=7)
    ```

    Commit it on a branch and run the branch without leaving yours:

    ```bash
    verdict run --code shorter-window -k 3 --label candidate
    ```

    Or keep the change in a separate folder and use `--app-root`, as `examples/refunds/candidate` does.
  </Step>

  <Step>
    ### Step 11: Compare [#step-11-compare]

    ```bash
    verdict compare BASELINE_RUN CANDIDATE_RUN --gate
    ```

    ```text
      Cases that changed:
        regressed refund-within-window                     3/3 → 0/3  (dev, p=0.1, adjusted 0.5)

    Verdict: fails the gate
      - 1 case(s) regressed: refund-within-window
    ```

    The overall interval crossed zero, so an average would have shipped this. The per-case flag caught it, and `--gate` exits 1.
  </Step>

  <Step>
    ### Step 12: See why [#step-12-see-why]

    ```bash
    verdict serve
    ```

    Open the failing trial. **Outcome** lists the three failed checks and points them at the refund agent. **Data changes** shows the order stayed `delivered`. **Trace** shows no refund tool step was ever made. See [Web UI](/docs/web-ui).
  </Step>

  <Step>
    ### Step 13: Gate CI [#step-13-gate-ci]

    Run both sides in CI and fail the job on the gate:

    ```bash
    verdict run --code origin/main -k 3 --label baseline
    verdict run -k 3 --label candidate
    verdict compare "$BASE" "$CAND" --gate --max-cost-increase 0.2
    ```

    See [Gate CI on compare](/docs/guides/gate-ci) for a complete job.
  </Step>
</Steps>

## Where to go next [#where-to-go-next]

* [Write checks](/docs/guides/write-checks) that survive better wording.
* [Instrument with OpenTelemetry](/docs/guides/opentelemetry) for traces and cost.
* [Keep costs down](/docs/guides/keep-costs-down) once your agent calls real models.


---

# Web UI

> A tour of the local web app for one project, from cases and runs to a single trial's trace and a code diff.

Source: https://mitej23.github.io/docs/web-ui



`verdict serve` opens a local web app for one project: its cases, runs, each trial's outcome and trace, comparisons and code versions. It reads the same store and case folders as the CLI, so both always agree on every number.

```bash
verdict serve              # http://127.0.0.1:8765
verdict serve --port 8791
verdict view               # the same, an alias
```

It needs the `web` extra and a one-time build of the React client from a checkout of the repository:

```bash
pip install "verdict-evals[web]"
cd web && npm install && npm run build
```

The server listens on 127.0.0.1 only and rejects requests from other origins. Runs started from the UI are separate `verdict run` processes, so stopping the server never stops a run. One run at a time.

Screenshots are of the refunds example after a baseline run and a run of the 7-day-window candidate.

## Cases [#cases]

The home page lists every case with its latest status, recent runs, pass rate, last run and cost per trial. **Run all cases** starts a run with the number of trials you pick; each row can run one case.

The run form shows how many trials the run takes and its estimated cost and time, from past trials of the same cases (`verdict run --estimate`). Multi-turn cases are tagged **conversation**.

For generated cases, **Coverage** counts cases per scenario value, so you can see which situations are thin. Cases that `verdict generate --pilot` set aside because they failed every pilot trial are listed under **To review**, with the reason, until a person moves them back or fixes them.

![Cases](https://mitej23.github.io/screenshots/cases.png)

*Every case, its recent results, and the button that starts a run.*

## Case [#case]

One case, in four tabs:

* **Overview:** the case's description, whether every check passed in the latest run, a matrix of recent results per check, and the latest (or a failing) trial against the known-good result. A banner says when the grader is wrong.
* **World:** the input, the clock, the tables each trial starts from, and the scripted services. A multi-turn case shows its **User** first: the scripted turns (with their `when` conditions), or the scenario a simulated user plays.
* **Checks:** the grader self-test as a table, every check against the known-good and the known-bad result (✓ or ✕ with the reason), the two results side by side, each check's kind and spec, and the raw `case.toml`, `world.json`, `oracle.json`, `known_bad.json` and `checks.py`.
* **Runs:** this case in every run that included it, with each trial's result and, for failures, the step behind them.

![Case: Checks](https://mitej23.github.io/screenshots/checks.png)

*The grader table: every check passes the known-good result; three fail the known-bad one, each with its reason.*

![Case: Overview](https://mitej23.github.io/screenshots/case.png)

*refund-within-window: the check matrix shows the candidate runs failing three checks; the latest run passes and matches the known-good result.*

## Runs [#runs]

Every run, newest first: label, status, cases passing every trial, mean reward, the change in mean reward against the run before it, trials passed (and how many errored), duration, cost, and the code version and commit it tested, tagged "uncommitted" when the checkout had changes. **Re-grade all runs** re-scores stored trials with the current checks, without model calls.

A run in progress shows a live banner with its progress on every page.

![Runs](https://mitej23.github.io/screenshots/runs.png)

*Every run with its mean reward, the change against the run before it, cost, and the code version and commit it tested.*

## Run [#run]

One run's results per case (trials, failing checks, reward, change against the previous run, average time, cost), its totals (mean reward, errored trials, tokens in and out, the models called), what it recorded (code fingerprint, git state, settings), and **Compare with** to pick a baseline. A running run can be cancelled; a failed or cancelled one shows the last lines of its log. If the app reached a host that nothing fakes, the run says which, and how to fix it.

![Run](https://mitej23.github.io/screenshots/run.png)

*The candidate run: mean reward 0.85; refund-within-window fails three checks and drops 0.75 against the previous run.*

## Trial [#trial]

One trial, with its siblings in the same run one click away. The summary line gives checks failed, reward, calls to services, model and tool calls, tokens, cost and time, and the models it called. Four tabs:

* **Outcome:** a multi-turn case starts with the **Conversation**, turn by turn, and how it ended. **Where it went wrong** lists each failed check and links it to the step behind it, with the same as one line (the `feedback` that `verdict trial --json` and the Python API give) and a button to copy it. **What the app did** lists the effects. The output and effects are set against the known-good result and against the same trial in the previous run, with a word-level text diff.
* **Trace:** the trial's steps as a tree (agents, model calls labelled with their decision, tools, events) with a timeline. Select a step to see its input, output, model, tokens and cost. Needs the `otel` extra.
* **Data changes:** per table, how many rows were added, changed and removed against the world; changed rows field by field, before and after; added and removed rows in full. Name the tables and the fields that identify a row with [`[ui] tables`](/docs/reference/verdict-toml#ui) in `verdict.toml`; by default rows match on `id`, `uuid`, `key` or `pk`.
* **Raw:** the full result the checks saw, as JSON.

![Trial: Outcome](https://mitej23.github.io/screenshots/trial.png)

*A failing candidate trial: the refund and email never happened, and the order stayed delivered.*

![Trial: Data changes](https://mitej23.github.io/screenshots/state.png)

*The Orders table (named in [ui] tables): order A100 went from delivered to refunded.*

![Trial: Trace](https://mitej23.github.io/screenshots/trace.png)

*payments-down: the refund tool call fails with the provider's 503, and the agent handles it.*

## Compare [#compare]

Pick two runs. The page shows the verdict and the gate's reasons (including cases graded by different checks on each side), cases left out because their input or world changed, dependency changes between the sides, trials where the app raised, a summary by split (all, dev, holdout), cost and time per trial, and a row per case with its flag, pass counts and adjusted p. **How this is decided** explains the method.

![Compare](https://mitej23.github.io/screenshots/compare.png)

*Baseline against the 7-day-window candidate: fails the gate, refund-within-window regressed.*

## Code [#code]

Every version of the app's code that Verdict knows about:

* `main`: the configured app root, with its uncommitted changes;
* every candidate folder a run has used (`--app-root`);
* every git worktree of the repository, including `--code` worktrees.

Each version shows what it's based on, its changes against `main`, its runs and trials passed. A version's page has two tabs: **Changes**, a per-file diff (side by side or unified), and **Runs**, a grid of every run of that version by case, marking runs whose code has since changed.

![Code: candidate](https://mitej23.github.io/screenshots/code-candidate.png)

*The candidate's one-line diff: REFUND_WINDOW from 30 to 7 days.*

## Rank [#rank]

Each attempted fix against one baseline: reward, interval, gate, the cases each improved and regressed, and cost change, best first (the same as `verdict rank`). Open it from Code: **Rank candidates** ranks each candidate's latest run against main's latest run.

## Setup [#setup]

**Project health** is `verdict doctor`, run in its own process: whether the app imports with every fake in place, the cases load, the graders pass their self-test, and which extras and model hosts are set. **The steps** lists using Verdict in order, as the commands that do each step. **Candidate checkouts** lists the worktrees `verdict run --code` made, with a button to remove the ones no run is using (`verdict worktrees --prune`).

## Theme [#theme]

Light, dark or auto (follows the system), from the switch in the sidebar.


---

# Contributing

> Set up the repository, run the tests, and follow the rules that keep Verdict trustworthy.

Source: https://mitej23.github.io/docs/project/contributing



Contributions are welcome. This page covers the setup, the tests and the few rules every change follows.

## Set up [#set-up]

```bash
git clone https://github.com/mitej23/verdict
cd verdict
uv sync --extra dev
uv run pytest -q
```

The test suite needs no API keys and no network. Run one file while you iterate:

```bash
uv run pytest tests/test_compare.py -q
```

Working on the web UI needs Node 20 or later:

```bash
cd web
npm install
npm run build        # writes into src/verdict/web/dist
```

Python-only contributors and CI can run every test without building it.

## The repository [#the-repository]

| Path                    | What it holds                                                                  |
| ----------------------- | ------------------------------------------------------------------------------ |
| `src/verdict/`          | the package; `cli.py` is the entry point                                       |
| `src/verdict/generate/` | the case generator                                                             |
| `src/verdict/web/`      | the JSON API behind `verdict serve`                                            |
| `web/`                  | the React client: the product UI, the landing page and this documentation site |
| `examples/refunds/`     | the deterministic example app used by tests and docs                           |
| `tests/`                | the test suite                                                                 |
| `docs/`                 | internal docs: product, architecture, decisions, research, learnings           |

## Rules [#rules]

* **Verdict never modifies the app under test.** Isolation happens at the edges: the state adapter, boundary fakes, the network guard and the clock. If a case needs the app changed to be testable, that's a design problem to record, not a patch.
* **Fail loud at the boundary.** An outbound call that no fake handles must raise, never pass through silently.
* **The case is the statistical unit, never the trial.** New metrics go through `compare.py`'s conventions: per-case paired differences, Holm-adjusted per-case tests, A/A detection. No averages that pool trials across cases.
* **Formats are a public contract.** Changing `case.toml`, `world.json`, the result JSON or `verdict.toml` needs a decision record and a migration note.
* **No LLM calls in tests.** Tests use `examples/refunds`, which is deterministic, and synthetic spans. The network guard is on during runner tests; a test that needs a host is a bug.
* **Fix bugs test-first.** Write the failing test that reproduces the bug, then fix it.
* **Standard library first.** New runtime dependencies need a decision record. Integrations stay optional extras.
* **Keep the plugin API stable.** Public functions in `boundary.py`, `state.py`, `checks.py` and `runner.py` are used by other people's apps: add keyword arguments with defaults rather than changing positional ones.

## Decisions and progress [#decisions-and-progress]

* Design decisions are numbered files in `docs/decisions/`, one per decision. A decision is never rewritten; a later one supersedes it.
* After a change in behaviour, add an entry to `docs/progress.md` saying what changed and how it was verified.
* `docs/features.json` lists features with a verify step. Mark one passing only after running its verify step.

## The docs site [#the-docs-site]

The site is part of the web app in `web/` (Vite, with Fumadocs for the docs layout). Pages are MDX files in `web/content/docs/`. Every page has a `title` and a one-line `description`. `cd web && npm run dev` serves it with hot reload at `/docs`; `npm run build` also writes the search index, `sitemap.xml`, `llms.txt` and a `.md` copy of every page.


---

# Release policy

> What is stable in Verdict 0.x, how file formats are versioned, and what a release must include.

Source: https://mitej23.github.io/docs/project/release-policy



Verdict is at 0.1.0. This page says what you can rely on before 1.0.

## What is stable [#what-is-stable]

* **File formats** are the public contract. `verdict.toml`, `case.toml`, `world.json`, the trial result JSON and the reward JSON each carry `schema = 1`. A file with a different schema fails loudly rather than being misread.
* **Optional additions** don't bump the schema. `[pricing]`, `[trace]`, `[ui]`, `[generator]`, `oracle.json`, `known_bad.json`, `generation.json` and the generator spec were added as optional parts of schema 1, so existing files stay valid.
* **The store** is additive: columns are added, never dropped.

## What may change [#what-may-change]

* **The Python API.** It's 0.x. The plugin functions in `verdict.boundary`, `verdict.state`, `verdict.checks` and `verdict.runner` are kept stable where possible, by adding keyword arguments with defaults rather than changing positional ones.
* **The web app's JSON API** under `/api` is not a public contract yet.
* **CLI output text.** Script against `--json` (on `compare` and `generate --estimate`) and exit codes, not the human-readable text.

## Format changes [#format-changes]

A change to a file format needs:

1. a decision record explaining why;
2. a migration note in the formats documentation;
3. a schema bump when old files would be misread.

## Releases [#releases]

* Every release gets an entry in the [changelog](/docs/changelog).
* The React client must be built before packaging (`cd web && npm ci && npm run build`, then `uv build`), because the wheel includes the built app.
* Semantic versioning starts with the public launch.

## Status of a public release [#status-of-a-public-release]

Verdict has not been published to PyPI or released publicly yet. Until it is, install from the repository (see [Installation](/docs/installation)).


---

# AI agents

> Read these docs from an AI agent or coding assistant with llms.txt, llms-full.txt and per-page Markdown.

Source: https://mitej23.github.io/docs/project/ai-agents



Every page of this site is also available as plain Markdown, so a coding assistant can read the docs without parsing HTML.

## llms.txt [#llmstxt]

[`/llms.txt`](/llms.txt) is an index of every page, grouped by section, with a one-line description and a link to the page's Markdown.

```bash
curl https://mitej23.github.io/llms.txt
```

## llms-full.txt [#llms-fulltxt]

[`/llms-full.txt`](/llms-full.txt) is the whole site in one Markdown file, in sidebar order. Paste it into a model's context, or point an agent at it.

## One page as Markdown [#one-page-as-markdown]

Add `.md` to any docs URL:

```bash
curl https://mitej23.github.io/docs/concepts/checks.md
```

Each page has a **Copy Markdown** button next to its title, and an **Open** menu to view the Markdown, open it on GitHub, or send it to an assistant.

Diagrams come through as fenced code blocks with the `mermaid` language. Components such as tabs and reference fields stay as MDX tags with their text inside.

## Writing an eval with an assistant [#writing-an-eval-with-an-assistant]

A good prompt for a coding assistant:

```text
Read https://mitej23.github.io/llms-full.txt. Then add Verdict to this app:
find the entry point production calls, the SQLAlchemy session factory, and every external
service client, and write verdict.toml plus one case for <a real incident>.
Don't change the app's code.
```

Ask it to run `verdict self-test` and `verdict run -k 1` and fix what fails. A `NetworkBlocked` error names a service it missed.


---

# License

> Verdict is licensed under the Apache License, Version 2.0.

Source: https://mitej23.github.io/docs/project/license



Verdict is open source under the [Apache License, Version 2.0](https://www.apache.org/licenses/LICENSE-2.0). The full text is in [`LICENSE`](https://github.com/mitej23/verdict/blob/main/LICENSE) at the root of the repository.

In short, you may use, modify and distribute Verdict, including in commercial and closed-source products, provided you:

* keep the license and copyright notices;
* state significant changes you make to its files;
* include the `NOTICE` file, if one is present, in what you distribute.

The license grants a patent license from contributors and provides the software without warranty. This summary is not legal advice; the license text governs.

## Your cases and code [#your-cases-and-code]

The license covers Verdict itself. Your app, your cases, worlds and generator specs are yours; running Verdict against them doesn't change their license.


---

# Reference

> Complete reference for the verdict CLI, every file format, the check kinds and the Python API.

Source: https://mitej23.github.io/docs/reference



Complete, exact reference. For explanations, see [Core concepts](/docs/concepts).

## Command line [#command-line]

<Cards>
  <Card title="CLI" href="/docs/reference/cli">
    `init`, `run`, `compare`, `self-test`, `regrade`, `serve`, `generate`, `runs`: every flag, default and exit code.
  </Card>
</Cards>

## Files [#files]

Every file carries `schema = 1`.

<Cards>
  <Card title="verdict.toml" href="/docs/reference/verdict-toml">
    The project config.
  </Card>

  <Card title="case.toml" href="/docs/reference/case-toml">
    One case: input and checks.
  </Card>

  <Card title="world.json" href="/docs/reference/world-json">
    Starting rows and scripted services.
  </Card>

  <Card title="oracle.json and known_bad.json" href="/docs/reference/oracle-and-known-bad">
    Results that prove the grader.
  </Card>

  <Card title="holdout.json" href="/docs/reference/holdout-json">
    The holdout split.
  </Card>

  <Card title="generator.toml" href="/docs/reference/generator-toml">
    The case generator's spec, and what it writes.
  </Card>

  <Card title="Generator hooks" href="/docs/reference/generator-hooks">
    `plan`

    , 

    `normalise`

    , 

    `build`

    , 

    `validate`

    , 

    `fake`

    .
  </Card>
</Cards>

## Grading [#grading]

<Cards>
  <Card title="Check kinds" href="/docs/reference/checks">
    All seven kinds and their fields.
  </Card>
</Cards>

## Python [#python]

<Cards>
  <Card title="Python API" href="/docs/reference/python-api">
    For fakes, checks, hooks and scripts.
  </Card>

  <Card title="Store schema" href="/docs/reference/store-schema">
    The SQLite results store.
  </Card>
</Cards>


---

# CLI

> Every verdict command, flag, default and exit code.

Source: https://mitej23.github.io/docs/reference/cli



Every command except `init` and `guide` reads a project config, `verdict.toml` in the current directory unless you pass `--config`.

```text
verdict [-h] [--version] COMMAND ...
```

| Command                                             | What it does                                                                                  |
| --------------------------------------------------- | --------------------------------------------------------------------------------------------- |
| [`init`](#verdict-init)                             | set up a project (optionally by reading its code), a generator spec, or a copy of the example |
| [`discover`](#verdict-discover)                     | read the app's code and propose its entry, state, boundaries and model hosts                  |
| [`doctor`](#verdict-doctor)                         | check the config, the app's import, extras, cases and graders                                 |
| [`run`](#verdict-run)                               | run cases, optionally against a candidate                                                     |
| [`compare`](#verdict-compare)                       | did a change help beyond noise, and did it break anything?                                    |
| [`runs`](#verdict-runs)                             | list runs                                                                                     |
| [`show`](#verdict-show)                             | one run, case by case                                                                         |
| [`cases`](#verdict-cases) / [`case`](#verdict-case) | the cases, and one case in full                                                               |
| [`trial`](#verdict-trial)                           | why a trial failed: checks, the step behind each, output, effects, state                      |
| [`trace`](#verdict-trace)                           | a trial's steps: agents, model calls, tool calls                                              |
| [`state`](#verdict-state)                           | a trial's database changes                                                                    |
| [`self-test`](#verdict-self-test)                   | check every grader against its known-good and known-bad results                               |
| [`regrade`](#verdict-regrade)                       | re-score stored trials with the current checks                                                |
| [`serve`](#verdict-serve)                           | the local web app                                                                             |
| [`generate`](#verdict-generate)                     | write cases from the project's generator spec                                                 |
| [`brief`](#verdict-brief)                           | what a coding agent needs to fix a run's failures                                             |
| [`candidate`](#verdict-candidate)                   | attempted fixes as named git worktrees                                                        |
| [`rank`](#verdict-rank)                             | rewards for candidate runs against a baseline, ranked                                         |
| [`lint`](#verdict-lint)                             | check verdict.toml and every case without importing the app                                   |
| [`schema`](#verdict-schema)                         | JSON Schema for a file format                                                                 |
| [`agent`](#verdict-agent)                           | install the skills and AGENTS.md block that let a coding agent set up and run Verdict         |
| [`worktrees`](#verdict-worktrees)                   | the candidate checkouts `run --code` made, and removing them                                  |
| [`guide`](#verdict-guide)                           | how a coding agent should use Verdict                                                         |

## For coding agents [#for-coding-agents]

Every command that reports something takes `--json`: one JSON object on stdout with `"schema": 1`, while progress and errors go to stderr. Wherever a command takes a run id, `latest` is the newest run and `latest~N` the one N runs before it. `verdict guide --install-skill` writes the workflow below as a Claude Code skill; `verdict guide` prints it for any other agent.

```bash
verdict run --label baseline --json
# ... change the app ...
verdict run --label "prompt v2" --json
verdict compare latest~1 latest --gate --json
verdict trial latest CASE_ID --json        # failed checks, the step behind each, effects, state
verdict trace latest CASE_ID 1 --json      # every model call and tool call
```

## Global options [#global-options]

<Fields>
  <Field name="--version">
    Print `verdict 0.1.0` and exit.
  </Field>

  <Field name="-h, --help">
    Show help for `verdict` or for a subcommand (`verdict run --help`).
  </Field>
</Fields>

## Exit codes [#exit-codes]

| Code | Meaning                                                                                                                                                                                                                                                                         |
| ---- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 0    | the command did what it was asked                                                                                                                                                                                                                                               |
| 1    | an error, printed to stderr as one line starting `verdict:` (a missing config, an app that doesn't import, an unknown run or case, a grader that fails its self-test), a failed gate (`compare --gate`), or `doctor` finding a problem. Set `VERDICT_DEBUG=1` for the traceback |
| 2    | invalid arguments (from the argument parser)                                                                                                                                                                                                                                    |

Failing trials never make `verdict run` exit non-zero. Gate on `verdict compare --gate`.

## verdict init [#verdict-init]

Set up a project in a directory.

```bash
verdict init [DIR] [--target MODULE:FUNCTION] [--discover] [--generator] [--example refunds]
```

<Fields>
  <Field name="DIR" type="path" default=".">
    The directory to set up. Created if missing.
  </Field>

  <Field name="--target" type="string">
    Your app's entry point, written into `[entry.default]`. Default: the placeholder `myapp.agent:handle`.
  </Field>

  <Field name="--discover" type="flag">
    Write `verdict.toml` from [`verdict discover`](#verdict-discover)'s proposal instead of the template, and a first case whose `[input]` has the keys the entry point reads.
  </Field>

  <Field name="--generator" type="flag">
    Also write `generator/` (`generator.toml`, `rules.md`, `scenarios.py`): a working spec for [`verdict generate`](#verdict-generate) to edit. Works in a project that already has `verdict.toml`.
  </Field>

  <Field name="--example" type="refunds">
    Copy a complete example project (app, candidate, cases, generator) instead. If DIR isn't empty, the copy goes in `DIR/refunds`.
  </Field>
</Fields>

Writes `verdict.toml`, `cases/example/case.toml` (one `no_error` check) and `cases/example/world.json`, and adds `.verdict/` to `.gitignore`: results, logs and worktrees go there once you run. Exits 1 if `verdict.toml` (or, with `--generator`, `generator/generator.toml`) already exists.

## verdict discover [#verdict-discover]

Read the app's code and propose its config. Nothing is imported or run.

```bash
verdict discover [DIR] [--json]
```

It scans the top-level packages in DIR (default `.`) and proposes:

* **the entry:** top-level functions taking one argument, ranked by name (`handle`, `respond` …) and module (`agent`, `app` …), with the input keys each reads (followed into same-module helpers it passes the input to);
* **state:** a SQLAlchemy `declarative_base()` or `DeclarativeBase`, plus a `sessionmaker`;
* **boundaries:** every public class or function that uses an HTTP, email, cloud or SaaS client (`httpx`, `requests`, `urllib.request`, `boto3`, `stripe`, `twilio` …);
* **network:** model API hosts implied by SDK imports (`openai`, `anthropic` …) and known model URLs in the code.

```text
Scanned refunds_app in /path/to/refunds

  entry      refunds_app.agent:handle   reads order_id
  state      sqlalchemy: refunds_app.db:Base.metadata, refunds_app.db:SessionLocal
  boundary   payments       refunds_app.services:PaymentsClient  (class, urllib.request)
  boundary   email          refunds_app.services:send_email  (function, urllib.request)
  network    no model hosts found
```

`--json` adds every candidate with its file and line, and the proposed `verdict_toml`. Anything it misses shows up on the first run: the network guard refuses the host, and `run` prints `! blocked HOST` with the fix.

## verdict doctor [#verdict-doctor]

Check that the project can run, without running it.

```bash
verdict doctor [--config PATH] [--port N] [--json]
```

Imports the app with the state adapter, boundaries and network guard installed, exactly as a run does, then checks the cases, graders (self-test), extras (clock, traces, web), allowed model hosts, the store, git, the generator spec, the `claude` CLI and the serve port. Each line is `✓` ok, `!` works with something missing, or `✕` a run would fail, with the fix.

**Exit codes:** 0 when nothing is `✕`; 1 otherwise.

## verdict run [#verdict-run]

Run cases and save the results to the store.

```bash
verdict run [--config PATH] [--cases IDS] [-k N] [--label TEXT] [--split dev|holdout]
            [-c N] [--app-root DIR | --code REF] [--skip-self-test]
```

<Fields>
  <Field name="--config" type="path" default="verdict.toml">
    The project config.
  </Field>

  <Field name="--cases" type="ids">
    Comma-separated case ids to run. Default: every case. An unknown id is an error.
  </Field>

  <Field name="-k, --trials" type="int">
    Trials per case. Default: the case's own `trials`, else `[run].trials`.
  </Field>

  <Field name="--label" type="string" default="&#x22;&#x22;">
    A note for the run, e.g. `baseline` or `prompt v2`. Shown by `verdict runs` and the web UI.
  </Field>

  <Field name="--split" type="dev | holdout">
    Only the holdout cases (listed in [`cases/holdout.json`](/docs/reference/holdout-json)), or only the others. Combines with `--cases`. Change your app while looking at dev; judge on holdout.
  </Field>

  <Field name="-c, --concurrency" type="int">
    Trials of one case that run at once. Default: `[run].concurrency`.
  </Field>

  <Field name="--app-root" type="path">
    Run the app code in this directory instead of `[app].root`: a candidate checkout. Cases, checks and the store still come from the config.
  </Field>

  <Field name="--candidate" type="name">
    Run a candidate made by [`verdict candidate new`](#verdict-candidate), as it is on disk (uncommitted edits included). The label defaults to `candidate NAME`.
  </Field>

  <Field name="--code" type="git ref">
    Run a branch or commit, checked out as a detached git worktree under `<store dir>/worktrees/<sha>`. The app runs from the same path relative to the repository root. Takes precedence over `--app-root`.
  </Field>

  <Field name="--skip-self-test" type="flag">
    Run even if a grader fails its self-test.
  </Field>

  <Field name="--budget" type="quick | standard | thorough">
    A preset in statistical terms. `quick` is k=1 on dev cases only (a smoke test), `standard` is the config's k on every case, and `thorough` is at least 5 per case, enough for per-case tests to resolve a 0/5 → 5/5 change. `-k` and `--split` override it. The choice is printed.
  </Field>

  <Field name="--max-cost" type="USD">
    Refuse to start (error VD200) if the estimate is over this. Default: `[run] budget_usd`, if set.
  </Field>

  <Field name="--yes" type="flag">
    Run even if the estimate is over the budget.
  </Field>

  <Field name="--estimate" type="flag">
    Print how many trials the run would take, and its cost and time, then stop. Cost is from the cases' trials in the five newest runs (the project's mean for cases that never ran); time is from recent runs' wall clock per trial. With nothing run yet, both are unknown (`null` in JSON).
  </Field>

  <Field name="--json" type="flag">
    Print progress to stderr and, at the end, one JSON object on stdout: `run_id`, `status`, `label`, `code`, `totals`, `stats`, and `cases` (each with `case`, `split`, `passed`, `trials`, `pass_k`, `failing_checks`).
  </Field>
</Fields>

Before running, `run` self-tests the selected cases. Output, per case:

```text
→ refund-within-window  (3 trials)
  ✓ ✓ ✓   3/3 passed

Run 20261008-191805-nogit-ba44 saved to /path/.verdict/verdict.db
Compare it: verdict compare <baseline-run> 20261008-191805-nogit-ba44
```

The run ends with a summary (trials passed, cases failing, cost, model calls, time) and, when judge checks called a model, a judge-usage line. It warns when no model calls were traced, or when the trace lost the spans some effects happened in (`trace_incomplete`). When a case failed a trial, the output ends with a `verdict trial RUN CASE` line for each. When the app reached a host that no boundary fakes, `! blocked HOST` follows that case's line.

**Exit codes:** 0 when the run finished, whatever the trials' results. 1 when a selected case's grader fails its self-test (without `--skip-self-test`), the config or a case is invalid, or the app fails to import (the run is saved as `failed`). A run stopped with Ctrl-C is saved with status `cancelled`; one that crashed, with status `failed` and the error.

## verdict compare [#verdict-compare]

Compare a baseline with a candidate, case by case.

```bash
verdict compare BASELINE CANDIDATE [--config PATH] [--cases IDS]
                [--max-cost-increase FRACTION] [--gate] [--json]
```

<Fields>
  <Field name="BASELINE" type="run ids">
    One or more run ids, comma-separated; `latest` and `latest~N` work. Trials of the same case are pooled.
  </Field>

  <Field name="CANDIDATE" type="run ids">
    One or more run ids, comma-separated.
  </Field>

  <Field name="--config" type="path" default="verdict.toml">
    The project config. Its store holds the runs; its cases folder holds `holdout.json`.
  </Field>

  <Field name="--cases" type="ids">
    Compare only these case ids.
  </Field>

  <Field name="--max-cost-increase" type="float">
    Fail the gate if mean cost per trial rises by more than this fraction, e.g. `0.2` for 20%.
  </Field>

  <Field name="--gate" type="flag">
    Exit 1 if the candidate fails the gate. For CI.
  </Field>

  <Field name="--json" type="flag">
    Print the full comparison as JSON instead of text: `schema`, `verdict`, `improved` and `regressed` (cases with a big swing, which with few trials may not be beyond noise yet; `fixed` and `broken` are the ones beyond noise), `gate_passed`, `reasons`, `same_code`, `fixed`, `broken`, `overall`, `dev`, `holdout`, `cost_per_trial`, `cost_change`, `seconds_per_trial`, `only_base`, `only_cand` and `rows` (one per case, with `flag`, `p`, `p_adj`, `significant`, `delta`, `split`, `base_rate`, `cand_rate`, `cost_delta`, `seconds_delta`).
  </Field>
</Fields>

Cases whose input, world, entry, clock or user changed between the two sides are left out and listed as not comparable (`case_changed`). The gate also fails when a case was graded with different checks on each side (`checks_changed`) until you run `verdict regrade`. Trials where the app raised are counted (`errors`), and a change in Python, lockfiles or model SDK versions between the sides is printed as a warning (`dependency_changes`).

The gate fails when a case regressed, a case is worse beyond noise, the overall interval is entirely below zero, cost rose past `--max-cost-increase`, or the sides share no cases. See [Compare and the gate](/docs/concepts/compare-and-the-gate).

**Exit codes:** 0 after comparing. 1 for an unknown run id, or with `--gate` when the gate fails. An A/A comparison (same code on both sides) never fails the gate.

## verdict self-test [#verdict-self-test]

Check every grader against its known-good and known-bad results.

```bash
verdict self-test [--config PATH] [--cases IDS] [--json]
```

<Fields>
  <Field name="--config" type="path" default="verdict.toml">
    The project config.
  </Field>

  <Field name="--cases" type="ids">
    Comma-separated case ids. Default: every case.
  </Field>
</Fields>

```text
—    large-order-escalates: no oracle.json or known_bad.json
ok   payments-down: known-good 1.0, known-bad 0.25
```

`FAIL` lines list each problem. `--json` prints `ok` and one row per case (`case`, `oracle`, `known_bad`, `problems`). **Exit codes:** 0 when no case has a problem; 1 when any does.

## verdict regrade [#verdict-regrade]

Re-score stored trials with the current checks. No model calls; your app doesn't run.

```bash
verdict regrade [--config PATH] [--runs IDS] [--json]
```

<Fields>
  <Field name="--config" type="path" default="verdict.toml">
    The project config.
  </Field>

  <Field name="--runs" type="run ids">
    Comma-separated run ids. Default: every finished, failed or cancelled run.
  </Field>
</Fields>

```text
Re-graded 2 run(s); 0 trial(s) changed result.
```

Updates trial grades, per-case summaries and run totals. **Exit code:** 0.

## verdict serve [#verdict-serve]

Serve the local web app for this project on 127.0.0.1.

```bash
verdict serve [--config PATH] [--port N] [--open]
verdict view  [--config PATH] [--port N] [--open]     # alias
```

<Fields>
  <Field name="--config" type="path" default="verdict.toml">
    The project config.
  </Field>

  <Field name="--port" type="int" default="8765">
    The port to listen on.
  </Field>

  <Field name="--open" type="flag">
    Open the app in your browser once it's up.
  </Field>
</Fields>

Needs the `web` extra. The released package includes the built app; from a source checkout, build it first (`cd web && npm install && npm run build`), or pages answer 503 with the instructions. **Exit codes:** 1 if the `web` extra is missing or the port is in use; otherwise it runs until stopped. See [Web UI](/docs/web-ui).

## verdict generate [#verdict-generate]

Write cases from the project's generator spec.

```bash
verdict generate [-n N] [--config PATH] [--spec PATH] [--estimate] [--dry-run]
                 [--budget USD] [--model ALIAS] [--review-model ALIAS] [--seed N] [--concurrency N] [--json]
```

<Fields>
  <Field name="-n" type="int" default="10">
    How many accepted cases to write. When some are rejected, more scenarios are sampled, for up to three rounds; it stops early if a round accepts none.
  </Field>

  <Field name="--config" type="path" default="verdict.toml">
    The project config. Cases are written to its `[run].cases` folder.
  </Field>

  <Field name="--spec" type="path">
    The generator spec. Default: `[generator].spec` in `verdict.toml`, else `generator/generator.toml` next to it.
  </Field>

  <Field name="--estimate" type="flag">
    Print the cost estimate only. No calls. Once this project has at least three generated cases, it's what they actually cost per case (rejections included); before that, it's from prompt sizes and runs low.
  </Field>

  <Field name="--dry-run" type="flag">
    Run the whole pipeline with the spec's free `fake()` writer instead of Claude. Needs a `fake` hook. The reviewer is skipped. The stand-in cases go to `<cases>/_dryrun/` (replaced each time), where runs never load them.
  </Field>

  <Field name="--keep" type="flag">
    With `--dry-run`, write the stand-ins into the cases folder like real ones (for demos).
  </Field>

  <Field name="--budget" type="float">
    Stop starting new cases once spend passes this many USD (API-equivalent). Default: `[budget].usd` in the spec.
  </Field>

  <Field name="--model" type="string">
    The writing model, a Claude Code alias such as `sonnet`. Default: `[model].write`.
  </Field>

  <Field name="--review-model" type="string">
    The reviewing model. Default: `[model].review`.
  </Field>

  <Field name="--seed" type="int" default="7">
    Sampling seed. The same seed draws the same scenarios.
  </Field>

  <Field name="--concurrency" type="int">
    Cases written at once. Default: `[budget].concurrency` in the spec.
  </Field>

  <Field name="--json" type="flag">
    Print the estimate (with `--estimate`) or the batch's summary and per-case results as JSON.
  </Field>
</Fields>

With `--pilot K`, each new case then runs K times on the current app. A case that fails every trial is more likely wrong than the app, so it moves to `cases/_review/<id>/` with an `issue.json` and stays out of runs until you move it back; one that passes some trials is kept and listed as flaky.

Real generation calls Claude Code headless (`claude -p`) and needs the `claude` CLI on your path.

```text
[1/3] case     gen-delivered-recently-ca132a  $0.000
[2/3] case     gen-delivered-long-ago-c4a8f0  $0.000
[3/3] case     gen-near-the-deadline-dbbf7b  $0.000

3 new case(s), 0 rejected, 0 skipped (budget); 0 needed a second attempt. $0.000 over 0 calls, 0.0 min.
```

**Exit codes:** 0 when the batch finished, including when some scenarios were rejected (see `cases/_rejected/`). 1 for a missing or invalid spec (scaffold one with `verdict init --generator`), a missing `fake` hook with `--dry-run`, or a missing `claude` CLI. See [Generating cases](/docs/concepts/generating-cases).

## verdict runs [#verdict-runs]

List the runs in the store, newest first.

```bash
verdict runs [--config PATH] [--limit N] [--json]
```

<Fields>
  <Field name="--config" type="path" default="verdict.toml">
    The project config.
  </Field>

  <Field name="--limit" type="int">
    Only the newest N runs.
  </Field>

  <Field name="--json" type="flag">
    One row per run: `id`, `status`, `label`, `created_at`, `code`, `fingerprint`, `cases`, `cases_passing`, `trials`, `passed`, `errors`, `mean_reward`, `delta` (against the run before), `cost_usd`, `error`.
  </Field>
</Fields>

```text
20261008-191805-nogit-8dd8           finished      12/15   code 4a5ffdf5d401  candidate
20261008-191805-nogit-ba44           finished      15/15   code 8cfd65598bb3  baseline
```

Columns: run id, status (`queued`, `running`, `finished`, `failed`, `cancelled`), trials passed / trials, code fingerprint, label. **Exit code:** 0.

## verdict show [#verdict-show]

One run, case by case, failing cases first.

```bash
verdict show [RUN] [--config PATH] [--json]
```

`RUN` defaults to `latest`. Each case shows its trial marks and the names of its failing checks; the last line is the `verdict trial` command for the first failing case. `--json` gives the run (status, label, totals, planned cases, git), `stats`, and one row per case with `pass_k`, `mean_reward`, `failing_checks`, `trials`, `split` and `delta` against the case's previous run.

```text
Run 20261009-032500-nogit-40ea  finished  cand
  code main   8/10 trials passed

  FAIL  refund-within-window                     ✕ ✕
          ✕ payments.refund(order_id=A100, amount=24.0) ×1
  pass  large-order-escalates                    ✓ ✓
```

## verdict cases [#verdict-cases]

Every case with its latest result: `failing`, `passing`, `never` run or `running`, its split, and its last pass count. `--json` adds tags, pass rate over recent runs and whether it has an oracle and a known-bad result.

```bash
verdict cases [--config PATH] [--json]
```

## verdict case [#verdict-case]

One case in full: title, description, folder, input, a summary of the world (tables, scripted services, clock), its checks, whether its grader passes its self-test, and its latest result.

```bash
verdict case CASE_ID [--config PATH] [--json]
```

## verdict trial [#verdict-trial]

Why a trial passed or failed, in one view.

```bash
verdict trial RUN CASE_ID [TRIAL] [--config PATH] [--json]
```

`TRIAL` defaults to the case's first failing trial in that run, else 1. Shows each failed check with what was seen and **the step behind it** (the tool call that made the offending effect, or the agent that produced the output), the passed checks, the output, every effect in call order, the state changes, and a summary of the trace.

```text
refund-within-window · trial 1 of run 20261009-032500-nogit-40ea (cand)
  FAILED   reward 0.25   trials ✕1 ✕2

Failed checks
  ✕ payments.refund(order_id=A100, amount=24.0) ×1
      saw: 0 matching (want 1); calls to it: no matching calls
      step: agent Refund agent — returns {}  (id 727d134d3b3c28a1)
```

In a multi-turn case, the conversation is shown first, with how it ended.

`--json` gives `score` (the fraction of checks passed), `feedback` (one string: each failed check, what was seen and the step behind it, the form reflective optimisers read), `run`, `case`, `trial`, `passed`, `reward`, `error`, `cost_usd`, `duration_s`, `siblings`, `input`, `failed` (each check with `name`, `detail`, `kind` and `step`), `passed_checks`, `output`, `output_text`, `effects` (`seq`, `service`, `action`, `args`, `ok`, `result`), `state_changes`, `usage`, `blocked_hosts`, `user`, `transcript`, `termination` and `trace` (counts, errors, warnings, total time, cost).

## verdict trace [#verdict-trace]

A trial's steps as a tree: agents, model calls (each labelled with its decision, such as `→ refund(order_id=A1)`), tool calls, events and errors, with what each step did to the outside world and which failed check is blamed on it.

```bash
verdict trace RUN CASE_ID [TRIAL] [--full] [--config PATH] [--json]
```

By default it shows agents, model calls, tools, warnings and errors; `--all` adds log lines and wrapper spans (JSON: `hidden_steps` says how many were left out). `--full` adds each model call's new input messages and its output, and each tool call's arguments and result. `--json` gives `summary` and `steps`, a flat list in tree order: `id`, `depth`, `type` (`agent`, `generation`, `tool`, `event`, `span`), `name`, `agent`, `level`, `label`, `effects`, `offset_ms`, `duration_ms`; model calls add `model`, `usage`, `cost`, `input`, `output` and `tools`; tool calls add `args` and `result`. Traces need the `otel` extra and an app instrumented with OpenTelemetry; see [Traces and observations](/docs/concepts/traces-and-observations).

## verdict state [#verdict-state]

The database rows a trial added, changed (field by field) and removed, against the case's world. Needs a `[state]` adapter.

```bash
verdict state RUN CASE_ID [TRIAL] [--config PATH] [--json]
```

## verdict brief [#verdict-brief]

What a coding agent needs to fix a run's failures, in one document that fits in its context.

```bash
verdict brief [RUN] [--cases IDS] [--max-chars N] [--config PATH] [--json]
```

`RUN` defaults to `latest`. The brief has:

* how to work (fix the app, never the cases; one candidate per fix; how to check and rank);
* the failure patterns: blamed steps and failed check kinds across cases;
* the product's rules (`generator/rules.md`, if there is one) and each agent's system prompt;
* per failing case: what it tests, its scenario, the input, and each failed check with what was seen and the step behind it. For a model call, that's the last messages it was sent and what it decided; for a tool call, its arguments and result.
* per case, also the error and where it was raised, the effects, the data changes, the conversation, and the commands for more detail.

It's Markdown by default and capped at `--max-chars` (default 32000, about 8k tokens): cases that don't fit are named at the end. `--json` gives the same, structured.

## verdict candidate [#verdict-candidate]

Attempted fixes, each in its own git worktree on its own branch, so the main checkout stays the baseline.

```bash
verdict candidate new NAME [--from REF]   # branch verdict/NAME in .verdict/worktrees/NAME; prints where to edit
verdict candidate list                    # each candidate's changes against the app and its runs
verdict candidate diff NAME               # the change, as a diff
verdict candidate remove NAME             # the worktree (the branch and runs are kept)
```

Run one with `verdict run --candidate NAME`: it runs as it is on disk, uncommitted edits included, and the run records its diff. Glue next to `verdict.toml` is used even if the candidate's branch doesn't have it.

## verdict rank [#verdict-rank]

Rewards for a set of candidates against one baseline.

```bash
verdict rank BASELINE CANDIDATE [CANDIDATE …] [--max-cost-increase F] [--config PATH] [--json]
```

Each argument is a run id (or `latest~N`), or several comma-separated, pooled. Each candidate is compared with the baseline as [`compare`](#verdict-compare) does, on the cases both ran. The **reward** is the mean per-case change in pass rate. Candidates are ranked by whether they pass the gate, then by reward, then by fewest regressions. Each row also lists the cases that improved and regressed (big swings) and that were fixed or broken (beyond noise). `best` is the top candidate if it passes the gate with a positive reward. The web app's Rank page (from Code: "Rank candidates") shows the same.

## verdict lint [#verdict-lint]

Check `verdict.toml` and every case without importing the app: fast enough to run after every file a person or an agent writes.

```bash
verdict lint [--config PATH] [--json]
```

Each problem has a code, where it is, what's wrong and the fix. Errors: `VD100` invalid config, `VD101` a target that isn't `package.module:name` (or an undefined entry), `VD102` a target's module not found in the app, next to `verdict.toml` or installed, `VD110` a case that doesn't load, `VD111` a check missing a field its kind needs, `VD121` world tables that aren't lists of rows, `VD130` unknown ids in `holdout.json`, `VD140` a `[user]` case on a non-turns entry. Warnings: `VD112` an effect check on an undeclared service, `VD120` a world script for an undeclared service, `VD150` a case with no `oracle.json` or `known_bad.json`. **Exit codes:** 1 if there's any error.

## verdict schema [#verdict-schema]

```bash
verdict schema config|case|world
```

Prints JSON Schema for `verdict.toml`, `case.toml` or `world.json`.

## verdict agent [#verdict-agent]

Install what lets a coding agent set up and run Verdict in this repo.

```bash
verdict agent setup [DIR] [--agents claude,codex] [--dry-run] [--json]
verdict agent update [DIR]                      # same, after upgrading Verdict
```

Writes:

* three skills, `verdict-setup` (maps the app, interviews you about what must work and what must never happen, writes the config, fakes and first cases, validates until green), `verdict-cases` and `verdict-check`, to `.claude/skills/` (Claude Code) and `.agents/skills/` (Codex, Cursor, Copilot, Gemini);
* two Claude Code subagents, `verdict-mapper` and `verdict-case-writer`, to `.claude/agents/`;
* a short block in `AGENTS.md` (between `verdict:start`/`verdict:end` markers) saying when to use each skill and the rules;
* `@AGENTS.md` in an existing `CLAUDE.md`, and `Bash(verdict *)` in `.claude/settings.json`.

Running it again changes only what's out of date. In Claude Code, start with `/verdict-setup`.

## verdict worktrees [#verdict-worktrees]

The checkouts `verdict run --code REF` made under `.verdict/worktrees/`, with each one's size, the runs that used it and its commit subject.

```bash
verdict worktrees [--prune] [--config PATH] [--json]
```

`--prune` removes every checkout not used by a run in progress (`git worktree remove`). The runs stay in the store with their code fingerprints and diffs; `--code REF` checks the commit out again when needed.

## verdict guide [#verdict-guide]

How a coding agent should use Verdict: the change loop, how to find out why a case failed, and the rules (fix the app, never the checks).

```bash
verdict guide                          # print it
verdict guide --install-skill [DIR]    # write DIR/.claude/skills/verdict/SKILL.md (default DIR: .)
```


---

# verdict.toml

> Every key in the project config, with its type and default.

Source: https://mitej23.github.io/docs/reference/verdict-toml



`verdict.toml` configures one project: where the app's code is, how to call it, what to isolate, and where cases and results live. Paths are relative to the file.

```toml title="examples/refunds/verdict.toml"
schema = 1

[app]
root = "."
paths = ["refunds_app"]

[entry.default]
target = "refunds_app.agent:handle"

[state]
adapter = "sqlalchemy"
metadata = "refunds_app.db:Base.metadata"
session = "refunds_app.db:SessionLocal"

[[boundary]]
service = "payments"
patch = "refunds_app.services:PaymentsClient"

[[boundary]]
service = "email"
patch = "refunds_app.services:send_email"

[network]
allow = []

[run]
cases = "cases"
store = ".verdict/verdict.db"
trials = 3
```

**Targets** are written `package.module:attribute`, and the attribute may be dotted (`Base.metadata`). They are imported with `[app].root` first on `sys.path`.

## Top level [#top-level]

<Fields>
  <Field name="schema" type="int" default="1">
    The format version. Any value other than `1` is rejected.
  </Field>
</Fields>

## \[app] [#app]

<Fields>
  <Field name="root" type="path" default="&#x22;.&#x22;">
    The directory holding the code under test. It goes first on `sys.path`. `verdict run --app-root` and `--code` replace it for one run.
  </Field>

  <Field name="paths" type="list of paths" default="[]">
    What the [code fingerprint](/docs/concepts/code-versions) covers, relative to `root`: every `.py` file under each path. Empty means the whole root. Also used to scope the uncommitted diff a run records.
  </Field>
</Fields>

## \[entry.\<name>] [#entryname]

At least one entry is required. Cases pick one by name with `entry`; the default name is `default`.

<Fields>
  <Field name="target" type="target">
    A callable `fn(input: dict) -> Any`, sync or async. Its return value is the trial's `output`; an exception becomes the trial's `error`.
  </Field>

  <Field name="mode" type="&#x22;function&#x22; | &#x22;turns&#x22;" default="&#x22;function&#x22;">
    `turns` makes it a chat entry, `fn(session: dict, message: str) -> Any`, called once per user turn of a [multi-turn case](/docs/guides/multi-turn-cases) within one trial. `session` holds `id`, `case`, `trial` and the case's `input`, and is the same dict every turn. The last reply is the trial's `output`.
  </Field>
</Fields>

```toml
[entry.default]
target = "myapp.agent:handle"

[entry.webhook]
target = "myapp.webhooks:on_message"
```

## \[users] [#users]

Optional. `simulator = "module:fn"` is the default simulator for simulated users, called as `fn(user: dict, transcript: list) -> str`; without it, simulated users are played by Claude Code (`claude -p`). A case's own `[user] simulator` wins.

## \[state] [#state]

Optional. Without it, trials have no database isolation and no `state` in their results.

<Fields>
  <Field name="adapter" type="string">
    Must be `"sqlalchemy"`. Other adapters are on the roadmap. Needs the `sqlalchemy` extra.
  </Field>

  <Field name="metadata" type="target">
    The app's SQLAlchemy `MetaData`, e.g. `"myapp.db:Base.metadata"`.
  </Field>

  <Field name="session" type="target">
    The session factory the app calls, e.g. `"myapp.db:SessionLocal"`. Swapped everywhere it's imported by name.
  </Field>
</Fields>

See [State adapters](/docs/concepts/state-adapters).

## \[\[boundary]] [#boundary]

Optional, repeatable: one table per external service.

<Fields>
  <Field name="service" type="string">
    The service name effects are recorded under and `world.json` scripts by, e.g. `"payments"`.
  </Field>

  <Field name="patch" type="target">
    A class (every public method is faked) or a function (swapped in every module under the app root that references it).
  </Field>

  <Field name="fake" type="target">
    A custom fake class whose methods take the same arguments as the real ones. Default: record each call and answer from `world.services`.
  </Field>
</Fields>

See [Boundary fakes](/docs/concepts/boundary-fakes).

## \[network] [#network]

<Fields>
  <Field name="allow" type="list of hosts" default="[]">
    Hosts the app may resolve, usually model APIs. A host matches an entry exactly or as a subdomain. Everything else raises `NetworkBlocked`: inside a trial, and for the whole run outside one too (the app importing, threads it starts), where only loopback also passes. `run` and `verdict trial` name every host that was refused.
  </Field>
</Fields>

## \[env] [#env]

Optional. What the app sees of the environment during a run, so production credentials never reach code under test. By default, variables matching tokens, secrets, passwords, DSNs and database URLs (`*TOKEN*`, `*SECRET*`, `*PASSWORD*`, `*_DSN`, `DATABASE_URL` …) are hidden; `*_API_KEY` passes, because the app calls its model. If `python-dotenv` is installed, the app's `load_dotenv()` and `dotenv_values()` get the same filter. `verdict doctor` lists what's hidden, and the environment is restored when the run ends.

<Fields>
  <Field name="set" type="table">
    Values to set, e.g. tripwires: `{ POSTGRES_HOST = "db.verdict.invalid" }`.
  </Field>

  <Field name="unset" type="list of names">
    More variables to hide.
  </Field>

  <Field name="forbid" type="list of patterns">
    Glob patterns added to the defaults, e.g. `["*_WEBHOOK*"]`.
  </Field>

  <Field name="allow" type="list of patterns">
    Let these through even if they match a forbidden pattern.
  </Field>
</Fields>

Verdict's own `env:` settings, such as `[trace.export]` headers, still read the real value.

## \[run] [#run]

<Fields>
  <Field name="cases" type="path" default="&#x22;cases&#x22;">
    The folder holding case folders.
  </Field>

  <Field name="store" type="path" default="&#x22;.verdict/verdict.db&#x22;">
    The SQLite results store. Created on first use. `--code` worktrees go in a `worktrees/` folder next to it.
  </Field>

  <Field name="trials" type="int" default="2">
    Trials per case, unless the case sets `trials` or the run passes `-k`.
  </Field>

  <Field name="concurrency" type="int" default="4">
    How many trials of one case run at once. Cases run one after another.
  </Field>

  <Field name="budget_usd" type="USD">
    Refuse to start a run whose estimate (`verdict run --estimate`) is over this, until `--yes`. Keeps a person or a coding agent from spending more than agreed.
  </Field>

  <Field name="background_timeout" type="seconds" default="30">
    After the entry returns, how long to wait for asyncio tasks it started and left running (fire-and-forget work such as an analysis after the reply). Their effects and state count; a task that raises becomes the trial's error, and one still running at the timeout is cancelled and reported.
  </Field>
</Fields>

## \[pricing] [#pricing]

Optional.

<Fields>
  <Field name="path" type="path">
    A project `pricing.toml`, merged model by model over the bundled prices. See [Keep costs down](/docs/guides/keep-costs-down#pricing).
  </Field>
</Fields>

## \[trace] [#trace]

Optional.

<Fields>
  <Field name="agent_names" type="table" default="{}">
    Display names for agents in the trace view, keyed by `gen_ai.agent.name` or span name: `{ refund_agent = "Refund agent" }`.
  </Field>

  <Field name="instrument" type="&#x22;auto&#x22; | list" default="&#x22;auto&#x22;">
    Model-call instrumentation Verdict turns on for a run. `"auto"` is `["pydantic_ai"]` when pydantic-ai is installed. Add `"openai"` for apps that call the OpenAI SDK directly (a span per `chat.completions.create` / `responses.create` with the model, tokens, messages and reply), unless the app already instruments OpenAI. Verdict collects spans from whatever tracer provider the app sets up, Logfire included. A run that traced no model calls while model hosts are allowed prints a warning.
  </Field>
</Fields>

## \[judge] [#judge]

Optional. Who answers [`judge` checks](/docs/reference/checks#judge).

<Fields>
  <Field name="model" type="string" default="&#x22;haiku&#x22;">
    The Claude Code model alias that answers, through `claude -p`.
  </Field>

  <Field name="grader" type="target">
    Your own: `fn(question: str, text: str) -> (answer: bool, reason: str)`. Use it to call another model, or in tests.
  </Field>
</Fields>

Answers are cached by (model or grader, question, text) in `judge-cache.json` next to the store.

### \[trace.export] [#traceexport]

Optional. Also send every trial's spans to an OTLP/HTTP endpoint (Logfire, Langfuse, Honeycomb, an OpenTelemetry collector), tagged `deployment.environment = "eval"` so they never mix with production traces. Spans are always kept locally too. Needs the `otel-export` extra. The endpoint's host is allowed through the network guard automatically.

```toml title="verdict.toml"
[trace.export]
endpoint = "https://logfire-api.pydantic.dev/v1/traces"
service = "refunds-evals"
headers = { Authorization = "env:LOGFIRE_EVALS_TOKEN" }
```

<Fields>
  <Field name="endpoint" type="url">
    The OTLP/HTTP traces endpoint, usually ending in `/v1/traces`.
  </Field>

  <Field name="headers" type="table" default="{}">
    Request headers. A value `"env:NAME"` is read from the environment variable `NAME`, so tokens stay out of the file; a missing variable is an error.
  </Field>

  <Field name="service" type="string">
    `service.name` on the exported spans.
  </Field>
</Fields>

## \[ui] [#ui]

Optional.

<Fields>
  <Field name="output_field" type="string">
    Which field of a dict output is the human-readable text the web UI shows and diffs. Default: the first string field among `reply`, `message`, `text`, `answer`, `response` and `output`, else the whole output as JSON.
  </Field>

  <Field name="tables" type="table" default="{}">
    Per state table, how **Data changes** matches and names rows: `key`, the field or fields that identify a row (default: the first of `id`, `uuid`, `key`, `pk`, else its position), and `label`, the table's heading. `{ orders = { key = ["id"], label = "Orders" }, events = { key = ["comm_id", "type"] } }`.
  </Field>
</Fields>

## \[generator] [#generator]

Optional.

<Fields>
  <Field name="spec" type="path" default="&#x22;generator/generator.toml&#x22;">
    The case generator's spec. `verdict generate --spec` overrides it. See [generator.toml](/docs/reference/generator-toml).
  </Field>
</Fields>

## Errors [#errors]

The config is validated on load, with messages that say how to fix it:

```text
No config at /path/verdict.toml. Create one with `verdict init`.
/path/verdict.toml: needs at least one [entry.<name>] with a target, e.g. [entry.default] target = "app.agent:handle"
/path/verdict.toml: every [[boundary]] needs `service` and `patch`
/path/verdict.toml: [state] adapter must be "sqlalchemy" (more adapters are on the roadmap)
```


---

# case.toml

> Every key in a case's case.toml, the file that defines one case's input and checks.

Source: https://mitej23.github.io/docs/reference/case-toml



`case.toml` defines one case: its identity, the input passed to the entry point, and the checks. It lives in `cases/<id>/`.

```toml title="cases/large-order-escalates/case.toml"
schema = 1
id = "large-order-escalates"
title = "A $450 order goes to a person, not an automatic refund"
tags = ["refund", "escalation"]
clock = "2026-10-08T10:00:00+00:00"

[input]
order_id = "A103"
message = "The sofa is damaged, refund please."

[[check]]
kind = "no_effect"
service = "payments"
action = "refund"

[[check]]
kind = "effect"
service = "email"
args = { to = "support@shop.example" }
count = 1

[[check]]
kind = "state"
table = "orders"
where = { id = "A103" }
equals = { status = "escalated" }
```

## Keys [#keys]

<Fields>
  <Field name="schema" type="int" default="1">
    The format version. Any other value is rejected.
  </Field>

  <Field name="id" type="string" default="the folder name">
    Unique and stable. Must match `^[a-z0-9][a-z0-9._-]*$`. Two cases with the same id are an error.
  </Field>

  <Field name="title" type="string" default="the id">
    A plain-English name for the situation.
  </Field>

  <Field name="description" type="string">
    What the case tests and its correct outcome, in a few sentences. Shown on the case's Overview tab. Generated cases get one from the writer.
  </Field>

  <Field name="tags" type="list of strings" default="[]">
    Free-form labels. Generated cases get `generated`.
  </Field>

  <Field name="entry" type="string" default="&#x22;default&#x22;">
    Which `[entry.<name>]` in `verdict.toml` to call. An undefined entry, or a case whose `[user]` table doesn't match the entry's `mode`, stops the run before it starts.
  </Field>

  <Field name="trials" type="int">
    Trials for this case, overriding `[run].trials`. `verdict run -k` overrides both.
  </Field>

  <Field name="clock" type="ISO 8601 datetime">
    Start time at this instant for each of the case's trials, e.g. `"2026-10-08T10:00:00+00:00"`; it then moves on normally, so traces keep real durations. Needs the `clock` extra.
  </Field>

  <Field name="[input]" type="table" default="{}">
    Passed to the entry point as a dict. A fresh deep copy per trial.
  </Field>

  <Field name="[user]" type="table">
    Makes this a multi-turn case: a scripted or simulated user talks to an entry with `mode = "turns"`. See [Multi-turn cases](/docs/guides/multi-turn-cases).

    * Scripted: `turns`, a list of messages; a turn written `{ when = "regex", say = "..." }` is sent only if the app's last reply matches.
    * Simulated: `reason_for_call`, `known_info`, `unknown_info`, `instructions`, `persona`, optional `model` (default `haiku`) and `simulator = "module:fn"`.
    * Either: `max_turns` (default 10).
  </Field>

  <Field name="[[check]]" type="array of tables">
    At least one. Each has a `kind` and that kind's fields. See [Check kinds](/docs/reference/checks).
  </Field>

  <Field name="[generator]" type="table">
    Written by `verdict generate`: the scenario's dimension values and `_dimensions` (every value of every dimension). Used for coverage in the web UI.
  </Field>
</Fields>

Other keys are kept and ignored at run time.

## Folder contents [#folder-contents]

| File                            |                                                                                      |
| ------------------------------- | ------------------------------------------------------------------------------------ |
| `case.toml`                     | required                                                                             |
| `world.json`                    | optional. See [world.json](/docs/reference/world-json)                               |
| `oracle.json`, `known_bad.json` | optional. See [oracle.json and known_bad.json](/docs/reference/oracle-and-known-bad) |
| `checks.py`                     | optional. Python checks referenced as `target = "checks:fn"`                         |
| `generation.json`               | written by `verdict generate`; not read at run time                                  |

## Loading rules [#loading-rules]

* Case folders are the direct subfolders of `[run].cases` that contain a `case.toml`.
* Folders whose names start with `_` or `.` are skipped.
* A check with an unknown `kind` is an error that lists the valid kinds.
* Errors name the file: `cases/x/case.toml: a case needs at least one [[check]]`.


---

# world.json

> The starting database rows and scripted service responses a case's trials begin from.

Source: https://mitej23.github.io/docs/reference/world-json



`world.json` is the state a case starts from: rows for each table and scripted responses for each faked service. Each trial gets its own deep copy. The file is optional; without it, the world is empty.

```json title="cases/payments-down/world.json"
{
  "schema": 1,
  "tables": {
    "orders": [
      {"id": "A104", "customer_email": "sam@example.com", "status": "delivered",
       "total": 80.0, "delivered_at": "2026-10-05T10:00:00"}
    ]
  },
  "services": {
    "payments": {"refund": {"raises": "503 provider unavailable"}}
  }
}
```

## Keys [#keys]

<Fields>
  <Field name="schema" type="int" default="1">
    The format version. Any other value is rejected.
  </Field>

  <Field name="tables" type="object" default="{}">
    `{table name: [row, ...]}`. Inserted into each trial's fresh database by the [state adapter](/docs/concepts/state-adapters). Every table must exist in the app's metadata. ISO strings are converted for date, datetime and time columns. Tables not listed start empty.
  </Field>

  <Field name="services" type="object" default="{}">
    `{service: {action: script}}`. The service name matches `[[boundary]].service`; the action is the method or function name.
  </Field>
</Fields>

## Scripts [#scripts]

Each action has one script:

<Fields>
  <Field name="returns" type="any">
    Returned on every call.
  </Field>

  <Field name="sequence" type="array">
    One value per call, in order. Once the list runs out, the last value repeats. An empty list returns `null`.
  </Field>

  <Field name="raises" type="string">
    The call raises `BoundaryError` with this message, and the effect is recorded with `ok: false` and the message as its `result`.
  </Field>
</Fields>

```json
{"services": {
  "payments": {
    "refund": {"returns": {"refund_id": "R1", "status": "ok"}},
    "lookup": {"sequence": [{"found": true}, {"found": false}]},
    "charge": {"raises": "card declined"}
  }
}}
```

An action with no script returns `None`. Returned values are deep copies, so the app can't change the world by mutating them.

Custom fakes read scripts with `scripted(service, action, default)`. See the [Python API](/docs/reference/python-api#verdictboundary).


---

# oracle.json and known_bad.json

> Known-good and known-bad results that prove a case's checks are right.

Source: https://mitej23.github.io/docs/reference/oracle-and-known-bad



`oracle.json` and `known_bad.json` are optional trial results stored with a case. The grader self-test requires the case's checks to pass every check on the first and fail at least one on the second.

```json title="cases/refund-within-window/oracle.json"
{"output": {"decision": "refunded", "reply": "Done: your refund of $24.00 is on its way."},
 "effects": [
   {"seq": 1, "service": "payments", "action": "refund", "args": {"order_id": "A100", "amount": 24.0}, "ok": true, "result": {"refund_id": "R-1"}},
   {"seq": 2, "service": "email", "action": "send_email", "args": {"to": "maya@example.com", "subject": "Your refund", "body": "..."}, "ok": true, "result": null}],
 "state": {"tables": {"orders": [{"id": "A100", "status": "refunded"}]}}}
```

```json title="cases/refund-within-window/known_bad.json"
{"output": {"decision": "outside_window", "reply": "Refunds are available within 7 days of delivery."},
 "effects": [],
 "state": {"tables": {"orders": [{"id": "A100", "status": "delivered"}]}}}
```

## Keys [#keys]

Both files have the shape of a [trial result](/docs/reference/store-schema#trial-result). Only these keys are read; missing ones take the defaults.

<Fields>
  <Field name="output" type="any" default="null">
    What the entry point returned.
  </Field>

  <Field name="effects" type="array" default="[]">
    Effects, each `{"seq", "service", "action", "args", "ok", "result"}`.
  </Field>

  <Field name="state" type="object" default="null">
    The final state, `{"tables": {name: [rows]}}`. Rows only need the columns your checks read.
  </Field>

  <Field name="error" type="string" default="null">
    An error message, for cases about failures.
  </Field>
</Fields>

## Rules [#rules]

| File             | Must                        |
| ---------------- | --------------------------- |
| `oracle.json`    | pass **every** check        |
| `known_bad.json` | fail **at least one** check |

`verdict self-test` reports each case; `verdict run` refuses to start when a selected case fails, unless `--skip-self-test`. Invalid JSON in either file is a load error.

## Also used by [#also-used-by]

* **The web UI** compares each trial's output and effects with `oracle.json`.
* **`verdict generate`** writes both files for every generated case and requires them to pass the self-test.

See [Grader self-test and regrade](/docs/concepts/self-test-and-regrade).


---

# holdout.json

> The list of cases compare reports separately as the holdout split.

Source: https://mitej23.github.io/docs/reference/holdout-json



`holdout.json` lists the cases that `verdict compare` reports as the holdout split. It sits in the cases folder, next to the case folders.

```json title="cases/holdout.json"
{"holdout": ["payments-down"]}
```

<Fields>
  <Field name="holdout" type="list of case ids" default="[]">
    Case ids in the holdout split. Every other case is in the dev split.
  </Field>
</Fields>

* The file is optional. Without it, every case is dev.
* Invalid JSON is an error.
* Holdout cases still count in the overall result and the gate. They are summarised on their own line, and each case row says its split.

```text
  all       5 cases   pass 100% → 80%   change -20% (95% CI -50% to +14%)   all trials pass 5 → 4
  dev       4 cases   pass 100% → 75%   change -25% (95% CI -60% to +15%)   all trials pass 4 → 3
  holdout   1 cases   pass 100% → 100%   change +0% (95% CI -46% to +47%)   all trials pass 1 → 1
```

Iterate against dev cases. Judge a change on holdout cases you haven't tuned against.

## Run one split [#run-one-split]

`verdict run --split` runs only one side, so you can iterate on dev cheaply and spend trials on holdout when you judge:

```bash
verdict run --split dev -k 2 --label "prompt v3"     # every case not in holdout.json
verdict run --split holdout -k 5 --label "prompt v3"   # only the holdout cases
```

`--split` combines with `--cases`. `--split holdout` without a `holdout.json`, or a split with no cases, is an error.


---

# generator.toml

> Every key in a case generator spec, plus the files verdict generate writes.

Source: https://mitej23.github.io/docs/reference/generator-toml



`generator.toml` is the case generator's spec: the scenario dimensions with their weights, the models, and the budget. By default it's `generator/generator.toml` next to `verdict.toml`; `[generator].spec` or `verdict generate --spec` point elsewhere. Paths in it are relative to it.

```toml title="examples/refunds/generator/generator.toml"
schema = 1
module = "scenarios.py"            # plan() decides each scenario's facts and checks
rules = "rules.md"                 # the policy in words, for the writer and the reviewer
example = "../cases/refund-within-window"
prefix = "gen"
id_from = "situation"              # gen-<situation>-<hash>

[model]
write = "sonnet"
review = "sonnet"

[budget]
usd = 2.0                          # stop starting new cases past this (API-equivalent USD)
per_call = 0.40
concurrency = 4

[dimensions.situation]
delivered_recently = 4             # within the 30-day window
delivered_long_ago = 2             # outside it
near_the_deadline = 1              # 29 days: still inside
still_in_transit = 2

[dimensions.order_size]
small = 4
large = 1

[dimensions.payments]
ok = 4
down = 1

[dimensions.tone]
polite = 3
angry = 2
terse = 2
```

## Top level [#top-level]

<Fields>
  <Field name="schema" type="int" default="1">
    The format version. Any other value is rejected.
  </Field>

  <Field name="module" type="path" default="&#x22;scenarios.py&#x22;">
    The hooks module. It must define `plan(values, rng)`. See [Generator hooks](/docs/reference/generator-hooks).
  </Field>

  <Field name="rules" type="path" default="&#x22;rules.md&#x22;">
    The product's rules in plain words, put in both the writer's and the reviewer's system prompt. Optional; a missing file means no rules.
  </Field>

  <Field name="example" type="path">
    A case folder shown to the writer for the format (its `case.toml`, `world.json`, `oracle.json` and `known_bad.json`), with an instruction not to copy its content.
  </Field>

  <Field name="prefix" type="string" default="&#x22;gen&#x22;">
    The first part of generated case ids.
  </Field>

  <Field name="id_from" type="string" default="the first dimension">
    The dimension whose value goes in the id: `<prefix>-<value>-<6-char hash>`, e.g. `gen-delivered-recently-ca132a`.
  </Field>

  <Field name="duplicate_threshold" type="float" default="0.9">
    A generated input this similar (0 to 1) to an existing case's input is rejected as a near-duplicate.
  </Field>
</Fields>

## \[model] [#model]

Models are Claude Code aliases or model ids, called through `claude -p`.

<Fields>
  <Field name="write" type="string" default="&#x22;sonnet&#x22;">
    The model that writes content. `--model` overrides it.
  </Field>

  <Field name="review" type="string" default="the write model">
    The independent reviewer. `--review-model` overrides it.
  </Field>
</Fields>

## \[budget] [#budget]

Costs are API-equivalent USD, as Claude Code reports them. On a Claude Code subscription the calls are not billed per token.

<Fields>
  <Field name="usd" type="float" default="5.0">
    Stop starting new cases once total spend passes this. Scenarios not started are reported as skipped. `--budget` overrides it.
  </Field>

  <Field name="per_call" type="float" default="0.50">
    The most one call may spend, passed to Claude Code as `--max-budget-usd`.
  </Field>

  <Field name="concurrency" type="int" default="4">
    Scenarios generated at once. `--dry-run` uses one.
  </Field>
</Fields>

## \[dimensions.\<name>] [#dimensionsname]

At least one dimension is required. Each maps a value to a positive relative weight.

```toml
[dimensions.order_size]
small = 4      # four times as likely as large
large = 1
```

A scenario takes one value per dimension. Sampling follows the weights, boosts values that have fewer generated cases than their share, and never draws the same combination twice. The `normalise()` hook can fold combinations that can't matter.

## \[\[require]] [#require]

Combinations that every batch draws first, until a case covers each, so the riskiest scenarios aren't left to chance:

```toml
[[require]]
slot = "same_service_clash"
clock = "evening_request"
```

Each table names one or more dimensions and a value of each; the rest are sampled as usual. A value that isn't in its dimension is an error, and so is one that `normalise()` would change.

## What generate writes [#what-generate-writes]

| Path                                                                  | When                                       |
| --------------------------------------------------------------------- | ------------------------------------------ |
| `cases/<id>/case.toml`, `world.json`, `oracle.json`, `known_bad.json` | a scenario passed every gate               |
| `cases/<id>/generation.json`                                          | with each generated case                   |
| `cases/_rejected/<id>.json`                                           | a scenario failed its gates twice          |
| `cases/_staging/`                                                     | temporary, removed at the end of the batch |

### generation.json [#generationjson]

An audit record; nothing reads it at run time.

<Fields>
  <Field name="scenario" type="object">
    `id`, `values`, `plan` and `seed`.
  </Field>

  <Field name="content" type="object">
    Claude's raw content, before `build()`.
  </Field>

  <Field name="attempts" type="array">
    Each attempt: `attempt`, `issues`, `cost_usd`, `model`, `tokens`, `self_test` (the oracle and known-bad rewards), `review` (the reviewer's answer) and `review_cost_usd`.
  </Field>

  <Field name="cost_usd" type="float">
    All attempts' write and review cost.
  </Field>

  <Field name="generated_at" type="ISO datetime">
    When the case was written.
  </Field>

  <Field name="settings" type="object">
    The models, budget, concurrency, dry-run flag and seed used.
  </Field>
</Fields>

### \_rejected/\<id>.json [#_rejectedidjson]

`{"scenario": ..., "attempts": [...]}`, with each attempt's issues. Files and folders starting with `_` are never loaded as cases.

### The \[generator] table in case.toml [#the-generator-table-in-casetoml]

Generated cases carry their scenario values and every dimension's values, for coverage:

```toml
[generator]
situation = "delivered_long_ago"
order_size = "n/a"
payments = "n/a"
tone = "polite"
_dimensions = { situation = ["delivered_recently", "delivered_long_ago", "near_the_deadline", "still_in_transit"], order_size = ["small", "large"], payments = ["ok", "down"], tone = ["polite", "angry", "terse"] }
```


---

# Generator hooks

> The functions and schema a generator spec's scenarios.py defines, with their arguments and return values.

Source: https://mitej23.github.io/docs/reference/generator-hooks



A generator spec's hooks module (`scenarios.py` by default) holds your product rules as code. `plan()` is required; the rest are optional.

| Hook                           | Required | Called                             | Returns                                          |
| ------------------------------ | -------- | ---------------------------------- | ------------------------------------------------ |
| `plan(values, rng)`            | yes      | once per sampled scenario          | facts, checks and the outcome                    |
| `normalise(values)`            | no       | on every sampled combination       | the values with irrelevant dimensions blanked    |
| `CONTENT_SCHEMA`               | no       | (a constant)                       | the JSON Schema Claude writes to                 |
| `build(values, plan, content)` | no       | after each write                   | the case: title, input, world, oracle, known_bad |
| `validate(values, plan, case)` | no       | after the built-in gates           | a list of problems                               |
| `fake(values, plan, rng)`      | no       | instead of Claude with `--dry-run` | content in the `CONTENT_SCHEMA` shape            |

In every hook, `values` is the scenario: `{dimension: value}`, after `normalise`.

## plan [#plan]

```python
def plan(values: dict, rng: random.Random) -> dict: ...
```

Decides what is correct. `rng` is seeded per scenario, so the same seed gives the same facts.

<Fields>
  <Field name="checks" type="list of dicts">
    The case's `[[check]]` tables, written into `case.toml` as is. See [Check kinds](/docs/reference/checks).
  </Field>

  <Field name="outcome" type="string">
    The correct outcome in words, for the writer and the reviewer.
  </Field>

  <Field name="facts" type="dict">
    Concrete values the case must use: ids, amounts, dates, names. The facts-used gate rejects a case whose input and world don't contain every string or number fact.
  </Field>

  <Field name="known_bad" type="string">
    The failure the known-bad result must show. Default: "a plausible failure that breaks at least one check".
  </Field>

  <Field name="clock" type="ISO datetime">
    The case's frozen clock. The reviewer judges dates against it.
  </Field>

  <Field name="tags" type="list of strings">
    Tags added after `generated`.
  </Field>
</Fields>

Any other keys are kept in the plan and passed to `build`, `validate` and `fake`.

```python title="generator/scenarios.py"
def plan(v, rng):
    outcome = _outcome(v)
    ...
    return {"facts": facts, "checks": checks, "outcome": words, "known_bad": bad,
            "clock": "2026-10-08T10:00:00+00:00", "tags": [outcome.replace("_", "-"), v["tone"]],
            "decision": outcome}
```

A plan without `checks` or `outcome` stops the batch with an error.

## normalise [#normalise]

```python
def normalise(values: dict) -> dict: ...
```

Set dimensions that can't change the outcome to a constant, such as `"n/a"`, so equivalent scenarios collapse into one and duplicates aren't drawn.

## CONTENT_SCHEMA [#content_schema]

A JSON Schema dict for what Claude writes. Default: a whole case:

```json
{"type": "object", "additionalProperties": false,
 "required": ["title", "input", "world", "oracle", "known_bad", "known_bad_failure"],
 "properties": {"title": {"type": "string", "maxLength": 90}, "input": {"type": "object"},
                "world": {"type": "object"}, "oracle": {...}, "known_bad": {...},
                "known_bad_failure": {"type": "string", "maxLength": 160}}}
```

Ask for less (only the text that needs language) and assemble the rest in `build`. The refunds spec asks for `title`, `message`, `good_reply`, `bad_reply` and `known_bad_failure`.

## build [#build]

```python
def build(values: dict, plan: dict, content: dict) -> dict: ...
```

Turns Claude's content into a case. Must return `title`, `input`, `world`, `oracle` and `known_bad`; it may return `clock`. Default: the content as is. An exception or a missing key fails the attempt with a message fed back to the writer.

## validate [#validate]

```python
def validate(values: dict, plan: dict, case: dict) -> list[str]: ...
```

Your own gate, run after the placeholder, facts-used and duplicate gates. Return problems as sentences; an empty list passes. Problems are fed back to the writer on the retry.

```python title="generator/scenarios.py"
def validate(v, plan, case):
    issues = []
    if plan["facts"]["item"] not in case["input"]["message"].lower():
        issues.append(f"the customer's message should mention the {plan['facts']['item']}")
    return issues
```

## fake [#fake]

```python
def fake(values: dict, plan: dict, rng: random.Random) -> dict: ...
```

A free stand-in for Claude, returning content in the `CONTENT_SCHEMA` shape. Required for `--dry-run`. `--estimate` also uses its output size (scaled by 3) to estimate output tokens; without it, the estimate assumes 1,800 output tokens per write.

## Gate order [#gate-order]

For each attempt: build → placeholders, facts used, duplicates → `validate` → the case loads → grader self-test → independent reviewer. A failed attempt is retried once with the issues as feedback; a second failure goes to `cases/_rejected/`. See [Generating cases](/docs/concepts/generating-cases#the-gates).


---

# Check kinds

> Every check kind, its fields, how it matches, and the label it generates.

Source: https://mitej23.github.io/docs/reference/checks



Each `[[check]]` in a `case.toml` has a `kind` and that kind's fields. There are seven kinds: `effect`, `no_effect`, `effect_count`, `state`, `output`, `no_error` and `python`.

Every kind accepts:

<Fields>
  <Field name="kind" type="string">
    One of the seven kinds. Anything else fails the case's load with the list of valid kinds.
  </Field>

  <Field name="name" type="string">
    A label that replaces the generated one in results and the UI.
  </Field>
</Fields>

## Matching [#matching]

`args`, `where` and `equals` use subset matching:

* a dict matches when every expected key is present and matches; extra keys are ignored;
* a list matches element by element and must have the same length;
* numbers match within 1e-9 (booleans are compared exactly);
* anything else must be equal.

## effect [#effect]

At least one matching effect, or exactly `count`.

<Fields>
  <Field name="service" type="string">
    Only effects on this service. Omit to match any.
  </Field>

  <Field name="action" type="string">
    Only this method or function name.
  </Field>

  <Field name="args" type="table" default="{}">
    Subset match on the call's arguments.
  </Field>

  <Field name="ok" type="bool">
    Only calls that succeeded (`true`) or raised (`false`).
  </Field>

  <Field name="count" type="int">
    Exactly this many matching calls. Omit for "at least one".
  </Field>
</Fields>

```toml
[[check]]
kind = "effect"
service = "payments"
action = "refund"
count = 1
args = { order_id = "A100", amount = 24.0 }
```

Label: `payments.refund(order_id=A100, amount=24.0) ×1`, or `… called` without `count`. On failure, the detail lists the calls made to that service and action.

## no_effect [#no_effect]

No matching effect. Same filters as `effect`: `service`, `action`, `args`, `ok`.

```toml
[[check]]
kind = "no_effect"
service = "payments"
```

Label: `no payments`. On failure, the detail lists the offending calls.

## effect_count [#effect_count]

The number of effects equals `count`.

<Fields>
  <Field name="count" type="int" default="0">
    The exact number of effects.
  </Field>

  <Field name="service" type="string">
    Count only this service's effects. Omit to count every effect.
  </Field>
</Fields>

```toml
[[check]]
kind = "effect_count"
count = 0
```

Label: `0 effects on all services`.

## state [#state]

A row in the final state matches, or with `absent`, doesn't exist. Needs a [state adapter](/docs/concepts/state-adapters).

<Fields>
  <Field name="table" type="string">
    The table name. A table missing from the final state fails the check.
  </Field>

  <Field name="where" type="table" default="{}">
    Subset match selecting rows. Empty selects every row.
  </Field>

  <Field name="equals" type="table" default="{}">
    Every selected row must match this subset. Empty passes if any row is selected.
  </Field>

  <Field name="absent" type="bool" default="false">
    Pass when no row matches `where`.
  </Field>
</Fields>

```toml
[[check]]
kind = "state"
table = "orders"
where = { id = "A100" }
equals = { status = "refunded" }

[[check]]
kind = "state"
table = "refunds"
where = { order_id = "A101" }
absent = true
```

Label: `orders[id=A100] status=refunded`. On failure: `no row matches`, or `got {'status': 'delivered'}`.

State checks can also require columns to be empty or set, since TOML has no null: `null = ["calendar_event_id"]`, `not_null = ["calendar_event_id"]`.

## output [#output]

The entry point's output, or a field of it, against one condition.

<Fields>
  <Field name="path" type="string">
    A dotted path into the output: dict keys, list indexes (`messages.0.text`) or attributes. Omit for the whole output.
  </Field>

  <Field name="contains" type="string">
    Case-insensitive substring.
  </Field>

  <Field name="not_contains" type="string">
    Case-insensitive substring that must be absent.
  </Field>

  <Field name="matches" type="regex">
    Case-insensitive regular expression, searched anywhere in the text.
  </Field>

  <Field name="not_matches" type="regex">
    Case-insensitive regular expression that must not match.
  </Field>

  <Field name="equals" type="any">
    The value, compared with subset matching (not converted to text).
  </Field>
</Fields>

Give exactly one condition. If several text conditions are set, only the first of `contains`, `not_contains`, `matches`, `not_matches` is used. Text conditions compare the value as text: a missing value is the empty string; a non-string is converted with `str()`. A check with no condition fails with an error in its detail.

```toml
[[check]]
kind = "output"
path = "decision"
equals = "not_delivered"

[[check]]
kind = "output"
path = "reply"
not_matches = "refund (of|is) .* on its way|refunded"
```

Label: `output.reply doesn't match 'refund (of|is) .* on its way|refunded'`. The detail is the first 160 characters of the text.

In a multi-turn case, `scope = "conversation"` searches every reply the app gave instead of only the last one (`path` is then ignored):

```toml
[[check]]
kind = "output"
scope = "conversation"
contains = "order number"
name = "asked for the order number"
```

## no_error [#no_error]

The entry point returned without raising.

```toml
[[check]]
kind = "no_error"
```

Label: `completed without error`. On failure, the detail is the error, e.g. `NetworkBlocked: …`.

## python [#python]

Your own function.

<Fields>
  <Field name="target" type="string">
    `checks:fn` for a function in the case folder's `checks.py` (any `<module>.py` in the case folder works), or an importable `package.module:fn`.
  </Field>
</Fields>

```python
def fn(result: dict, case: verdict.case.Case) -> tuple[bool, str] | bool: ...
```

`result` is the trial result (`output`, `error`, `traceback`, `effects`, `state`, `clock`). Return `(passed, detail)` or a bool.

```toml
[[check]]
kind = "python"
target = "checks:reply_mentions_amount"
name = "reply states the refunded amount"
```

Label: the target, unless `name` is set.

## nothing_else [#nothing_else]

Nothing happened that the case's other checks don't account for, the way τ-bench compares the whole final database: an effect that no `effect` check matches fails it, and so does a row added, changed or removed that no `state` check names. A changed field passes only if a `state` check matching that row lists it in `equals`.

<Fields>
  <Field name="ignore" type="list of strings" default="[]">
    What may change freely: a table (`"audit_log"`), a field (`"orders.updated_at"`), a service (`"email"`) or a service action (`"email.send_email"`).
  </Field>
</Fields>

```toml
[[check]]
kind = "nothing_else"
ignore = ["orders.updated_at"]
```

Label: `nothing else changed`. The detail lists what was extra, e.g. `payments.refund(order_id=B7, amount=9.0); orders id=A100 changed total`. Opt-in: add it to cases where an extra write is as bad as a missing one.

## judge [#judge]

A yes/no question about what the app said, answered by a model: for rules a pattern can't express, such as "the reply doesn't confirm a time when nothing was booked" (a regex misreads "is already booked").

<Fields>
  <Field name="question" type="string">
    The question about the text, e.g. `"Does the reply tell the vendor the site visit is confirmed?"`.
  </Field>

  <Field name="expect" type="bool" default="true">
    The right answer for this case.
  </Field>

  <Field name="path" type="string">
    Dotted path into the output (default: the whole output). `scope = "conversation"` judges every reply in a multi-turn case instead.
  </Field>
</Fields>

```toml
[[check]]
kind = "judge"
question = "Does the reply say the refund went through?"
expect = false
path = "reply"
```

The detail is the model's answer and its one-line reason. Empty text passes only when `expect = false`. Configure who answers with [`[judge]`](/docs/reference/verdict-toml#judge). It's slower and noisier than the other kinds: keep effects and state in deterministic checks, and give the case an `oracle.json` and `known_bad.json` so the grader self-test judges both.

## Results [#results]

Each check produces:

```json
{"name": "orders[id=A100] status=refunded", "passed": false, "detail": "got {'status': 'delivered'}",
 "kind": "state"}
```

`service` and `action` are copied onto effect checks' results so the web UI can find the step behind a failure. A check that raises fails with `check raised <Error>: <message>`. A multi-turn trial that hits `max_turns`, or whose app raises mid-conversation, gets one more failing check, `the user ended the conversation`.

A trial's reward is the fraction of its checks that passed; it passes only when all of them do.


---

# Python API

> The functions and classes you use from Python, in custom fakes, Python checks, generator hooks and scripts.

Source: https://mitej23.github.io/docs/reference/python-api



Most projects only need the CLI and the file formats. You use the Python API when you write a custom fake, a Python check or generator hooks, or drive Verdict from a script or a test.

<Callout title="0.x">
  The API may change before 1.0. The plugin functions in `verdict.boundary`, `verdict.state`, `verdict.checks` and `verdict.runner` are kept stable where possible. See the [release policy](/docs/project/release-policy).
</Callout>

## In your fakes and checks [#in-your-fakes-and-checks]

### verdict.boundary [#verdictboundary]

#### record [#record]

```python
def record(service: str, action: str, args: dict | None = None, ok: bool = True, result: Any = None) -> dict
```

Append an effect to the current trial and return it. Custom fakes call this for every call they handle, so every effect has the same shape: `{"seq", "service", "action", "args", "ok", "result"}` (plus `span_id` when OpenTelemetry is active). `args` and `result` are stored as JSON; values that aren't JSON-serialisable are converted to strings. Raises `BoundaryError` outside a trial.

#### scripted [#scripted]

```python
def scripted(service: str, action: str, default: Any = None) -> Any
```

The case's scripted response for this call, from `world.services[service][action]`: the `returns` value, the next `sequence` value, or `default` when nothing is scripted. Raises `BoundaryError(message)` for a `raises` script. Returns a deep copy. Raises `BoundaryError` outside a trial.

```python title="my_evals/fakes.py"
from verdict.boundary import record, scripted

class Payments:
    def refund(self, order_id, amount):
        result = scripted("payments", "refund", default={"status": "ok"})
        record("payments", "refund", {"order_id": order_id, "amount": amount}, ok=True, result=result)
        return result
```

#### Errors [#errors]

<Fields>
  <Field name="BoundaryError" type="RuntimeError">
    Raised by a fake: a scripted failure, or a fake called outside a Verdict trial.
  </Field>

  <Field name="NetworkBlocked" type="RuntimeError">
    The app tried to resolve a host that no fake covers and `[network].allow` doesn't list.
  </Field>
</Fields>

### verdict.context [#verdictcontext]

#### trial [#trial]

```python
def trial() -> Trial | None
```

The trial running in the current context, or `None` outside one. Trials of a case run concurrently; this is a `ContextVar` lookup, so it's correct in async tasks and in the threads Verdict runs sync entry points in.

`Trial` fields: `case_id`, `number`, `world` (this trial's copy), `effects` (recorded so far), `calls` (sequence positions), `fakes` (custom fake instances), `state` (the state adapter's handle).

### verdict.case.Case [#verdictcasecase]

The case passed to Python checks and returned by `load_case`.

<Fields>
  <Field name="id, title" type="str">
    From `case.toml`.
  </Field>

  <Field name="path" type="Path">
    The case folder.
  </Field>

  <Field name="input" type="dict">
    `[input]`.
  </Field>

  <Field name="checks" type="list[dict]">
    The `[[check]]` tables.
  </Field>

  <Field name="world" type="dict">
    `world.json`, or `{}`.
  </Field>

  <Field name="entry" type="str">
    The entry name, default `"default"`.
  </Field>

  <Field name="tags" type="list[str]">
    Tags.
  </Field>

  <Field name="trials" type="int | None">
    The case's own trial count.
  </Field>

  <Field name="clock" type="str | None">
    The frozen clock.
  </Field>

  <Field name="meta" type="dict">
    Every key in `case.toml`, including `[generator]`.
  </Field>

  <Field name="oracle, known_bad" type="dict | None">
    The parsed `oracle.json` and `known_bad.json`.
  </Field>
</Fields>

### Python checks [#python-checks]

```python
def my_check(result: dict, case: Case) -> tuple[bool, str] | bool
```

Referenced from a case as `target = "checks:my_check"`. See [Check kinds](/docs/reference/checks#python).

## Driving Verdict from Python [#driving-verdict-from-python]

### verdict.api.Project [#verdictapiproject]

The high-level API: what the CLI does, as methods that return the same dicts as `--json`. Start here for scripts, notebooks, and optimisers that need a judge.

```python
from verdict.api import Project

p = Project("verdict.toml")
base = p.run(label="baseline", trials=3)                  # returns the run id
cand = p.run(label="prompt v2", trials=3, code="prompt-v2")   # a branch, checked out in its own worktree
p.compare(base, cand)["verdict"]                          # "fails the gate", "better beyond noise", …
p.trial(cand, "refund-within-window")["feedback"]         # why it failed, in one string
```

| Method                                                                                           | Returns                                                                            |
| ------------------------------------------------------------------------------------------------ | ---------------------------------------------------------------------------------- |
| `cases(split=None, ids=None)`                                                                    | `list[Case]`, optionally only `"dev"` or `"holdout"`                               |
| `run(cases=None, trials=None, label="", code=None, app_root=None, concurrency=None, quiet=True)` | the run id                                                                         |
| `evaluate(cases=None, trials=2, code=None, app_root=None, allow_holdout=False)`                  | `{run_id, score, cases: [{case, score, passed, trials, mean_reward, feedback}]}`   |
| `compare(baseline, candidate, max_cost_increase=None)`                                           | the `compare --json` object; each side is a run id, `latest`/`latest~N`, or a list |
| `runs()`, `show(run="latest")`                                                                   | as `runs --json` and `show --json`                                                 |
| `trial(run, case, number=None)`                                                                  | the `trial --json` report, with `score` and `feedback`                             |
| `trace(run, case, number=1)`                                                                     | `{steps, blame}` as `trace --json`                                                 |
| `brief(run="latest", cases=None, max_chars=24000)`                                               | the `brief --json` object                                                          |
| `new_candidate(name, from_ref="HEAD")`                                                           | `{name, path, branch, app_root, …}`; run it with `run(candidate=name)`             |
| `rank(baseline, candidates)`                                                                     | the `rank --json` object; each candidate a run id or a list                        |

`evaluate` is shaped for optimisers such as DSPy's GEPA or MIPROv2: per case, a score (the pass rate) and the feedback a reflective optimiser reads, from runs where every side effect stays inside the boundary. The overall `score` is the mean of per-case pass rates. It refuses holdout cases unless `allow_holdout=True`, because a proposer that sees them overfits to the test; keep the holdout for `compare` at the end.

### verdict.config [#verdictconfig]

```python
def load(path: str | Path) -> Config
def resolve(target: str) -> Any          # import "package.module:attr.sub"
```

`load` reads and validates `verdict.toml`, resolving paths relative to it; it raises `ConfigError` with a fix-it message. `Config` fields: `path`, `app_root`, `paths`, `entries`, `entry_modes`, `state`, `boundaries`, `network_allow`, `cases_dir`, `store_path`, `trials`, `concurrency`, `pricing_path`, `agent_names`, `raw`. `config.with_app_root(path)` returns the same config pointed at a candidate checkout.

### verdict.case [#verdictcase]

```python
def load_case(folder: Path) -> Case
def load_cases(cases_dir: Path, ids: list[str] | None = None) -> list[Case]
def holdout_ids(cases_dir: Path) -> set[str]
```

Raise `CaseError` for invalid or duplicate cases, or unknown ids.

### verdict.runner [#verdictrunner]

```python
def run(config: Config, case_ids: list[str] | None = None, trials: int | None = None, label: str = "",
        store: Store | None = None, progress: Callable[[str], None] = print, run_id: str | None = None) -> str
```

Run cases and return the run id. Doesn't self-test first (the CLI does). Each run must be in its own process or run after the previous one finished: the sandbox patches process-wide state.

```python
from verdict import config, runner

cfg = config.load("examples/refunds/verdict.toml")
base = runner.run(cfg, trials=2, label="baseline")
cand = runner.run(cfg.with_app_root(cfg.app_root / "candidate"), trials=2, label="candidate")
```

### verdict.compare [#verdictcompare]

```python
def compare(store, base_ids: list[str], cand_ids: list[str], cases: list[str] | None = None,
            max_cost_increase: float | None = None, holdout: Iterable[str] = ()) -> dict
def render(result: dict) -> str
```

`compare` returns the same structure as `verdict compare --json` (plus each row's raw trial lists). `render` formats it as the CLI prints it.

```python
from verdict.case import holdout_ids
from verdict.compare import compare, render
from verdict.store import Store

result = compare(Store(cfg.store_path), [base], [cand], holdout=holdout_ids(cfg.cases_dir))
print(result["verdict"], result["gate_passed"])
```

### verdict.checks [#verdictchecks]

```python
def grade(result: dict, case: Case) -> dict    # {"passed", "reward", "checks": [...]}
```

### verdict.selftest [#verdictselftest]

```python
def self_test(cases: Iterable[Case]) -> list[dict]     # [{"case", "oracle", "known_bad", "problems"}]
def regrade(store, cases: Iterable[Case], run_ids: list[str] | None = None) -> dict   # {"runs", "trials_changed"}
```

### verdict.store.Store [#verdictstorestore]

```python
store = Store(path)
store.runs()                          # every run, newest first
store.run(run_id)                     # one run, or None
store.run_cases(run_id)               # per-case summaries
store.trials(run_id, case_id=None)    # trials, with result_json and reward_json
store.trial(run_id, case_id, number)
store.spans(trial_id)
```

Rows are `sqlite3.Row`. See [Store schema](/docs/reference/store-schema).

### verdict.candidates [#verdictcandidates]

```python
def fingerprint(root: Path, paths: list[str]) -> dict          # {"code": "8cfd65598bb3", "parts": {path: hash}}
def diff(root: Path, paths: list[str]) -> str                  # uncommitted changes, new files included
def git_info(root: Path, paths: list[str]) -> dict             # {"sha", "branch", "dirty"}
def worktree_for(ref: str, repo: Path, worktrees: Path) -> Path
```

### verdict.tracing [#verdicttracing]

```python
def available() -> bool                # opentelemetry-sdk is installed
def install(export: dict | None = None) -> bool   # add the span collector; with [trace.export] settings, an OTLP exporter too
def flush() -> None                    # send exported spans still queued (the runner calls it at the end of a run)
def trial_span(case_id: str, number: int)   # context manager; yields {"spans": [...]} filled on exit
```

The runner calls these; you rarely need them.

### verdict.observations [#verdictobservations]

```python
def build(spans, prices: dict | None = None, agent_names: dict | None = None, effects: list | None = None) -> dict | None
def blame(trace: dict | None, failed_checks: list[dict], effects: list[dict]) -> list[dict]
```

`build` turns stored spans into the trace view model (rows of typed steps, counts, cost). `blame` returns `{"check", "obs"}` for each failed check, where `obs` is the step behind it.

### verdict.pricing [#verdictpricing]

```python
def load(path: Path | None = None) -> dict     # bundled prices, with a project file merged over
def cost_of(prices: dict, model: str, tokens_in: int, tokens_out: int, cached: int = 0) -> float | None
def usage(spans, prices: dict) -> dict         # {"llm_calls", "tokens_in", "tokens_out", "tokens_cached", "cost_usd"}
```

### verdict.generate [#verdictgenerate]

```python
from verdict.generate import pipeline, spec

s = spec.load(Path("generator/generator.toml"))          # raises spec.SpecError
settings = pipeline.Settings.from_spec(s, dry_run=True, seed=7)
out = pipeline.run(s, cfg.cases_dir, n=5, settings=settings)   # {"summary": {...}, "results": [...]}
est = pipeline.estimate(s, 25, settings)                       # the --estimate numbers
```

`spec.sample(spec, n, seed=7, existing=())` draws scenarios without generating them.


---

# Store schema

> The SQLite tables Verdict saves runs, per-case summaries, trials and spans in.

Source: https://mitej23.github.io/docs/reference/store-schema



Results live in one SQLite file, `[run].store` (default `.verdict/verdict.db`). The schema is additive: columns are added, never dropped. Open it with any SQLite client, or through [`verdict.store.Store`](/docs/reference/python-api#verdictstorestore).

```mermaid
erDiagram
  runs ||--o{ run_cases : summarises
  runs ||--o{ trials : contains
  trials ||--o{ spans : records
```

## runs [#runs]

One row per run.

| Column                      | Type              |                                                                                                                                           |
| --------------------------- | ----------------- | ----------------------------------------------------------------------------------------------------------------------------------------- |
| `id`                        | TEXT, primary key | `YYYYMMDD-HHMMSS-<sha or nogit>-<4 hex>`                                                                                                  |
| `created_at`, `finished_at` | TEXT              | ISO timestamps (UTC)                                                                                                                      |
| `label`                     | TEXT              | `--label`                                                                                                                                 |
| `status`                    | TEXT              | `queued`, `running`, `finished`, `failed` or `cancelled`                                                                                  |
| `pid`                       | INTEGER           | the process running it                                                                                                                    |
| `expected_trials`           | INTEGER           | trials planned                                                                                                                            |
| `config_json`               | TEXT              | what the run tested: Verdict version, config path, app root, paths, git info, code fingerprint, the uncommitted diff, and trials per case |
| `totals_json`               | TEXT              | `{"cases", "trials", "passed"}`                                                                                                           |
| `error`                     | TEXT              | why a run failed or was cancelled                                                                                                         |

## run_cases [#run_cases]

One row per case per run.

| Column              | Type              |                                                      |
| ------------------- | ----------------- | ---------------------------------------------------- |
| `run_id`, `case_id` | TEXT, primary key |                                                      |
| `title`             | TEXT              | the case's title at run time                         |
| `k`                 | INTEGER           | trials run                                           |
| `passed_trials`     | INTEGER           | trials that passed                                   |
| `pass_k`            | INTEGER           | 1 when every trial passed                            |
| `mean_reward`       | REAL              | mean fraction of checks passed                       |
| `cost_usd`          | REAL              | total cost, or NULL when any trial's cost is unknown |

## trials [#trials]

One row per trial.

| Column                       | Type                 |                                             |
| ---------------------------- | -------------------- | ------------------------------------------- |
| `id`                         | INTEGER, primary key |                                             |
| `run_id`, `case_id`, `trial` |                      | unique together; `trial` counts from 1      |
| `passed`                     | INTEGER              | 1 when every check passed                   |
| `reward`                     | REAL                 | fraction of checks passed                   |
| `duration_s`                 | REAL                 | wall time                                   |
| `error`                      | TEXT                 | the app's error, `Type: message`            |
| `cost_usd`                   | REAL                 | model cost from spans, or NULL when unknown |
| `result_json`                | TEXT                 | the trial result (below)                    |
| `reward_json`                | TEXT                 | `{"passed", "reward", "checks": [...]}`     |

### Trial result [#trial-result]

```json
{"schema": 1, "case": "refund-within-window", "trial": 1,
 "output": {"decision": "refunded", "reply": "Done: your refund of $24.00 is on its way."},
 "error": null, "traceback": null,
 "effects": [{"seq": 1, "service": "payments", "action": "refund", "args": {"order_id": "A100", "amount": 24.0},
              "ok": true, "result": {"refund_id": "R-1", "status": "ok"}, "span_id": "56ec21373e40af1a"}],
 "state": {"tables": {"orders": [{"id": "A100", "customer_email": "maya@example.com", "status": "refunded",
                                  "total": 24.0, "delivered_at": "2026-09-28T12:00:00"}]}},
 "clock": "2026-10-08T10:00:00+00:00", "duration_s": 0.006,
 "usage": {"llm_calls": 0, "tokens_in": 0, "tokens_out": 0, "tokens_cached": 0, "cost_usd": 0.0}}
```

## spans [#spans]

One row per OpenTelemetry span kept for a trial (with the `otel` extra).

| Column                       | Type        |                                         |
| ---------------------------- | ----------- | --------------------------------------- |
| `trial_id`, `span_id`        | primary key |                                         |
| `trace_id`, `parent_span_id` | TEXT        | hex ids                                 |
| `name`                       | TEXT        | span name                               |
| `kind`                       | TEXT        | `agent`, `llm`, `tool`, `log` or `span` |
| `start_ns`, `end_ns`         | INTEGER     | nanosecond timestamps                   |
| `status`, `status_message`   | TEXT        | OpenTelemetry status                    |
| `attributes_json`            | TEXT        | every span attribute                    |

## Indexes [#indexes]

`trials_by_case` on `trials(case_id)` and `run_cases_by_case` on `run_cases(case_id)`.


---

# Changelog

> What changed in each Verdict release.

Source: https://mitej23.github.io/docs/changelog



## 0.1.0 (unreleased) [#010-unreleased]

The first version: the core harness, the comparison method, the web app and the case generator, proven on the bundled `examples/refunds` app. Not yet published to PyPI.

### Formats (schema 1) [#formats-schema-1]

* `verdict.toml` with `[app]`, `[entry.<name>]`, `[state]`, `[[boundary]]`, `[network]` and `[run]`, plus optional `[pricing]`, `[trace]`, `[ui]` and `[generator]`.
* Case folders: `case.toml`, `world.json`, optional `checks.py`, `oracle.json` and `known_bad.json`, and `holdout.json` in the cases folder.
* The trial result and reward JSON.

### Isolation [#isolation]

* Boundary fakes for classes (every public method) and functions (swapped everywhere they're referenced, including copies imported by name after the swap), with scripted `returns`, `sequence` and `raises` responses and custom fake classes.
* The network guard: an unhandled outbound host raises `NetworkBlocked`.
* The SQLAlchemy state adapter: a fresh in-memory SQLite database per trial from `world.tables`, with Postgres-only types mapped.
* A frozen clock per case (the `clock` extra).

### Running and grading [#running-and-grading]

* Seven check kinds: `effect`, `no_effect`, `effect_count`, `state`, `output`, `no_error`, `python`, each with a detail that says why it failed.
* The runner: trials per case run concurrently; app errors are recorded, not raised.
* A local SQLite store of runs, per-case summaries, trials and spans.
* Code fingerprints, the uncommitted diff, and candidates from a folder (`--app-root`) or a git ref in its own worktree (`--code`).
* The grader self-test (`verdict self-test`, and before every `verdict run`) and `verdict regrade`.
* `verdict run --split dev|holdout` and `-c/--concurrency`; `verdict view` as an alias of `serve`.
* An optional case `description`, shown on the case's Overview.

### The CLI for coding agents [#the-cli-for-coding-agents]

* `verdict doctor`, `show`, `cases`, `case`, `trial`, `trace`, `state` and `guide`: everything the web app shows about a run, a case or a trial, from the terminal.
* `--json` on every reporting command (one object with `"schema": 1` on stdout; progress and errors on stderr), and `latest` / `latest~N` wherever a run id is taken.
* `verdict guide --install-skill` writes the agent workflow as a Claude Code skill.
* `verdict init --target`, `--generator` (a working generator spec) and `--example refunds` (the example now ships in the package); `verdict serve --open`; an `all` extra.
* Errors are one line (`verdict: …`, exit 1); an entry point that doesn't import says which module is missing, and its run is saved as `failed` instead of staying `running`.

### Setup, multi-turn cases and the Python API [#setup-multi-turn-cases-and-the-python-api]

* `verdict discover` and `verdict init --discover` read the app's code and propose the entry (with the input keys it reads), the SQLAlchemy state, a boundary per outbound client, and model hosts. `init` adds `.verdict/` to `.gitignore`.
* A run names every host the network guard refused (`! blocked HOST`), and `verdict trial` shows them.
* Multi-turn cases: `[entry.<name>] mode = "turns"` and a case's `[user]` table, scripted (`turns`, with `when` branches) or simulated (played by `claude -p`, or your own `simulator`). The transcript and how the conversation ended are in the trial.
* The `nothing_else` check: anything the other checks don't account for fails the case. The `output` check gains `scope = "conversation"`.
* Runs record each case's version and the app's dependencies. `compare` leaves out cases whose setup changed, fails the gate on checks that changed until `regrade`, counts errors, and warns on dependency changes.
* `verdict run --estimate` and `--budget quick|standard|thorough`; `verdict generate --pilot K`; `verdict worktrees [--prune]`.
* `verdict.api.Project`: run, evaluate, compare, show, trial and trace from Python; trial reports carry `score` and `feedback`.
* The example gains a chat front end and a multi-turn case, `chat-asks-for-the-order`.

### Fixes from running Verdict on a production app (ReatAI) [#fixes-from-running-verdict-on-a-production-app-reatai]

* The SQLAlchemy adapter handles `server_default=func.now()` and other function defaults, world rows that leave out different optional columns, UUIDs the app generates, and timezone-aware columns (no more phantom changes).
* A run imports the app from one checkout only. Glue next to `verdict.toml` works for candidates, dotted python checks resolve in `self-test` and `run`, and a candidate run that would load the baseline's code stops instead of mislabelling the result.
* Tasks the entry leaves running (fire-and-forget work) are awaited, up to `[run] background_timeout`.
* A case's `clock` now starts time there and lets it tick, so traces have real durations and order.
* The `[env]` policy: tokens, secrets, passwords, DSNs and database URLs are hidden from the app by default, including from `load_dotenv()`. The network guard covers the whole run, not only trials.

### The fix loop [#the-fix-loop]

* `verdict brief`: the context a coding agent needs to fix a run's failures, size-capped.
* `verdict candidate new|list|diff|remove` and `verdict run --candidate NAME`: each attempted fix in its own git worktree.
* `verdict rank`: rewards for candidates against a baseline, ranked; the Rank page; `Project.brief()`, `new_candidate()`, `rank()`.
* The `verdict-fix` skill.

### Comparing [#comparing]

* `verdict compare`: per-case paired differences, a two-stage bootstrap interval with Beta draws, Holm-adjusted Fisher tests per case, regressed and watch flags, A/A detection, a holdout split, cost and time per trial.
* `--gate` for CI, with `--max-cost-increase`, and `--json` output.

### Traces [#traces]

* OpenTelemetry span collection per trial (the `otel` extra), with kinds inferred from GenAI attributes.
* Observations: spans as typed steps (agent, generation, tool, event), each model call labelled with its decision, and each failed check pointed at the step behind it.
* Model pricing from a bundled `pricing.toml`, overridable per project.
* Per-trial usage: model and tool calls, tokens, models used.
* `[trace.export]`: also send spans to an OTLP/HTTP backend (the `otel-export` extra), with headers read from the environment.

### Web app [#web-app]

* `verdict serve` (the `web` extra): a JSON API and a React client, built from `web/`.
* Pages: Cases; Case (Overview, World, Checks, Runs); Runs; Run; Trial (Outcome, Trace, Data changes, Raw); Compare; Code; a code version (Changes, Runs). Live progress while a run is going.
* The Checks tab's grader table: every check against the known-good and known-bad results.
* Run figures: mean reward, change against the previous run, errored trials, tokens, models, average time per case, commit.
* Data changes grouped per table, with row keys and labels from `[ui] tables`.

### Case generation [#case-generation]

* `verdict generate`: spec-driven case generation through Claude Code (`claude -p`). `plan()` decides each scenario's facts and checks; Claude writes the content; gates (placeholders, facts used, duplicates, the spec's `validate()`, load, the grader self-test) and an independent reviewer decide what is kept. Failed scenarios retry once with feedback, then go to `cases/_rejected/`.
* Gap-filling sampling, `--estimate`, `--dry-run` and `--concurrency`. Generated cases get a description.
* A generator spec for `examples/refunds`.

### Fixed during development [#fixed-during-development]

* A candidate checkout inside the baseline's folder was treated as the baseline's code, so modules weren't re-imported.
* A function imported by name after a fake was installed kept the fake after undo, so the next run inherited a stale fake.
* The generator's facts gate was case-sensitive.
* The generator's reviewer wasn't shown the case's clock, so it rejected correct relative dates.

### Verified [#verified]

* The test suite passes with no API keys or network.
* On `examples/refunds`, the baseline passes 15/15 trials at `k = 3`; the candidate with a 7-day refund window fails the gate with exit 1, flagging `refund-within-window` (3/3 → 0/3) as regressed.
* A real generation batch of 3 cases cost $0.068 API-equivalent over 6 calls in 18 seconds; all 3 were accepted on the first attempt and pass on the real app.
