Metadata-Version: 2.5
Name: verdict-evals
Version: 0.1.0
Summary: Regression testing for AI agent products: run your real code with side effects contained, grade outcomes, and get an honest verdict on every change.
License-Expression: Apache-2.0
License-File: LICENSE
Requires-Python: >=3.11
Provides-Extra: all
Requires-Dist: fastapi>=0.110; extra == 'all'
Requires-Dist: opentelemetry-exporter-otlp-proto-http>=1.24; extra == 'all'
Requires-Dist: opentelemetry-sdk>=1.24; extra == 'all'
Requires-Dist: sqlalchemy>=2.0; extra == 'all'
Requires-Dist: time-machine>=2.14; extra == 'all'
Requires-Dist: uvicorn>=0.29; extra == 'all'
Provides-Extra: clock
Requires-Dist: time-machine>=2.14; extra == 'clock'
Provides-Extra: dev
Requires-Dist: fastapi>=0.110; extra == 'dev'
Requires-Dist: httpx>=0.27; extra == 'dev'
Requires-Dist: opentelemetry-exporter-otlp-proto-http>=1.24; extra == 'dev'
Requires-Dist: opentelemetry-sdk>=1.24; extra == 'dev'
Requires-Dist: pytest>=8; extra == 'dev'
Requires-Dist: python-dotenv>=1.0; extra == 'dev'
Requires-Dist: sqlalchemy>=2.0; extra == 'dev'
Requires-Dist: time-machine>=2.14; extra == 'dev'
Requires-Dist: uvicorn>=0.29; extra == 'dev'
Provides-Extra: otel
Requires-Dist: opentelemetry-sdk>=1.24; extra == 'otel'
Provides-Extra: otel-export
Requires-Dist: opentelemetry-exporter-otlp-proto-http>=1.24; extra == 'otel-export'
Requires-Dist: opentelemetry-sdk>=1.24; extra == 'otel-export'
Provides-Extra: sqlalchemy
Requires-Dist: sqlalchemy>=2.0; extra == 'sqlalchemy'
Provides-Extra: web
Requires-Dist: fastapi>=0.110; extra == 'web'
Requires-Dist: uvicorn>=0.29; extra == 'web'
Description-Content-Type: text/markdown

# Verdict

Regression testing for AI agent products. Verdict runs your app's real code path against a
recorded world, with the database, external APIs, messages and clock contained at the edges.
It grades what your app *did*, and tells you whether a code change made things better or worse,
beyond noise.

```bash
uv sync --extra dev
cd examples/refunds
uv run verdict run -k 3 --label baseline
uv run verdict run -k 3 --app-root candidate --label candidate
uv run verdict compare <baseline-run> <candidate-run> --gate     # exit 1: the candidate regressed

uv run verdict generate --dry-run -n 5     # write cases from the generator spec (drop --dry-run to use claude -p)
(cd ../../web && npm install && npm run build) && uv run verdict serve   # the web app on :8765
```

- **Your code, our edges.** Nothing in your app is mocked except its boundaries.
- **Outcomes over wording.** Checks run on effects and final state.
- **Honest statistics.** The case is the unit: an A/A noise floor, Holm-adjusted per-case
  tests, a holdout split and a regression gate.
- **Graders you can trust.** Each case can carry a known-good and a known-bad result; a grader
  that passes the bad one or fails the good one stops the run.
- **Traces.** OpenTelemetry spans per trial as agent, model and tool steps, with the step
  behind every failed check; optionally exported to your own backend.
- **Cases at scale.** `verdict generate`: code decides each scenario's correct outcome, Claude
  Code writes the content, gates and a reviewer decide what is kept.

| Where | What |
|---|---|
| `src/verdict/` | the package and the `verdict` CLI |
| `web/` | the front end: the product UI, the landing page and the user docs (`web/content/docs`), one Vite app built into the package |
| `docs/` | product, architecture, decisions and learnings, for contributors |
| `examples/refunds/` | a deterministic example app with cases and a generator spec |

Status: early (0.1.0, unreleased), APIs may change. Licence: Apache-2.0.
