Metadata-Version: 2.4
Name: pytest-xharness-eval
Version: 0.10.0
Summary: pytest plugin for running the same evaluation suite across AI agent harnesses (cross-harness eval).
Author: Josh Peak
Author-email: Josh Peak <neozenith.dev@gmail.com>
License-Expression: MIT
License-File: LICENSE
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: Framework :: Pytest
Classifier: Topic :: Software Development :: Testing
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Requires-Dist: pytest>=8.3
Maintainer: Josh Peak
Maintainer-email: Josh Peak <neozenith.dev@gmail.com>
Requires-Python: >=3.12
Project-URL: Homepage, https://github.com/neozenith/pytest-xharness-eval
Project-URL: Documentation, https://github.com/neozenith/pytest-xharness-eval
Project-URL: Repository, https://github.com/neozenith/pytest-xharness-eval.git
Project-URL: Issues, https://github.com/neozenith/pytest-xharness-eval/issues
Project-URL: Changelog, https://github.com/neozenith/pytest-xharness-eval/releases
Description-Content-Type: text/markdown

# pytest-xharness-eval 🧪🤖

<p align="center">
    <!-- CICD / Publishing Health -->
    <a href="https://github.com/neozenith/pytest-xharness-eval/actions/workflows/cicd.yml"><img src="https://github.com/neozenith/pytest-xharness-eval/actions/workflows/cicd.yml/badge.svg" alt="CICD Checks"></a>
    <a href="https://github.com/neozenith/pytest-xharness-eval/actions/workflows/publish.yml"><img src="https://github.com/neozenith/pytest-xharness-eval/actions/workflows/publish.yml/badge.svg" alt="Build Status"></a>
    <!-- coverage-badge -->
    <img src="https://img.shields.io/badge/coverage-97%25-brightgreen.svg" alt="Coverage">
    <!-- coverage-badge -->
</p>
<p align="center">
    <!-- project development health -->
    <a href="https://github.com/neozenith/pytest-xharness-eval/graphs/commit-activity"><img alt="GitHub commit activity" src="https://img.shields.io/github/commit-activity/m/neozenith/pytest-xharness-eval"/></a>
    <a href="https://github.com/neozenith/pytest-xharness-eval/issues"><img alt="GitHub open issues" src="https://img.shields.io/github/issues/neozenith/pytest-xharness-eval"/></a>
    <a href="https://github.com/neozenith/pytest-xharness-eval/pulls"><img alt="GitHub open pull requests" src="https://img.shields.io/github/issues-pr/neozenith/pytest-xharness-eval"/></a>
</p>
<p align="center">
    <!-- License and latest info -->
    <a href="https://github.com/neozenith/pytest-xharness-eval/blob/main/LICENSE"><img alt="License" src="https://img.shields.io/github/license/neozenith/pytest-xharness-eval"/></a>
    <a href="https://github.com/neozenith/pytest-xharness-eval/releases"><img src="https://img.shields.io/github/release/neozenith/pytest-xharness-eval" alt="Latest Release"></a>
    <a href="https://pypi.org/project/pytest-xharness-eval/"><img src="https://img.shields.io/pypi/v/pytest-xharness-eval" alt="PyPI"></a>
</p>

<p align="center">pytest plugin for <b>cross</b> AI agent <b>harness eval</b>uation.</p>
<p align="center"><i>Write the eval once. Run it against every harness and every model.</i></p>

<!--TOC-->

- [pytest-xharness-eval 🧪🤖](#pytest-xharness-eval-)
  - [What it does](#what-it-does)
  - [Quickstart](#quickstart)
  - [Narrow a run](#narrow-a-run)
  - [Configuration](#configuration)
  - [How it works](#how-it-works)
  - [What it does not do](#what-it-does-not-do)
  - [Development](#development)
  - [Read next](#read-next)

<!--TOC-->

## What it does

A pytest plugin that runs the `claude` and `codex` CLIs headlessly against a fixture
workspace, captures each run's own session log, prices it, and grades what the agent
left behind. A skill opts in by adding an `evals/` directory; pytest does the rest.

Every eval cell is live and costs money. There is no replay mode. Preview the spend
with `--dry-run` before a sweep. The design rationale lives in
[ARCHITECTURE.md](ARCHITECTURE.md) and the decision log in
[docs/adrs/](docs/adrs/index.md); agents start at [AGENTS.md](AGENTS.md).

----

## Quickstart

1. Install the plugin into the repository that holds your skills. The `pytest11`
   entry point registers it; no `conftest.py` wiring is needed:

   ```sh
   uv add --dev pytest-xharness-eval
   ```

2. Pin pytest's rootdir to the repository root, so the plugin finds `skills/` from
   any argument path. An empty `[tool.pytest.ini_options]` table is enough:

   ```toml
   [tool.pytest.ini_options]
   ```

3. Add an eval beside the skill. Files and functions both carry the `eval_` prefix,
   as `test_` does for pytest. Fixtures are seed workspaces copied fresh for every
   cell; every run output lands under one git-ignored cache root, never in the
   skills tree (ADR 0032):

   ```text
   skills/<skill>/
     SKILL.md
     evals/
       eval_<suite>.py
       fixtures/<name>/
       goldens/<name>/                                  # optional known-good output (ADR 0046)
       treatments/<name>[__<harness>]/                  # optional overlays swept beside the control (ADR 0055)
   .xharness_eval_cache/
     build/                                             # per-cell workspaces
     pricing/prices-YYYYMMDD.toml                       # rates priced live at collection (ADR 0060)
     results/{skill}/{harness}/{model}[--{effort}][+{treatment}]/{run}/{session}/  # log.jsonl, result.json, history.json
     report/                                            # report.json + the aggregated microsite
   ```

   ```python
   from pytest_xharness_eval import CaseOutput, evalcase
   from pytest_xharness_eval.verify import check_files_written, check_rollout

   @evalcase(task="...", skill="<skill>", fixture="<name>")
   def eval_<case>(output: CaseOutput) -> None:
       check_rollout(output)                     # real session, billed, priced
       check_files_written(output, "OUTPUT.md")  # this run is what produced it
       assert "the thing" in output.read("OUTPUT.md")
   ```

   `task` is what a user types *after* naming the skill, it never names the skill, a
   CLI, or where a `SKILL.md` lives. Each harness renders its own invocation around it:
   `/<skill> <task>` for `claude`, `$<skill> <task>` for `codex` (ADR 0044). The full
   grader surface, every field you can assert on, the bundled verifiers, and the goldens
   convention, is [docs/rollout.md](docs/rollout.md).

4. Preview the matrix. Nothing is invoked:

   ```sh
   uv run pytest skills/<skill>/evals --dry-run
   ```

   ```text
   xharness-eval: skills root = /repo/skills, cache = /repo/.xharness_eval_cache
   xharness-eval: matrix = plugin default (11 of 14 catalogued models, output rate below $50/MTok); a case's models= overrides it
   collected 11 items
   skills/<skill>/evals/eval_<case>.py sssssssssss

   ============================ agent eval report ============================
     dry-run          -  skills/<skill>/evals/eval_<case>.py::eval_<case>[claude/claude-haiku-4-5-20251001]
     dry-run          -  skills/<skill>/evals/eval_<case>.py::eval_<case>[claude/claude-sonnet-5]
     ...
     dry-run          -  skills/<skill>/evals/eval_<case>.py::eval_<case>[codex/gpt-6.1-sol]
     total spend: $0.0000 across 11 cell(s)
     report: /repo/.xharness_eval_cache/report/report.json
   ```

5. Run it live, with `-v` so every cell reports its verdict, USD, context, wall
   clock, turns, and tool calls as it lands. Add `-n 2` to run cells in parallel.
   This spends money:

   ```sh
   uv run pytest skills/<skill>/evals -v
   ```

   ```text
   skills/<skill>/evals/eval_<case>.py::eval_<case>[claude/claude-opus-5] PASSED  est $0.5762 (harness $0.5773)  352,451 accumulative_billed_tokens  23,898 baseline_tokens  76.0s  9 turns  8 tools
   ```

   Read the status word as: this plugin's estimate from its price table (and the
   harness CLI's own figure, where it reports one), every billed token summed over
   all turns (the cached prefix is re-read each turn), the harness's own prompt on
   turn 1, wall clock, model calls, tool calls. Every estimate records the rates it
   used and where they came from (`rates_applied`).

   Each cell leaves its verbatim session log (`log.jsonl`), a normalised
   `result.json` with a per-turn ledger, and one `history.json` metrics record in
   its own `results/{skill}/{harness}/{model}[--{effort}]/{run}/{session}/` directory, no two
   cells share a file, so parallel workers never contend (ADR 0032). At session end
   the one combine step aggregates everything under `results/`, every skill, every
   run, into `report/`: `report.json`, the accumulated `history.jsonl`, and a
   browsable `report.html` with its glossary (`XHARNESS-REPORT-GLOSSARY.md`)
   beside it. Serve it with `python3 -m http.server --directory .xharness_eval_cache`
   and open `/report/report.html`; it fetches the JSON beside it.

`run` is a `RunResult`: session id, log path, token usage by tier, tool calls, files
written, and USD cost. The reference case, with its assertions written as a tutorial,
is
`eval_palette_mandate.py`, the reference case kept beside the skill it grades in the consuming repository.

----

## Narrow a run

The matrix is the spend dial. These options are the plugin's own; everything else is
stock pytest (`-k`, `-x`, `-m eval`, node ids).

| Option | Effect | Example |
|--------|--------|---------|
| path | One skill or all of them | `pytest skills/x/evals`, `pytest skills/*/evals` |
| `--harness <name>` | Only cells for that harness (`claude` or `codex`), repeatable | `pytest skills/x/evals --harness codex` |
| `--model <substring>` | Only cells whose model id contains the string, or one exact `harness/model`, repeatable | `pytest skills/x/evals --model opus` |
| `--effort <rung>` | Only cells at that reasoning rung, repeatable. Matches the *resolved* rung, so `--effort max` selects claude's `max` and codex's `xhigh` alike | `pytest skills/x/evals --effort max` |
| `--treatment <name>` | Only cells under that treatment, repeatable. `control` names the untreated cell (ADR 0055) | `pytest skills/x/evals --treatment control` |
| `-k <expr>` | Boolean slices over cell ids and case names (stock pytest) | `-k "opus or sol"`, `-k "codex and not sol"` |
| `--xharness-timeout <s>` | Seconds one cell's CLI may run before it is killed (default 600). Raise it for the top effort rungs, which think for longer by design | `pytest skills/x/evals --xharness-timeout 1800` |
| `--xharness-keep-workspaces` | Leave each finished cell's build workspace in `.xharness_eval_cache/build/` for inspection. By default it is removed once its evidence is captured and graded | `pytest skills/x/evals --xharness-keep-workspaces` |
| `--dry-run` | Enumerate cells and validate pricing, invoke nothing | `pytest skills/x/evals --dry-run` |
| `--collect-only -q` | List cell node ids (stock pytest) | `pytest --collect-only -q skills/x/evals` |

Do not run `pytest skills` from the root: it walks into every skill's `scripts/`
directory and collects their unit tests too. `skills/*/evals` is the full matrix.

A case overrides the matrix with `@evalcase(..., models=["codex/gpt-5.6-sol"])`.

### The effort axis

A matrix entry may name a third component, the reasoning budget the CLI is asked for:

```toml
[tool.pytest.ini_options]
xharness_matrix = """
claude/claude-opus-5/low
claude/claude-opus-5/max
codex/gpt-5.6-sol/mid
"""
```

Each line is a separately graded cell, so one model at two rungs is the cost-versus-quality
comparison the axis exists for. An entry with no third component is unchanged: it leaves
the CLI on whatever default its own configuration gives it.

Each harness declares its own ladder, and three portable aliases name a *position* on it
rather than a level:

| You write | Position | On `claude` | On `codex` |
|-----------|----------|-------------|------------|
| `min` | first rung | `low` | `low` |
| `mid` | middle rung | `high` | `high` |
| `max` | last rung | `max` | `max` |

Both shipped CLIs happen to declare the same five rungs (`low, medium, high, xhigh, max`),
so the aliases resolve identically on each today. That is a fact about these two CLIs and
not a rule: the ladder lives on the harness class, so a third CLI may declare any rungs it
likes and the aliases keep working by position.

Resolution happens once, at collection, so a node id, an evidence directory and a report
row all carry the rung that was actually sent. A rung no harness has
(`claude/claude-opus-5/minimal`) stops the sweep at collection, before anything is spent.
That check is not theoretical: `minimal` appears in codex-cli's own local enum, and a paid
sweep found that no gpt-5.6 model accepts it — the CLI forwards it, the API answers 400,
and the run exits having produced nothing (ADR 0049).

### The treatment axis

A treatment is a directory of files copied over a case's fixture: an `AGENTS.md`, a
`CLAUDE.md`, anything the agent should find in its workspace. It answers "does this change
to the agent's standing instructions change what the same cell costs and how well it does?"

```text
evals/treatments/cheap_eval_subagents/AGENTS.md          # every harness gets this
evals/treatments/cheap_eval_subagents__claude/CLAUDE.md  # claude also gets this: "@AGENTS.md"
```

Name it on the case, or for every case with the `xharness_treatments` ini key:

```python
@evalcase(task=TASK, skill=SKILL, fixture=FIXTURE, treatments=["cheap_eval_subagents"])
```

The axis is opt-in, and every treatment is swept beside its *control*, the same cell with
no treatment. So the case above collects `claude/claude-sonnet-5` and
`claude/claude-sonnet-5+cheap_eval_subagents` as two separately graded cells.
`--treatment control` or `--treatment cheap_eval_subagents` narrows to one arm.

The `__<harness>` directory exists because the CLIs read different files: codex reads
`AGENTS.md`, claude reads `CLAUDE.md`. Each cell reads its workspace's own file and nothing
above it. A treatment with no files for a harness the case sweeps stops collection, because
that cell would be its control billed twice under a second name (ADR 0055).

----

## Configuration

The matrix has three scopes, highest precedence first: a case's `models=`, the
project's `xharness_matrix` ini key, and the plugin's bundled default. The report
header names which one applied.

The bundled default is every model in the **model catalogue** whose output rate is below
`xharness_output_rate_limit` (default `50` USD per million tokens). Today that keeps every
catalogued model except the apex ones, Fable and Astra. Raise the limit to opt in to their
cost, or name them in `xharness_matrix` or a case's `models=`, which the limit never filters
(ADR 0058). Preview the default with `--dry-run` before a first paid run.

### The model catalogue

Every supported model is listed in one config file,
[`src/pytest_xharness_eval/derive/prices/models.toml`](src/pytest_xharness_eval/derive/prices/models.toml),
beside the dated price records. Each model carries three facts that every record stores
(ADR 0057, ADR 0059):

| Fact | Example | Meaning |
|------|---------|---------|
| `line` | `opus`, `sol` | The provider's own product line |
| `family_tier` | `3` | The model's role in its lineup, 1 the smallest, curated by hand. A number rather than a name, so it survives a lineup change, and frozen at release, so history stays comparable |
| `released` | `2026-07-24` | The release date |

Today's tiers: 1 is Haiku and Luna, 2 is Sonnet and Terra, 3 is Opus and Sol, 4 is Fable
and Astra. A tier is a role, never a price, which is what lets a report compare every tier 3
model across providers.

A matrix entry naming a model the catalogue does not list stops collection before anything
is spent. Add a new model before a plugin release with one ini line:

```ini
xharness_models =
    codex/gpt-6.2-sol: line=sol tier=3 released=2026-10-20
```

### Live pricing

A catalogued model with no bundled price row is **priced live** at collection. Its rates are
looked up in LiteLLM's price feed by exact first-party id, and saved as a dated record under
`.xharness_eval_cache/pricing/`. The sweep and every later replay read that record, so a run
priced live is re-priced identically. A model the feed does not price either still stops
collection, and nothing is guessed (ADR 0060). With every model priced, nothing is fetched.

The bundled records are the offline default. In this repository, `make test` first curates a
new bundled snapshot (`make prices`) whenever `models.toml` lists a model they do not price.

The ini keys, paths relative to pytest's rootdir:

| Key | Default | Purpose |
|-----|---------|---------|
| `xharness_matrix` | (plugin default) | Project matrix: `harness/model` or `harness/model/effort` entries every case sweeps unless it sets `models=` |
| `xharness_output_rate_limit` | `50` | The plugin default matrix sweeps only catalogued models whose output rate, in USD per million tokens, is below this. Raise it to opt in to apex models (ADR 0058) |
| `xharness_price_feed` | LiteLLM's feed | Where a model with no price row is priced live from at collection: a URL or a local path (ADR 0060) |
| `xharness_models` | (none) | Model catalogue rows that add or correct a model before a plugin release: `<harness>/<model>: line=<line> tier=<n> released=YYYY-MM-DD [effort=false]`; `effort=false` marks a model whose CLI ignores a reasoning rung, so a matrix entry naming one is refused (ADR 0057, ADR 0063) |
| `xharness_treatments` | (none) | Treatment names under each suite's `evals/treatments/`, swept beside the untreated control unless a case sets `treatments=` (ADR 0055) |
| `xharness_skills_dir` | `skills` | Directory holding `<skill>/evals/` trees |
| `xharness_cache_dir` | `.xharness_eval_cache` | The git-ignored root for build workspaces, results and the report (ADR 0032) |
| `xharness_skill_ignore` | (none) | gitignore-style patterns for skill files that are not decision surface; a bare pattern applies to every skill, `<skill>: <pattern>` to the skills matching the selector (ADR 0026) |
| `xharness_report_design_tokens` | bundled | design tokens JSON that themes `report/report.html` (flag: `--xharness-report-design-tokens FILE`) |
| `xharness_report_inline` | `false` | embed every result, log and the tokens into `report/report.html` so it opens over `file://` (flag: `--xharness-report-inline`) |
| `xharness_timeout_s` | `600` | Seconds one cell's CLI may run before it is killed (flag: `--xharness-timeout SECONDS`). A killed run is captured and priced: it **fails** if its session was still active in the 5 minutes before the limit (it ran out of time), and **errors** if it had been silent longer (it stalled) (ADR 0064) |
| `xharness_keep_workspaces` | `false` | Leave each finished cell's build workspace in place for inspection instead of removing it once its evidence is captured and graded (flag: `--xharness-keep-workspaces`, ADR 0062) |
| `xharness_prices` | (none) | Price rows that add to or override the bundled price records: `<harness>/<model>: input=<usd/MTok> output=<usd/MTok> [cache_read=..] [cache_write=..] [cache_write_1h=..] [long_context_above=<prompt tokens> long_input=.. long_output=.. [long_cache_read=..] [long_cache_write=..] [long_cache_write_1h=..]] [from=YYYY-MM-DD] [to=YYYY-MM-DD]` (ADR 0030, ADR 0050, ADR 0051) |

```toml
[tool.pytest.ini_options]
xharness_matrix = [
    "claude/claude-opus-5",
    "claude/claude-sonnet-5",
    "claude/claude-haiku-4-5-20251001",
    "codex/gpt-5.6-luna",
    "codex/gpt-5.6-terra",
    "codex/gpt-5.6-sol",
]
```

An unpriced model stops the sweep at collection, before any spend. Add a price row
to the same ini block, naming the harness that runs the model, in USD per million
tokens (ADR 0030, ADR 0050). A row with `from=`/`to=` applies only to runs stamped
inside `[from, to)`:

```toml
xharness_prices = [
    "codex/gpt-5.6-luna: input=0.20 output=1.20 cache_read=0.02 long_context_above=272000 long_input=0.40 long_output=1.80",
    "claude/claude-sonnet-5: input=3.00 output=15.00 from=2026-10-01",
]
```

Some providers bill a long prompt at higher rates: OpenAI prices a call whose prompt
exceeds 272K tokens (cached input included) at the long-context rates for the whole
call. State that tier with `long_context_above` and the `long_*` keys. Every call is
priced on its own prompt, from the per-call ledger, so one long call in a run is billed
correctly beside many short ones (ADR 0051).

The bundled rates are dated records, one `derive/prices/prices-YYYYMMDD.toml` per
interval, curated from LiteLLM's feed with `make prices`. Every run is priced from the
record in effect on the day it ran, so a replay reproduces the bill it had then. Each
estimate's `rates_applied` names the record, its interval and any long-context tier,
and `long_context_calls` counts the calls billed at that tier.

----

## How it works

```mermaid
flowchart LR
    CASE["eval_*.py case"]
    PLUG["plugin/<br/>collect, expand matrix"]
    WS["model/workspace.py<br/>pristine copy"]
    RUN["harness/<br/>ClaudeHarness | CodexHarness"]
    LOG["session log<br/>this run's own"]
    NORM["SessionLog.to_result<br/>RunResult"]
    PRICE["runtime/pipeline.derive<br/>price, coverage, case"]
    GRADE["case assertions"]
    REP["report.json"]

    CASE --> PLUG --> WS --> RUN --> LOG --> NORM --> PRICE --> GRADE --> REP

    classDef new fill:#7c3aed,color:#fff
    classDef data fill:#0f766e,color:#fff
    classDef good fill:#047857,color:#fff
    class CASE,PLUG,WS,RUN new
    class LOG,NORM data
    class PRICE,GRADE,REP good
```

One cell flows left to right: a case is expanded into cells, each cell gets a fresh
workspace, the CLI runs, its own log is located and normalised, priced, graded, and
reported. The hard part is the middle: the two CLIs need different contracts to tie a
verdict to the right log. [ARCHITECTURE.md](ARCHITECTURE.md) explains both.

----

## What it does not do

- It does not replay recorded runs. Every cell invokes the real CLI (ADR 0002).
- It does not give the agent a git repository. The workspace is a plain copy of the
  fixture, so skills that read git history are out of scope (ADR 0004).
- It does not mock either CLI, in tests or in evals.
- It does not price an unknown model as zero. It refuses to run (ADR 0007).
- It does not throttle providers. `-n N` runs N cells at once; each cell is
  isolated (own workspace, own `CODEX_HOME`, own Claude session). If a provider
  rate-limits you, `-n 2 --dist loadgroup` keeps each harness's cells on one
  worker (parallel across harnesses, serial within one).

----

## Development

Build, test and release instructions live in [CONTRIBUTING.md](CONTRIBUTING.md).

----

## Read next

- [docs/rollout.md](docs/rollout.md): what a rollout leaves you, the `CaseOutput` a grader is handed, every `RunResult` field it can assert on, the bundled `check_*` verifiers, and the goldens convention
- [docs/token-accounting.md](docs/token-accounting.md): how `accumulative_billed_tokens` (billed across turns) and `peak_context_tokens` (the largest prompt) are derived from what each provider reports, with a worked session
- [ARCHITECTURE.md](ARCHITECTURE.md): why the two CLIs need different capture
  contracts, how pricing works, and the vocabulary the code uses.
- [AGENTS.md](AGENTS.md): operating instructions and hard boundaries for agents.
- [docs/adrs/index.md](docs/adrs/index.md): the decision index, generated from the records (ADR 0047).
