# Autobench Full Documentation

> Complete, LLM-readable documentation for Autobench: define how an application should be measured, then run, record, compare, and replay its benchmarks in one consistent format.

Canonical documentation: https://vcoderun.github.io/autobench/

This file follows the site navigation order. Each section includes the canonical page URL followed by its complete Markdown source.

---

## Home

Canonical page: https://vcoderun.github.io/autobench/

# Autobench

**Define what better means once. Autobench evaluates every version by the same rules, compares the
results, and keeps the full history ready for inspection.**

Applications change: a team may switch models, revise a prompt, replace a tool, tune an algorithm,
or ship a new configuration. To decide whether the change is actually better, they commonly write
a benchmark script. That script runs representative inputs, checks the outputs, records values
such as correctness, latency, token usage, or cost, and compares one version with another.

The script works, but every project tends to build this machinery again. Results use incompatible
formats, measurement and scoring logic become mixed with application code, and an old result is
often impossible to inspect without rerunning the original program.

We built Autobench to solve this problem. You describe the inputs to test, the variants to compare,
the application task, and the meaning of success in YAML or Python. Autobench then runs the full
matrix, collects measurements and traces, evaluates each result, records the application assets
that affected it, and stores an immutable experiment record. That same record can be replayed,
reported, compared, or exported later without calling the application again.

```bash
uv add autobench
autobench validate autobench.yaml
autobench run autobench.yaml --record runs/latest
```

After the run, inspect the recorded experiment without executing the application again:

```bash
autobench replay runs/latest
autobench report runs/latest
autobench compare runs/latest --baseline current --candidate proposed
autobench export runs/latest --format csv --path analysis/runs.csv
```

## The Framework Loop

```text
BenchmarkSpec
  Dataset[Case] x Variant[Factor]
    -> task(ctx, case)
    -> observations + ABP trace + artifacts + asset versions
    -> scorers + per-run derivation
  -> cross-run derivation + policies
  -> immutable RunRecords
  -> replay + Rich reports + comparison + exports + optimization feedback
```

The task is the only application-specific part. Autobench owns the repeated infrastructure around
it: matrix planning, context propagation, instrumentation, scoring, derivation, persistence,
reporting, and replay.

## Why Semantic Evidence Matters

Raw names such as `prompt_tokens`, `input_tokens`, `accuracy`, and `answer_quality` are local
conventions. Autobench observations can also declare stable meaning:

```text
llm.tokens.input
llm.tokens.output
llm.model.name
quality.correctness
time.latency
money.cost
agent.tool.argument.correctness
```

That semantic layer lets reports, pricing derivation, policy checks, and optimization systems use
evidence from different applications without guessing what every local metric name means.

## What You Can Benchmark

Autobench is optimized for AI systems but does not require one:

| System | Cases | Variants | Evidence |
| --- | --- | --- | --- |
| LLM application | prompts and expected answers | model, prompt, temperature | quality, tokens, latency, cost |
| Agent | user goals and expected actions | instructions, tools, model | action selection, arguments, sequence, completion |
| Search or retrieval | queries and relevant items | index, reranker, limits | recall, precision, latency |
| Service/API | requests and expected responses | release, configuration | correctness, errors, throughput, SLA |
| Algorithm | input fixtures | implementation | correctness, repeated timings, speedup |
| Data pipeline | source batches | parser or policy | coverage, validity, loss, runtime |

See [Use Cases](use-cases.md) for complete patterns.

## Core Capabilities

| Area | Included |
| --- | --- |
| Definition | Human-readable YAML DSL, typed Python builder, JSON Schema completion |
| Data | Inline/file/glob datasets, defaults, attachments, generated and production cases |
| Execution | Sync/async tasks, deterministic matrices, bounded concurrency, failure isolation |
| Evidence | Semantic observations, checks, events, artifacts, measurements, ABP traces |
| Evaluation | Built-in and custom scorers, expected actions, policies, metric packs |
| Derivation | Token cost, tiered pricing, paired baselines, comparison verdicts |
| Instrumentation | Manual spans, method instrumentation, Pydantic AI, pydantic-gepa, OpenAI, Agents, HTTPX |
| Lineage | Explicit and automatic prompt/tool/schema/capability/agent asset versioning |
| Persistence | Immutable YAML records, source hashes, environment metadata, portable artifacts |
| Analysis | Replay, Rich reports, leaderboards, matrices, distributions, comparisons, exports |

## Choose A Starting Point

| Goal | Read |
| --- | --- |
| Install the right extras | [Installation](installation.md) |
| Run a complete benchmark | [First Benchmark](getting-started.md) |
| Find a pattern for your system | [Use Cases](use-cases.md) |
| Understand ownership and data flow | [Architecture](architecture.md) |
| Author the full DSL | [YAML Spec](yaml-spec.md) |
| Compose benchmarks in Python | [Python API](python-api.md) |
| Instrument an existing SDK application | [Native Instrumentation](native-instrumentation.md) |
| Record an optimizer run and candidate lineage | [Pydantic-GEPA Instrumentation](pydantic-gepa-instrumentation.md) |
| Collect prompt/tool/schema lineage automatically | [Automatic Asset Discovery](automatic-asset-discovery.md) |
| Inspect all shipped features | [Capability Map](capabilities.md) |

## Project Boundaries

Autobench records and evaluates evidence. It does not own your application, make causal claims from
confounded runs, keep provider pricing permanently current, or choose an optimization algorithm.
Those boundaries keep the core usable for arbitrary systems while allowing pydantic-gepa,
autoptimize, or another consumer to build on stable experiment records.

---

## Installation

Canonical page: https://vcoderun.github.io/autobench/installation/

# Installation

Autobench supports Python 3.11 through 3.14. The base package includes the benchmark DSL, runtime,
semantic evidence models, evaluation, recording, replay, reports, CLI, and manual ABP spans.

## Base Package

=== "uv"

    ```bash
    uv add autobench
    ```

=== "pip"

    ```bash
    python -m pip install autobench
    ```

Verify the installation:

```bash
autobench --help
python -c "import autobench; print(autobench.__version__)"
```

## Optional SDK Integrations

Native ABP instrumentors are optional so a generic benchmark does not install AI SDKs.

```bash
uv add 'autobench[instrumentation]'
```

The instrumentation extra supplies the supported Pydantic AI, OpenAI Python, and HTTPX integration
environment. OpenAI Agents support has its own extra:

```bash
uv add 'autobench[openai-agents]'
```

Install only the native pydantic-gepa optimizer integration with:

```bash
uv add 'autobench[pydantic-gepa]'
```

It records optimizer lifecycle, evaluation evidence, budgets, candidate lineage, and component
asset versions. See [Pydantic-GEPA Instrumentation](pydantic-gepa-instrumentation.md).

Inspect what the current environment can instrument:

```bash
autobench instrumentation doctor
```

The command reports compatibility rather than failing because an optional SDK is absent.

To export immutable ABP records to an OTLP HTTP/protobuf backend, install the independent exporter
extra:

```bash
uv add 'autobench[otlp]'
```

Collection still uses ABP. The extra is needed only on the process that performs
`autobench telemetry export`; see [OTLP Export](otlp-export.md).

## Development Checkout

From the repository root:

```bash
uv sync --extra dev --extra instrumentation --extra openai-agents --extra otlp
make prod
```

Useful targets:

| Command | Purpose |
| --- | --- |
| `make tests` | Test suite with source line and branch coverage |
| `make check` | Ruff, ty, and basedpyright |
| `make docs` | LLM bundles and strict Zensical build |
| `make examples` | Offline end-to-end example matrix |
| `make prod` | Full supported-Python and release quality gates |
| `make pre-commit` | Repository-wide hooks |

## Editor Setup For YAML

Every exported Autobench YAML document starts with a `yaml-language-server` schema directive.
Versioned schemas are shipped under `schemas/<autobench-version>/` and installed to the user schema
directory when the schema helpers run.

For a repository-local benchmark:

```yaml
# yaml-language-server: $schema=./schemas/0.3.0/benchmark_schema.json
benchmark:
  smoke-test:
    cases: []
```

Use the schema matching the Autobench version that validates and executes the file. This provides
completion for scorer variants, policy operators, instrumentation settings, report configuration,
and semantic registry entries.

## Credential Handling

Autobench itself does not require model credentials. Live examples read provider configuration from
the relevant SDK environment. For example:

```bash
export OPENROUTER_API_KEY=...
export OPENROUTER_MODEL=openrouter:openai/gpt-5.6-luna
```

Do not place credentials in benchmark specs, cases, artifacts, or capture policies. Use the
[capture policy](automatic-asset-discovery.md#privacy-and-capture-policy) to prevent sensitive SDK
inputs from being retained. Runtime evidence defaults to metadata, but versioned behavioral assets
default to full content and are stored in `artifacts/asset-content.sqlite3`; set
`asset_default_level: hash` when that local registry must not retain prompt, tool, or schema bodies.

## Next Step

Continue with [First Benchmark](getting-started.md), which creates a task, dataset, variant matrix,
score, record, report, comparison, and export.

---

## First Benchmark

Canonical page: https://vcoderun.github.io/autobench/getting-started/

# First Benchmark

This guide builds a complete deterministic benchmark. It compares two text transformations, scores
their outputs, records every case and variant, and replays the result.

## Project Layout

```text
text-benchmark/
  autobench.yaml
  benchmark_task.py
```

The task remains ordinary application code. The YAML file describes how Autobench should execute
and evaluate it.

## Write The Task

Create `benchmark_task.py`:

```python
from __future__ import annotations

from typing import Literal

from pydantic import BaseModel, TypeAdapter

from autobench import Case, RunContext

Transform = Literal["upper", "title_upper"]
TRANSFORM = TypeAdapter(Transform)


class TextInput(BaseModel):
    text: str


class TextOutput(BaseModel):
    text: str


def run(ctx: RunContext, case: Case) -> TextOutput:
    sample = TextInput.model_validate(case.input)
    transform = TRANSFORM.validate_python(ctx.factor("transform"))

    with ctx.span(
        "transform_text",
        kind="workflow",
        input=sample.model_dump(),
        attributes={"transform": transform},
    ) as span:
        text = sample.text.upper()
        if transform == "title_upper":
            text = sample.text.title().upper()
        output = TextOutput(text=text)
        span.set_output(output.model_dump())
        return output
```

The required task signature is `task(ctx, case)`: `RunContext` is always first and `Case` is always
second. Sync and async functions are both supported. Span duration is measured by Autobench.

## Define The Benchmark

Create `autobench.yaml`:

```yaml
# yaml-language-server: $schema=./schemas/0.3.0/benchmark_schema.json
benchmark:
  text-transform:
    description: Compare deterministic text transformations.
    cases:
      - id: greeting
        input:
          text: hello autobench
        expected:
          text: HELLO AUTOBENCH
      - id: whitespace
        input:
          text: release ready
        expected:
          text: RELEASE READY
    run:
      python: benchmark_task:run
    variants:
      current:
        label: Current implementation
        factors:
          transform: upper
      proposed:
        label: Proposed implementation
        factors:
          transform:
            value: title_upper
            optimize: true
    score:
      exact_text:
        exact:
          actual: output.text
          expected: case.expected.text
        semantic: quality.correctness
        goal: maximize
        role: objective
    report:
      leaderboard:
        show:
          correctness:
            metric: quality.correctness
            aggregate: ratio_true
      matrix:
        metric: quality.correctness
      compare:
        current -> proposed:
          show:
            correctness:
              metric: quality.correctness
              aggregate: ratio_true
```

This produces four runs: two cases multiplied by two variants.

## Validate Before Running

From `text-benchmark/`:

```bash
autobench validate autobench.yaml
```

Validation parses the DSL, resolves the task and source files relative to the spec, loads external
datasets and pricing files, verifies unique IDs, and displays the planned matrix. It does not call
the task.

## Run And Record

```bash
autobench run autobench.yaml --record runs/text-transform
```

Autobench renders Rich terminal tables and writes:

```text
runs/text-transform/
  experiment.yaml
  summary.yaml
  cases/
    greeting/current/run.yaml
    greeting/proposed/run.yaml
    whitespace/current/run.yaml
    whitespace/proposed/run.yaml
  artifacts/
```

The actual run filenames use stable run IDs inside the case and variant directories. The records
include the case snapshot, factors, output, observations, score, trace, source hashes, environment,
and status.

## Replay And Analyze

```bash
autobench replay runs/text-transform
autobench report runs/text-transform
autobench compare runs/text-transform --baseline current --candidate proposed
```

These commands load recorded evidence. They do not import `benchmark_task.py` and do not execute the
subject again. Comparison reports factor changes and metric deltas but does not claim that a
confounded difference is causal.

## Export A Projection

```bash
autobench export runs/text-transform \
  --format yaml \
  --path analysis/text-transform.yaml

autobench export runs/text-transform \
  --format csv \
  --path analysis/text-transform.csv
```

Terminal output stays human-oriented and uses Rich tables. YAML, CSV, and Markdown are file export
formats.

## Add Runtime Evidence

Tasks can emit evidence that is not part of the return value:

```python
ctx.metric(
    "characters",
    len(output.text),
    semantic_type="text.characters",
    unit="count",
)
ctx.check("not_empty", bool(output.text), reason="The transformed text must not be empty.")
ctx.artifact("output", output.model_dump(), media_type="application/yaml")
```

Use scores for evaluation results, observations for runtime facts, and artifacts for payloads that
must remain inspectable.

## Run Concurrently

```bash
autobench run autobench.yaml \
  --concurrency 4 \
  --record runs/text-transform-concurrent
```

The matrix order and run IDs remain deterministic. ABP context is task-local, so concurrent runs do
not share parent spans or evidence.

## Next Steps

- Move cases to a file: [Datasets And Variants](datasets-and-variants.md)
- Add quality, cost, and policy gates: [Scoring And Derivation](scoring-and-derivation.md)
- Instrument an SDK automatically: [Native Instrumentation](native-instrumentation.md)
- Track prompt and tool versions: [Automatic Asset Discovery](automatic-asset-discovery.md)
- Select a complete pattern: [Use Cases](use-cases.md)

---

## Use Cases

Canonical page: https://vcoderun.github.io/autobench/use-cases/

# Use Cases

The same Autobench runtime supports deterministic functions, services, LLM applications, agents,
and performance experiments. The patterns below show where domain code ends and framework
infrastructure begins.

## Choose A Pattern

| Need | Core primitives |
| --- | --- |
| Compare implementations | cases, variants, exact/pass scorers, comparison |
| Measure noisy performance | `measure_callable`, sample artifacts, paired baseline |
| Compare LLM quality and cost | semantic token metrics, pricing derivation, policies |
| Evaluate agent behavior | ABP tool spans, expected actions, span selectors |
| Instrument an existing AI app | `instrument_all()`, native SDK instrumentors |
| Track prompts/tools/schemas | explicit tracking or automatic asset discovery |
| Turn production failures into regressions | `ProductionSample`, sampling policy, reviewed cases |
| Feed an optimizer | objectives, constraints, factors, asset versions, feedback records |
| Publish evidence to telemetry | immutable records, optional ABP-to-OTLP export |

## Application Regression Benchmark

Use a file-backed dataset when the benchmark is a maintained regression suite:

```yaml
# yaml-language-server: $schema=./schemas/0.3.0/benchmark_schema.json
benchmark:
  support-routing:
    dataset:
      source: file://datasets/tickets.yaml
      version: "2026-08-06"
      defaults:
        tags: [regression]
    run:
      python: benchmark_tasks:route_ticket
    variants:
      production:
        factors:
          routing_profile: v3
      candidate:
        factors:
          routing_profile:
            value: v4
            optimize: true
    score:
      route:
        exact:
          actual: output.queue
          expected: case.expected.queue
        semantic: quality.correctness
        goal: maximize
        role: objective
      handled:
        pass: output.handled
        semantic: result.success
        role: constraint
```

Keep routing logic in `benchmark_tasks.py`. Autobench handles matrix expansion, status isolation,
score projection, and comparison. This pattern also fits parsers, validators, ranking functions,
API clients, and data transformations.

## Repeated Performance Measurement

Do not hand-roll warmup, repetition budgets, percentiles, or sample artifacts:

```python
from autobench import Case, RunContext, Semantic, measure_callable


def run(ctx: RunContext, case: Case) -> dict[str, bool]:
    values = list(case.input["values"])
    target = int(case.input["target"])
    strategy = str(ctx.factor("strategy"))

    def execute() -> None:
        if strategy == "linear":
            target in values
        else:
            target in set(values)

    measurement = measure_callable(
        execute,
        warmup=3,
        repetitions=25,
        max_seconds=2.0,
    )
    ctx.record_measurement(
        "lookup_latency",
        measurement,
        semantic_type=Semantic.TIME_LATENCY,
        include_samples_artifact=True,
    )
    return {"found": target in values}
```

Derive candidate speedup only after both matched runs exist:

```yaml
post_derive:
  - kind: paired_baseline
    baseline_variant: linear
    match_on: case_id
    metric: time.latency
    formula: baseline_over_candidate
    include_baseline: true
    output:
      name: speedup
      semantic_type: performance.speedup
      unit: ratio
      direction: maximize
      role: objective
```

Correctness should remain a constraint. A faster wrong implementation is not a successful
candidate.

## LLM Quality, Usage, And Cost

Instrumentors or tasks record usage as semantic observations:

```python
ctx.metric(
    "input_tokens",
    usage.input_tokens,
    semantic_type="llm.tokens.input",
    unit="token",
)
ctx.metric(
    "output_tokens",
    usage.output_tokens,
    semantic_type="llm.tokens.output",
    unit="token",
)
ctx.factor_observation("model", model_name, semantic_type="llm.model.name")
ctx.factor_observation("provider", provider, semantic_type="llm.provider")
```

Cost remains a derivation instead of being hard-coded into Autobench instrumentation:

```yaml
derive:
  - kind: token_cost
    pricing: file://pricing/models.yaml
    output:
      name: request_cost
      semantic_type: money.cost
      unit: usd
      direction: minimize
      role: constraint
policies:
  - name: quality-floor
    metric: quality.correctness
    must_greater_equal: 0.9
  - name: per-request-budget
    metric: money.cost
    must_less_equal: 0.01
```

The pricing file can normalize provider-specific model identifiers, aliases, cache prices, and
tiered input/output rates. Price sources are convenience adapters into this format; Autobench does
not become a live pricing service.

## Existing Pydantic AI Application

For a Pydantic AI application, automatic instrumentation removes task-level telemetry:

```python
from autobench import Benchmark, Case, ExactScorer, Semantic

benchmark = (
    Benchmark("support-agent")
    .dataset(
        [
            Case(
                id="order-status",
                input="Where is order A-42?",
                expected={"status": "delayed"},
            )
        ]
    )
    .variants(
        [
            {
                "id": "luna",
                "factors": {
                    "model": "openrouter:openai/gpt-5.6-luna",
                },
            }
        ]
    )
    .task("support_benchmark:run")
    .scoring(
        [
            ExactScorer(
                name="status",
                actual="output.status",
                expected="case.expected.status",
                semantic_type=Semantic.QUALITY_CORRECTNESS,
            )
        ]
    )
    .instrument_all()
)
```

The task can contain only the agent call. Compatible instrumentors collect Pydantic AI agent/model/
tool/validation activity, the OpenAI-compatible client layer, and HTTPX transport evidence. The run
also receives automatically discovered prompt, tool, output-schema, capability, and agent versions
when those values cross supported SDK boundaries.

Use `exclude={"httpx"}` to avoid transport spans or select a narrower asset family:

```python
benchmark.instrument_all(
    exclude={"httpx"},
    assets={
        "representations": ["definition", "effective"],
        "include": ["prompt", "tool", "output_schema"],
    },
)
```

## Agent Tool Selection And Arguments

Agent evaluation should use execution evidence, not only final text. Declare expected actions in
the case:

```yaml
cases:
  - id: refund-order
    input:
      message: Refund order A-42
    expected:
      actions:
        - tool: lookup_order
          args:
            order_id: A-42
          order: 1
        - tool: issue_refund
          args:
            order_id: A-42
          order: 2
```

Then score the tool spans:

```yaml
score:
  tool_selection:
    expected_action:
      metric: selection
      span:
        kind: tool
    semantic: agent.tool.selection.correctness
    goal: maximize
    role: objective
  tool_arguments:
    expected_action:
      metric: arguments
      span:
        kind: tool
    semantic: agent.tool.argument.correctness
    goal: maximize
    role: objective
  tool_sequence:
    expected_action:
      metric: sequence
      span:
        kind: tool
    semantic: agent.tool.sequence.correctness
    goal: maximize
    role: constraint
```

This works with manually recorded tool spans and native SDK traces. It does not require an LLM judge
for deterministic action contracts.

## Custom SDK Without Application Changes

When an SDK is not built in, instrument a stable method and declare both evidence and assets:

```python
from autobench import (
    InstrumentAssetSpec,
    InstrumentMetricSpec,
    Semantic,
    SpanKind,
    instrument_method,
)

instrument_method(
    WorkflowClient,
    "execute",
    span="workflow.execute",
    span_kind=SpanKind.WORKFLOW,
    metrics=[
        InstrumentMetricSpec(
            name="confidence",
            semantic_type=Semantic.QUALITY_SCORE,
            value_path="result.confidence",
        ),
    ],
    assets=[
        InstrumentAssetSpec(
            kind="prompt",
            local_id="instructions",
            value_path="kwargs.instructions",
            name="routing_instructions",
        ),
        InstrumentAssetSpec(
            kind="tool",
            local_id="tools",
            value_path="kwargs.tools",
            many=True,
        ),
        InstrumentAssetSpec(
            kind="output_schema",
            local_id="output",
            value_path="kwargs.output_type",
            name="routing_output",
        ),
    ],
)
```

Serializable configurations use `value_path` or an import target. Typed Python integrations may use
`value_factory` for extraction that cannot be represented as a path. Keep domain computation in the
application; instrumentation should describe stable boundaries and evidence extraction.

## Production Failures As Regression Cases

Convert selected production samples into cases without coupling the benchmark to a production
database:

```python
from autobench import (
    ProductionSample,
    SampleReason,
    SamplingPolicy,
    samples_to_cases,
)

samples = [
    ProductionSample(
        id="trace-1842",
        input={"message": "Refund order A-42"},
        expected={"route": "billing"},
        reason=SampleReason.FAILURE_ONLY,
        privacy_tags=("customer_text",),
    )
]

cases = samples_to_cases(
    samples,
    policy=SamplingPolicy(
        reasons=(SampleReason.FAILURE_ONLY,),
        max_samples=100,
    ),
)
```

Review state, source reason, timestamp, trace identity, and privacy tags become metadata. Promote
reviewed cases into a versioned YAML dataset before using them as a release gate.

## Synthetic Case Generation With Provenance

Autobench does not prescribe a model-based generator, but it owns the typed preparation and
generated-data lineage boundary:

```python
from pathlib import Path

from autobench import (
    Case,
    CaseGeneratorInput,
    GeneratedCaseBatch,
    generate_dataset_sync,
    write_generation_result,
)

result = generate_dataset_sync(
    lambda request: GeneratedCaseBatch(
        cases=(Case(id="edge-1", input={"message": "..."}),),
        generator_asset_version="prompt.generate_cases@82ab39",
        model_provider="openrouter",
        model_name="openai/gpt-5.6-luna",
    ),
    CaseGeneratorInput(seed=17),
    generator_id="generation:generate_cases",
    dataset_id="routing-edge-cases",
)
write_generation_result(result, Path("datasets/routing-edge-cases.yaml"))
```

Candidate, accepted, and rejected states remain visible in the generation manifest. Generation is a
separate operation, so review and freezing happen before variants see the dataset. See
[Generated Datasets](generated-datasets.md).

## CI Regression Gate

A typical CI job validates, executes, stores artifacts, and checks policy state:

```bash
set -e
autobench validate benchmarks/release.yaml
autobench run benchmarks/release.yaml \
  --concurrency 4 \
  --record artifacts/autobench-release
autobench report artifacts/autobench-release
autobench export artifacts/autobench-release \
  --format csv \
  --path artifacts/autobench-runs.csv
```

Persist the whole record directory, not only the CSV. The CSV is a projection; the immutable YAML
records and artifacts contain replay, lineage, source, and diagnostic evidence.

## Optimization Handoff

Autobench marks metrics by role and direction:

- objective: improve this metric;
- constraint: do not violate this boundary;
- diagnostic: explain behavior without becoming an objective.

Factors can set `optimize: true`, and tracked assets identify the exact prompt/tool/schema versions
used. Convert a recorded run into compact feedback:

```python
from pathlib import Path

from autobench import build_optimization_feedback_input, load_run_record

record = load_run_record(
    Path("runs/latest/cases/refund-order/candidate/run.yaml"),
    root_dir=Path("runs/latest"),
)
feedback = build_optimization_feedback_input(record)
```

An optimizer should propose candidates and run controlled validation experiments. Autobench supplies
evidence and comparison; it does not claim that independently best assets can be mixed safely.

## Replay-Only Analysis

Recorded evidence supports analysis in an environment without the application or provider SDKs:

```python
from pathlib import Path

from autobench import build_report, replay_experiment

experiment = replay_experiment(Path("runs/latest"))
report = build_report(experiment)
```

This is the correct boundary for dashboards, offline reports, audits, post-hoc extraction, and
optimizer data ingestion.

To publish the same immutable evidence into an OTLP-compatible operations backend without rerunning
the subject:

```bash
autobench telemetry export runs/latest \
  --endpoint https://collector.example/v1/traces \
  --service-name routing-benchmark
```

The outbound adapter preserves experiment/run/trace identity and semantic events. It is not an
alternative record format; see [OTLP Export](otlp-export.md).

---

## Example Projects

Canonical page: https://vcoderun.github.io/autobench/examples/

# Example Projects

The repository examples use the public Autobench runtime. They are ordered by the amount of
framework surface they demonstrate, not by whether the subject is AI-based.

## Offline Release Matrix

These examples are credential-free and run in `make examples`:

| Example | Subject | Main features |
| --- | --- | --- |
| `minimal` | text transformation | inline cases, variants, exact score, matrix, comparison |
| `basic` | support routing | file dataset, spans, checks, artifacts, Rich reports |
| `mid` | response generation | semantic usage, pricing, cost, policies, distributions |
| `advanced` | search implementations | repeated samples, noise, paired speedup |
| `abp_manual` | ticket router | manual span plus method instrumentation |
| `abp_concurrent` | async workers | task-local trace context and concurrent runs |
| `automatic_assets` | Pydantic AI and custom SDK | automatic behavioral asset lineage |
| `generated_dataset` | support-routing case preparation | typed generator, request YAML, review state, frozen dataset and provenance manifest |
| `otlp_export` | immutable ABP record | offline OTLP hierarchy mapping through an injected exporter |
| `pydantic_gepa` | Optimize Anything Pipeline | optimizer lifecycle, engine branches, evaluation budgets, candidate lineage, and component asset versions |

Run all offline examples:

```bash
make examples
```

## Minimal: Learn The Matrix

```bash
autobench validate examples/minimal/autobench.yaml
autobench run examples/minimal/autobench.yaml --record /tmp/autobench-minimal
autobench replay /tmp/autobench-minimal
```

Read `examples/minimal/autobench.yaml` together with `minimal_benchmark.py`. This is the shortest
complete `case x variant -> task -> score -> record -> report` implementation.

## Generated Dataset: Prepare Before Planning

```bash
cd examples/generated_dataset
autobench dataset generate generator:generate_routing_cases \
  --request request.yaml \
  --output generated-cases.yaml \
  --id routing-generated \
  --version v1
```

This example is deterministic and credential-free. It writes a normal dataset and a separate
generation manifest, showing the boundary between data preparation and benchmark execution.

## Basic: Application Evidence

```bash
autobench run examples/basic/autobench.yaml --record /tmp/autobench-basic
autobench report /tmp/autobench-basic \
  --format markdown --layout bundle --output /tmp/autobench-basic-report
```

The task validates typed input, reads a factor, opens a workflow span, stores its output as an
artifact, and lets declarative scorers evaluate correctness and handling. The candidate fixes a
known routing failure, so the case matrix and comparison contain a visible behavioral delta.
Its YAML also publishes `reports/benchmark.md` inside the record before the manifest is sealed; the
second command demonstrates a post-hoc bundle generated from replayed evidence.

## Mid: Quality, Cost, And Constraints

```bash
autobench run examples/mid/autobench.yaml --record /tmp/autobench-mid
```

This example records input/output tokens and latency, resolves a local model pricing table, derives
`money.cost`, checks success and cost policies, and configures leaderboard and distribution views.
It is the best starting point for an LLM benchmark that already has a task implementation.

## Advanced: Measurement And Paired Baselines

```bash
autobench run examples/advanced/autobench.yaml --record /tmp/autobench-advanced
```

The task uses `measure_callable()` and `ctx.record_measurement()` instead of custom timing loops.
The post-deriver matches runs by case and computes candidate speedup against the baseline while
correctness remains a constraint.

## Pydantic AI: Live Layered Instrumentation

```bash
uv sync --extra instrumentation
export OPENROUTER_API_KEY=...
export OPENROUTER_MODEL=openrouter:openai/gpt-5.6-luna
uv run python examples/pydantic_ai/openrouter_instrument_all.py \
  --record /tmp/autobench-openrouter
```

The program makes a real OpenRouter request through Pydantic AI, uses a tool, streams structured
output, and calls `Benchmark.instrument_all()`. The task has no manual metrics, spans, or tracking
decorators. Autobench collects layered Pydantic AI, OpenAI client, and HTTPX evidence plus prompt,
tool, output-schema, and agent versions.

Inspect it afterward:

```bash
autobench instrumentation trace /tmp/autobench-openrouter
autobench replay /tmp/autobench-openrouter
```

`examples/pydantic_ai/agent_benchmark.py` is provider-neutral and accepts any configured Pydantic AI
model identifier through `PYDANTIC_AI_MODEL`.

## Pydantic-GEPA: Optimizer Evidence

```bash
uv run autobench run examples/pydantic_gepa/autobench.yaml \
  --record /tmp/autobench-pydantic-gepa
uv run autobench report /tmp/autobench-pydantic-gepa
```

The directory contains four credential-free benchmarks: standard GEPA, Optimize Anything Omni,
multi-component prompt/tool/output-schema optimization, and staged checkpoint/resume. Native
instrumentation records engine contenders, selection evidence, resource-specific budgets,
candidate lineage, effective component versions, and a replayable typed projection.

```bash
for name in standard autobench multi_component resume; do
  uv run autobench run "examples/pydantic_gepa/$name.yaml" \
    --record "/tmp/autobench-pydantic-gepa-$name"
done
```

An optional live example layers Pydantic AI, OpenAI-compatible OpenRouter, and HTTPX evidence under
the optimizer evaluation without duplicate accounting:

```bash
export OPENROUTER_API_KEY=...
uv run python examples/pydantic_gepa/live_pydantic_ai.py \
  --record /tmp/autobench-pydantic-gepa-live
```

It uses `openrouter:openai/gpt-5.6-luna` and is intentionally excluded from offline CI. See
[Pydantic-GEPA Instrumentation](pydantic-gepa-instrumentation.md).

## Automatic Asset Discovery

```bash
uv run python examples/automatic_assets/pydantic_ai_discovery.py \
  --record /tmp/autobench-pydantic-assets

uv run python examples/automatic_assets/custom_sdk_discovery.py \
  --record /tmp/autobench-custom-assets
```

Both are offline. The first uses a real Pydantic AI `Agent`, `AbstractCapability`, tool, and Pydantic
output model with no explicit tracking. The second adds prompt, tools, and output-schema extraction
to an arbitrary method with `InstrumentAssetSpec`.

## ABP Manual And Concurrent

```bash
autobench run examples/abp_manual/autobench.yaml --record /tmp/abp-manual
autobench run examples/abp_concurrent/autobench.yaml \
  --concurrency 2 \
  --record /tmp/abp-concurrent
```

Use the manual example to learn `RunContext.span()` and `instrument_method()`. Use the concurrent
example to inspect sibling span parentage and task-local context under async execution.

## OpenAI Streaming

```bash
uv sync --extra instrumentation
uv run python examples/abp_openai/run_openai_streaming.py
```

This uses the official OpenAI client and a real streaming parser over an offline HTTPX mock
transport. It demonstrates first-chunk and stream-completion evidence without network access.

## OpenAI Agents

```bash
uv sync --extra openai-agents
uv run python examples/abp_openai_agents/run_openai_agents.py
```

The example sends real OpenAI Agents workflow/function/custom trace events through the Autobench
trace processor. It requires no model request.

## Replay And Extraction

```bash
uv run python examples/abp_replay/replay_and_extract.py /tmp/recorded-experiment
```

The script loads records without provider SDKs and creates extraction-derived records with explicit
parent lineage.

## Offline OTLP Export

```bash
uv run autobench run examples/abp_manual/autobench.yaml --record /tmp/abp-manual
uv run python examples/otlp_export/export_record.py /tmp/abp-manual
```

The example maps a real recorded experiment to OTel SDK spans through an injected in-memory
exporter, so it verifies hierarchy and delivery without a collector or network request. Production
delivery uses `autobench telemetry export`; see [OTLP Export](otlp-export.md).

## CodeMode: Migrating A Real Benchmark Runner

```bash
export OPENROUTER_API_KEY=...
uv run python examples/codemode/run_benchmark.py --only parse_cron \
  --record /tmp/autobench-codemode
```

The CodeMode example replaces a bespoke benchmark script with cases, model-pair factors, a task,
semantic coverage/success/latency evidence, generated-spec artifacts, recording, and reports. Its
task still owns Vowel CodeMode calls; Autobench remains generic. The external CodeMode runtime and
network credentials are required.

## What To Copy

Copy the pattern, not generated run directories:

- task signature and typed input/output from `minimal` or `basic`;
- pricing and policies from `mid`;
- measurement and paired comparison from `advanced`;
- automatic SDK setup from `pydantic_ai`;
- optimizer evidence from `pydantic_gepa`;
- custom instrumentation from `automatic_assets`;
- replay processing from `abp_replay`.

For combinations not represented by one project, use [Use Cases](use-cases.md) and the
[Capability Map](capabilities.md).

---

## Architecture

Canonical page: https://vcoderun.github.io/autobench/architecture/

# Architecture

Autobench separates application execution from experiment infrastructure. This is the central
design constraint: the framework can benchmark any system because it does not own the system.

## Layered Model

| Layer | Owns | Does not own |
| --- | --- | --- |
| Definition | benchmark, dataset, variants, evaluation and report configuration | application implementation |
| Runtime | matrix planning, context, task invocation, concurrency, statuses | provider event loops or business orchestration |
| ABP | trace context, signals, spans, measurements, capture, source provenance | OpenTelemetry or hosted trace storage |
| Evaluation | scorers, derivation, policies, paired comparisons | domain truth that only the application can supply |
| Tracking | behavioral asset identity, versions, representations, diffs, uses | source control or deployment promotion |
| Records | immutable run and experiment evidence, artifacts, source hashes | mutable operational databases |
| Reports | semantic projections, aggregation, comparison and exports | causal inference from uncontrolled changes |
| Outbound adapters | projections such as ABP-to-OTLP delivery | canonical evidence or application instrumentation |

## One Canonical Spec

YAML and the Python builder converge on `BenchmarkSpec`:

```text
YAML DSL -----------+
                    +--> BenchmarkSpec --> BenchmarkPlan --> ExperimentResult
Benchmark builder --+
```

The builder is ergonomic composition; it is not a second runtime. `Benchmark.to_spec()` returns the
same model loaded by `load_benchmark_spec()`.

## Execution Lifecycle

For each case x variant pair, Autobench:

1. creates a stable run ID and `RunContext`;
2. activates task-local ABP context;
3. invokes `task(ctx, case)` synchronously or asynchronously;
4. preserves evidence even when the task fails;
5. evaluates built-in and Python scorers;
6. projects scores into semantic observations;
7. derives per-run metrics such as token cost;
8. finalizes status and trace state.

The context tracks these phases explicitly. With durable recording, `await ctx.checkpoint(name)`
commits a frozen partial snapshot through the active record session. Cooperative cancellation
finalizes partial trace state, commits a reserved terminal checkpoint, and propagates the same
`CancelledError`; concurrent cancellation drains sibling cleanup within a fixed bound before the
experiment session aborts.

As runs finish, an optional recorder commits their execution snapshots. After all runs finish, the
runtime applies cross-run derivation and policy evaluation, then
materializes immutable YAML records and referenced artifacts.

Hard termination is intentionally weaker: `SIGKILL` and power loss preserve only the last staging
manifest commit under the selected durability mode. Autobench checkpoints evidence, while the
subject application remains responsible for resumable execution state.

## Evidence Model

`Observation` is the common query and aggregation unit. Its local `name` explains the metric in the
application; `semantic_type` explains what the value means across applications. Source and
provenance distinguish task observations, scores, derived values, and trace extraction.

ABP preserves richer execution evidence as an ordered signal stream and a materialized `Trace`.
Useful trace values can be extracted into observations without discarding their span or source-map
lineage.

```text
SDK call
  -> native instrumentor
  -> ABP signals
  -> canonical Trace
  -> semantic extraction
  -> Observation
  -> report / policy / optimizer
```

## Definition And Effective Assets

Behavioral components have two useful representations:

- **definition**: what application code configured, such as a prompt template or Python tool;
- **effective**: what an SDK sent to a model or downstream system after normalization.

Automatic asset discovery can record both and link their versions. This makes lineage explain not
only that a tool changed, but also how its model-facing schema changed.

## Immutability And Replay

A `RunRecord` is evidence, not a cache entry. Replay never mutates it and never silently executes
the task. Rescoring, recanonicalization, or trace extraction creates new derived records with parent
lineage.

This supports three distinct workflows:

- **report replay**: render new views over unchanged evidence;
- **evidence replay**: run a versioned extractor or canonicalizer over stored ABP data;
- **execution rerun**: intentionally execute a new experiment against the current application.

## Extension Seams

Choose the narrowest seam that owns the behavior:

| Need | Extension |
| --- | --- |
| Call application code | Python task |
| Evaluate domain output | Python scorer |
| Compute from same-run observations | Deriver |
| Compare matched runs | Post-deriver |
| Enforce an acceptance rule | Policy |
| Collect a stable SDK boundary | Instrumentor |
| Map vendor fields to semantics | Source map / extractor |
| Add domain defaults | Metric pack |
| Version a behavioral component | Tracking or asset discovery |

Application-specific logic belongs in tasks and scorers. Generic SDK behavior belongs in an
instrumentor. This prevents core Autobench from accumulating one-off integrations disguised as
framework concepts.

An external framework may retain ownership of its telemetry backend. Its adapter can use
`InstrumentationRuntime.span()` to join the active ABP tree, while the external package owns backend
multiplexing and conditional restoration. Autobench core therefore supplies context and evidence
semantics without importing the framework or replacing an existing telemetry destination.

After recording, the optional OTLP adapter can map immutable experiment, run, trace, and span
evidence to an external telemetry backend. This happens outside benchmark execution and never
turns OTLP into storage, replay lineage, or canonical semantics.

## Optimization Boundary

Autobench produces optimization-grade evidence: objectives, constraints, diagnostics, factors,
asset versions, candidate feedback, and replayable run lineage. It deliberately does not select
mutation strategies or promote candidates. Consumers such as pydantic-gepa and autoptimize can use
the records without Autobench becoming coupled to one optimizer.

---

## Core Concepts

Canonical page: https://vcoderun.github.io/autobench/concepts/

# Core Concepts

Autobench models a benchmark as a deterministic experiment over cases and variants. The concepts
below appear in both the YAML DSL and Python API.

## BenchmarkSpec

The canonical definition of one benchmark. It contains metadata, capture policy, dataset, task,
variants, scoring, derivation, policies, instrumentation, report configuration, and a semantic
registry.

## Case And Dataset

A `Case` is one input and its optional expectation:

```python
from autobench import Case

case = Case(
    id="refund-request",
    input={"message": "Refund order 42"},
    expected={"route": "billing"},
    tags=["regression", "routing"],
    metadata={"language": "en"},
)
```

A `DatasetSpec` adds identity, version, defaults, source provenance, and attachments to a case
collection.

## Variant And Factor

A variant is one concrete configuration of the subject. Factors are independent variables such as
model, prompt version, implementation strategy, or feature flag.

```text
case: refund-request
variant: candidate
factors: model=gpt-x, prompt=refund-v4, temperature=0
```

Autobench records factors and their semantics. It does not assume that changing several factors at
once proves which one caused a metric delta.

## Task

The task adapts a case and variant to the system being benchmarked:

```python
def run(ctx, case):
    model = ctx.factor("model")
    return application.execute(case.input, model=model)
```

The task owns application calls. Autobench owns invocation, timing context, evidence preservation,
and status classification.

## Observation

An `Observation` is a typed fact produced during a run. It has a kind, name, value, optional
semantic type and unit, source, role, direction, and provenance.

Kinds include metrics, factors, events, diagnostics, outcomes, checks, and artifacts. Sources let
projection distinguish a task-emitted metric from a scorer or derived value with the same semantic
type.

## Semantic Type

A semantic type is a stable string such as `quality.correctness`, `llm.tokens.input`, or
`time.latency`. It lets generic components consume meaning rather than application-local names.

The registry carries aliases, parent relationships, aggregation hints, cardinality, privacy, and
stability metadata. Applications may extend it without replacing built-ins.

## Score

A `ScoreRecord` is evaluator output. Scores can be objectives, constraints, or diagnostics and can
include reasons, errors, and selected span provenance. They are projected into observations with
score precedence for reporting and policies.

## Derivation

A deriver computes a metric from evidence:

- per-run derivation uses one run, such as tokens + model pricing -> cost;
- post-derivation uses the experiment, such as matched baseline/candidate latency -> speedup.

Derived observations preserve their input references and source.

## Policy

A policy is a pass/fail requirement over semantic metrics. Operators are explicit fields such as
`must_equal`, `must_greater_equal`, `must_less_equal`, and `must_be_between`. Policies affect
evaluation status without hiding the underlying metric.

## ABP Trace

The Autobench Protocol (ABP) is the native execution evidence model. Instrumentors and manual spans
emit immutable signals that materialize into a trace containing spans, measurements, events, links,
references, errors, stream state, and diagnostics.

ABP is not an OpenTelemetry wrapper. The optional outbound OTLP adapter can project recorded ABP
evidence into OTel spans, but Autobench controls its semantic, replay, accounting, and optimization
contracts.

## Tracked Asset

A tracked asset is a behavioral component whose exact version matters to a run: prompt, tool,
output schema, type, capability, agent, guardrail, handoff, policy, toolset, or arbitrary config.

Assets have stable logical IDs and content-addressed versions. Their history records normalized
state, source hashes, parent versions, changed fields, and diffs. An `AssetUse` binds the version and
representation actually used to a run and optional span.

## RunRecord And ExperimentRecord

`RunRecord` is the immutable evidence for one case x variant execution. `ExperimentRecord` describes
the plan, source hashes, environment, report configuration, run paths, and aggregate statuses for
the whole matrix.

Both complete and partial evidence can be immutable. A run records `partial` and `end_reason`; an
experiment records one terminal state, post-processing completeness, and its planned, recorded, and
missing run identities. This lifecycle metadata is separate from benchmark outcomes: an experiment
can complete even when individual runs fail.

Records are human-readable YAML views backed by strict typed models and versioned JSON Schemas. A
content manifest is validated before replay, and final experiment directories are published only
after every record and summary file exists.

## Report

A report is a replay-time projection, not the source of truth. Leaderboards, case matrices,
comparisons, distributions, and run tables all derive from RunRecords. New report configuration can
therefore analyze existing evidence without running the subject.

## Optimization Feedback

Feedback records compact failed scores, policy violations, task and span errors, factors, and asset
versions. They give optimizers structured evidence without forcing them to scrape terminal output or
infer semantics from metric names.

---

## Datasets And Variants

Canonical page: https://vcoderun.github.io/autobench/datasets-and-variants/

# Datasets And Variants

Autobench expands a dataset against variants to create a deterministic run matrix. Cases describe
what is evaluated; variants describe what changes.

## Cases

Each `Case` has a stable ID and may carry arbitrary input, expected output, metadata, tags, and
attachments:

```python
from autobench import Case

case = Case(
    id="refund-request",
    input={"message": "I need a refund for order 42"},
    expected={"route": "billing", "priority": "normal"},
    metadata={"tenant": "demo"},
    tags=["routing", "smoke"],
)
```

Inputs and expected values are intentionally generic. They can be strings, mappings, Pydantic
models serialized by the task, structured multimodal references, or domain-specific payloads.
Attachments use `ArtifactRef` values when a case depends on external material.

## Dataset Sources

Cases may be authored inline:

```yaml
dataset:
  version: v1
  cases:
    - id: refund-request
      input:
        message: I need a refund
      expected:
        route: billing
```

Or loaded relative to the benchmark file:

```yaml
dataset:
  source: file://datasets/cases.yaml
  version: v1
```

File-backed datasets use the same DSL representation. Glob-backed sources can combine separate
case files while duplicate case IDs remain validation errors.

## Case Defaults

Defaults reduce repeated metadata without hiding the final case payload:

```yaml
dataset:
  defaults:
    metadata:
      locale: en-US
    tags: [regression]
  cases:
    - id: ticket-1
      input: {message: Reset my password}
      tags: [authentication]
```

Mapping values are merged, tags are deduplicated, and explicit scalar case values override
defaults.

## Variants And Factors

A variant is one concrete factor set:

```yaml
variants:
  baseline:
    label: Current production route
    factors:
      model:
        value: openrouter:openai/gpt-5.6-luna
        semantic: llm.model.name
        optimize: true
      prompt_version:
        value: route-v3
        semantic: prompt.version
        optimize: true
      temperature: 0
```

`value` is the runtime value. `semantic` tells downstream consumers what the factor means.
`optimize` is a hint that the factor is a candidate optimization axis; Autobench records it but
does not choose search strategies.

The Python form is equivalent:

```python
from autobench import FactorValue, Semantic, Variant

variant = Variant(
    id="baseline",
    label="Current production route",
    factors=[
        FactorValue(
            name="model",
            value="openrouter:openai/gpt-5.6-luna",
            semantic_type=Semantic.LLM_MODEL_NAME,
            optimize=True,
        ),
        FactorValue(name="temperature", value=0),
    ],
)
```

Tasks read factors through `ctx.factor(name)`. Factors are also copied into RunRecords and report
variant-configuration tables.

## Generated And Production Cases

The data layer preserves where generated examples came from:

- `ProductionSample` models a source sample and review state.
- `sample_to_case` and `samples_to_cases` convert samples without losing provenance.
- `mark_generated_case` records generation metadata.
- `generated_batch_from_cases` creates a `GeneratedCaseBatch` with generator, model, and source
  details;
- `generate_dataset` and `generate_dataset_sync` execute a typed generator before benchmark
  planning;
- `write_generation_result` publishes a normal dataset plus a separate provenance manifest.

Autobench owns the preparation and evidence contract, not the model or sampling strategy. Complete
generation freezes accepted and candidate cases into ordinary dataset YAML; rejected cases remain
in the manifest with their reasons. Incomplete generation writes only an incomplete sidecar and
cannot replace a benchmark dataset. See [Generated Datasets](generated-datasets.md).

## Identity And Reproducibility

- Case IDs and variant IDs must be unique.
- Dataset content hashes depend on normalized content rather than filesystem location.
- Matrix order is deterministic.
- Run IDs are stable for a given plan position, case, and variant.
- Dataset version may also be emitted as `dataset.version` semantic evidence.

---

## Generated Datasets

Canonical page: https://vcoderun.github.io/autobench/generated-datasets/

# Generated Datasets

Autobench can prepare generated cases before a benchmark begins, preserve how they were produced,
and publish them as an ordinary reviewable dataset. Generation is deliberately separate from the
case x variant execution loop: every variant sees the same frozen cases.

## Lifecycle

```text
generation request
  -> sync or async CaseGenerator
  -> GeneratedCaseBatch
  -> review/provenance normalization
  -> complete dataset YAML + generation manifest
  -> normal benchmark dataset.source
```

This separation prevents a model, random seed, provider retry, or partially completed generation
job from changing matrix identity while variants are already running.

## Generator Contract

A generator receives one `CaseGeneratorInput` and returns `GeneratedCaseBatch`. It may be
synchronous or asynchronous:

```python
from autobench import (
    Case,
    CaseGeneratorInput,
    GeneratedCaseBatch,
    GeneratedCaseReview,
    GenerationCost,
    GenerationDeterminism,
    GenerationUsage,
    ReviewStatus,
)


def generate_cases(request: CaseGeneratorInput) -> GeneratedCaseBatch:
    route = str(request.settings.get("route", "billing"))
    cases = (
        Case(
            id="generated-refund",
            input={"message": "Refund a duplicate charge"},
            expected={"route": route},
        ),
    )
    return GeneratedCaseBatch(
        cases=cases,
        generator_asset_version=request.prompt_asset_version,
        model_provider="openrouter",
        model_name="openai/gpt-5.6-luna",
        determinism=GenerationDeterminism.NOT_GUARANTEED,
        usage=GenerationUsage(input_tokens=120, output_tokens=45, requests=1),
        cost=GenerationCost(amount=0.0012, currency="usd"),
        reviews=(
            GeneratedCaseReview(
                case_id="generated-refund",
                status=ReviewStatus.ACCEPTED,
            ),
        ),
    )
```

Autobench does not prescribe the model, sampler, prompt framework, or review service. The callable
is the application-owned adapter; its typed output is the portable evidence boundary.

## Request YAML

The CLI can pass a human-authored request to the generator:

```yaml
# yaml-language-server: $schema=schemas/0.3.0/generation_request_schema.json
generation:
  request:
    seed: 17
    prompt:
      content: Generate privacy-safe support routing cases.
      asset_version: prompt.routing-generator@v1
    settings:
      route: billing
      count: 20
    metadata:
      owner: evaluation
    seed_cases:
      - id: reviewed-refund
        input:
          message: Refund for a duplicate charge
        expected:
          route: billing
```

`settings` and `metadata` accept portable serialized values. `seed_cases` provide reviewed examples
without requiring the generator to load a Pydantic Evals dataset or an Autobench benchmark spec.

## Generate From The CLI

```bash
autobench dataset generate generator:generate_cases \
  --request generation-request.yaml \
  --output datasets/generated-routing.yaml \
  --id generated-routing \
  --version v1
```

A complete invocation writes:

- `datasets/generated-routing.yaml`: normal dataset DSL consumed by `dataset.source`;
- `datasets/generated-routing.generation.yaml`: provider, model, prompt version, seed, settings,
  usage, cost, review state, timestamps, and content hashes.

Existing outputs are protected. Use `--force` only when intentionally replacing both artifacts.

## Generate From Python

```python
from pathlib import Path

from autobench import (
    CaseGeneratorInput,
    generate_dataset_sync,
    write_generation_result,
)

request = CaseGeneratorInput(
    seed=17,
    prompt="Generate privacy-safe support routing cases.",
    prompt_asset_version="prompt.routing-generator@v1",
    settings={"route": "billing", "count": 20},
)
result = generate_dataset_sync(
    generate_cases,
    request,
    generator_id="generator:generate_cases",
    dataset_id="generated-routing",
    version="v1",
)
written = write_generation_result(result, Path("datasets/generated-routing.yaml"))
```

Use `await generate_dataset(...)` inside an async application. `GenerationResult` validates the
request hash, case records, review projection, dataset content, timestamps, and dataset hash before
publication.

## Review Semantics

Every generated case has one of three states:

| State | Published in complete dataset? | Meaning |
| --- | --- | --- |
| `candidate` | Yes | Generated and not yet explicitly accepted or rejected |
| `accepted` | Yes | Explicitly reviewed and accepted |
| `rejected` | No | Excluded; a nonempty rejection reason is required |

The generation manifest retains all three states, including the complete rejected case and its
reason. The dataset retains `source`, review status, generator/model provenance, and its own content
hash in each included case's metadata. Teams that permit only accepted data can review or filter
candidates in their generator before returning a complete batch.

## Incomplete Generation

A generator that reaches a budget, provider, or review boundary can return:

```python
return GeneratedCaseBatch(
    complete=False,
    incomplete_reason="manual review required",
    cases=partial_cases,
)
```

Autobench writes only `<output-stem>.incomplete.yaml` and exits the CLI with status `2`. It does not
write or replace the requested dataset, even with `--force`. The sidecar preserves partial cases for
inspection without presenting them as benchmark-ready truth. An exception is a failed generation
operation, not a partial result; return an explicit incomplete batch when partial evidence exists.

## Determinism And Hashes

`GenerationDeterminism` is an evidence claim made by the generator adapter:

- `guaranteed`: the provider and adapter contract guarantee the same generated content;
- `not_guaranteed`: repeated requests may differ;
- `unknown`: no reliable guarantee is available.

Autobench never upgrades this claim based only on a seed. The request hash covers seed cases,
prompt, prompt version, settings, and metadata. Each generated case and the frozen dataset have
separate SHA-256 identities. For a guaranteed generator, identical normalized output produces
byte-identical dataset YAML; the provenance manifest still has new execution timestamps.

The manifest stores the prompt hash and asset version rather than copying prompt content. Keep the
request file or tracked prompt asset when the exact historical text must be reconstructed.

## Use In A Benchmark

After generation, consume the output like any other file-backed dataset:

```yaml
benchmark:
  routing:
    dataset:
      source: file://datasets/generated-routing.yaml
      version: v1
    run:
      python: benchmark_task:run
    variants:
      current: {}
      candidate: {}
```

Generation never runs implicitly during `autobench run`. Regenerate deliberately, review the diff,
then start a new experiment so every run has one immutable dataset identity.

---

## Tasks And Runtime

Canonical page: https://vcoderun.github.io/autobench/tasks-and-runtime/

# Tasks And Runtime

The task is the only application-specific execution boundary required by Autobench. It receives a
runtime context and a case, invokes the subject, records evidence, and returns the output that
scorers evaluate.

## Task Contract

```python
from autobench import Case, RunContext


def run_case(ctx: RunContext, case: Case) -> dict[str, object]:
    model = ctx.factor("model")
    result = call_application(case.input, model=model)
    ctx.outcome(result.ok)
    return {"answer": result.answer, "ok": result.ok}
```

The positional contract is always `task(ctx, case)`. Tasks may be synchronous or asynchronous:

```python
async def run_case(ctx: RunContext, case: Case) -> dict[str, object]:
    result = await call_application(case.input)
    return {"answer": result.answer}
```

YAML resolves the callable relative to the benchmark file before falling back to import paths:

```yaml
run:
  python: benchmark_tasks:run_case
```

## RunContext

`RunContext` owns evidence for one case x variant run:

| Method | Use |
| --- | --- |
| `factor(name)` | Read a variant factor |
| `span(...)` | Open a timed nested operation |
| `metric(...)` / `metrics(...)` | Record one or many metrics |
| `factor_observation(...)` | Record a runtime-discovered factor |
| `event(...)` | Record a discrete event |
| `diagnostic(...)` | Record non-objective diagnostic evidence |
| `outcome(...)` | Record semantic run success |
| `check(...)` | Record a boolean correctness check with an optional reason |
| `record_measurement(...)` | Record summary statistics and optional sample artifact |
| `artifact(...)` | Attach a structured or file-like payload |
| `error(...)` | Attach a structured error without losing collected evidence |
| `attach_tracked_asset(...)` | Bind a tracked asset version to the run |

Context evidence remains available even when the task raises. The runtime captures the exception,
preserves observations and artifacts already emitted, and records a structured error.

## Matrix Execution

`build_benchmark_plan` validates and counts the matrix before execution. `expand_matrix` produces
one `MatrixRunSpec` per case x variant pair. The CLI renders the same plan during validation.

```bash
autobench validate autobench.yaml
autobench run autobench.yaml --concurrency 4 --record runs/latest
```

Concurrency bounds the number of active runs. Result ordering stays deterministic even when task
completion order differs.

## Failure And Status Model

Autobench separates three status layers:

- `TaskStatus`: whether application execution completed, failed, or was skipped.
- `EvaluationStatus`: whether scoring and constraints completed.
- `RunStatus`: final passed, failed, errored, or skipped state.

This distinction prevents a policy failure from looking like an application exception and lets
reports separate execution reliability from evaluation quality.

## Progress Events

`ProgressEvent` is a live, typed observer surface for terminals, service runners, and UIs. Pass one
or more synchronous or asynchronous handlers to any execution entry point:

```python
from autobench import ProgressEvent, ProgressEventKind, run_benchmark_spec


async def publish(event: ProgressEvent) -> None:
    if event.kind is ProgressEventKind.RUN_FINISHED:
        await send_status(event.run_id, event.run_status, event.sequence)


result = await run_benchmark_spec(spec, progress_handlers=(publish,))
```

Runtime delivery has these guarantees:

- `benchmark_started` is emitted once after Autobench owns execution.
- Each emitted `run_started` receives exactly one `run_finished` on cooperative success, failure,
  error, skip, or cancellation.
- `run_finished.run_status` is the final status after cross-run derivation and policies.
- Failed policies emit `policy_violation`; passing policies do not create noise.
- `benchmark_finished.experiment_status` is `completed`, `cancelled`, or `aborted`.
- `sequence` is unique and monotonic for one benchmark execution. Concurrent producers are
  serialized by the dispatcher; final run events follow logical matrix order.

Handlers run in registration order. A synchronous handler runs inline. An asynchronous handler is
awaited before the next handler or event, so a slow handler applies deliberate backpressure to the
benchmark.

Library execution is strict by default. A handler that raises is disabled, remaining handlers still
receive terminal events, durable recording is finalized, and Autobench then raises
`ProgressDispatchError`. CLI progress explicitly uses `ProgressErrorPolicy.BEST_EFFORT` and reports
renderer failures to stderr:

```python
from autobench import ProgressErrorPolicy, run_benchmark_spec

result = await run_benchmark_spec(
    spec,
    progress_handlers=(optional_dashboard,),
    progress_error_policy=ProgressErrorPolicy.BEST_EFFORT,
    progress_error_handler=report_dashboard_failure,
)
```

Progress is not persistence. The recorder stages evidence through its own lifecycle even when a
progress handler fails. A hard process death cannot emit terminal events; inspect durable staging
to recover the last committed evidence.

Cooperative cancellation cannot interrupt a completed run halfway through its durable stage.
Recorder operations are shielded and session-owned; cancellation is propagated only after the
commit reaches a terminal state or is retained as explicitly active cleanup. Abort and close run
after outstanding stages, checkpoints, artifact transfers, and final publication work, never in
parallel with them. Cleanup callbacks that ignore cancellation remain tracked and their eventual
exceptions are reported through the event loop instead of becoming unobserved task warnings.

## Python Builder

The builder compiles to the same `BenchmarkSpec` used by YAML:

```python
from autobench import Benchmark, Case, FactorValue, PassFailScorer, Semantic, Variant

result = (
    Benchmark("routing")
    .dataset([Case(id="refund", input={"message": "Refund order 42"})])
    .variants(
        [
            Variant(
                id="baseline",
                factors=[FactorValue(name="route", value="v1")],
            )
        ]
    )
    .task("benchmark_tasks:run_case")
    .scoring(
        [
            PassFailScorer(
                name="success",
                path="output.ok",
                semantic_type=Semantic.RESULT_SUCCESS,
            )
        ]
    )
    .run()
)
```

Use YAML for portable benchmark definitions and the builder when a Python application needs to
compose specs programmatically. Both execute through the same planner and runtime.

---

## Observations And Semantics

Canonical page: https://vcoderun.github.io/autobench/observations-and-semantics/

# Observations And Semantics

An observation is Autobench's atomic evidence unit. Raw names remain useful to humans, while
semantic types make evidence portable across applications, reports, and optimizers.

## Observation Model

An `Observation` carries:

- stable ID and local name
- kind: metric, factor, event, diagnostic, or artifact
- value and optional unit
- semantic type
- optimization direction and role
- source and optional span ID
- tags, case ID, and variant ID

```python
from autobench import Direction, ObservationRole, Semantic

ctx.metric(
    "answer_accuracy",
    0.94,
    semantic_type=Semantic.QUALITY_CORRECTNESS,
    direction=Direction.MAXIMIZE,
    role=ObservationRole.OBJECTIVE,
)
```

The local name may be `answer_accuracy`, `judge_score`, or `coverage`; the semantic type tells the
framework whether those values share meaning.

## Built-In Semantic Families

| Family | Examples |
| --- | --- |
| LLM | `llm.tokens.input`, `llm.tokens.output`, `llm.request.count`, `llm.model.requested`, `llm.model.response`, `llm.provider.name` |
| Cost | `money.cost`, `serving.cost`, `optimization.cost`, `lifetime.cost` |
| Time | `time.latency`, `time.first_chunk`, `time.critical_path` |
| Result | `result.success` |
| Quality | `quality.score`, `quality.correctness`, `coverage.ratio` |
| Agent | task completion, plan quality/adherence, step efficiency, tool selection/arguments/sequence, output correctness |
| Assets | `prompt.version`, `agent.tool.version`, `agent.version`, `dataset.version` |
| Operations | count, maximum depth/fan-out, incomplete work, parallelism, retries, recovered retries, first-attempt success |
| Workflow | validation failures, approval count/wait, tool-call success/failure, message growth, evidence-reference counts |

`Semantic` exposes completion-friendly constants. `SemanticType` remains extensible so domain
metrics can use names such as `retrieval.recall` or `business.conversion`.

## Registry And Aliases

`SemanticRegistry` stores definitions, aliases, and parent relationships. A custom registry can be
embedded in a benchmark spec and is merged with built-ins:

```yaml
semantic_registry:
  version: 1
  types:
    business.conversion:
      description: Whether the workflow produced a qualified conversion.
      parent: result.success
      unit: boolean
  aliases:
    conversion: business.conversion
```

Parent relationships let a query request a broad semantic category while preserving specific
metrics. Aliases prevent local naming differences from fragmenting evidence.

## Roles And Directions

Roles describe how a metric participates in evaluation:

- objective: something to optimize
- constraint: something that must remain acceptable
- diagnostic: explanatory evidence

Directions are `maximize` or `minimize`. Factors, events, and artifacts cannot declare an
optimization direction because they are not outcomes.

## Sources And Projection

The same semantic metric can be emitted by a task, scorer, deriver, policy, or adapter. Raw
observations are never discarded. Projection chooses a canonical value using explicit source
priority and ABP accounting scope. A derived aggregate summary is preferred to same-source direct
measurements for single-value reporting, while direct observations remain queryable. Logical
operation IDs correlate equivalent framework/client evidence; equal-priority disagreements are
marked ambiguous instead of silently picking one.

Use `ObservationQuery` for raw or projected lookup and `filter_observations` for selectors such as
semantic type, role, source, or span.

```python
from autobench import ObservationQuery

query = ObservationQuery(observations=list(result.observations))
costs = query.values("money.cost", projected=False)
```

Reports, policies, and derivation use this semantic projection layer rather than relying on local
metric names.

## Metric Packs

A `MetricPack` bundles reusable semantic defaults without forcing every metric into core:

- semantic registry additions
- scorer factory references
- default report metrics
- feedback extractors

Built-in packs cover `agentic`, `structured_output`, `llm_usage`, and `performance`. Applications
can register their own packs through `MetricPackRegistry` while keeping the RunRecord contract
unchanged.

---

## Scoring And Derivation

Canonical page: https://vcoderun.github.io/autobench/scoring-and-derivation/

# Scoring And Derivation

Scorers evaluate one run. Derivers compute new metrics from collected evidence. Post-derivers work
across runs after the complete experiment exists. Policies turn semantic metrics into explicit
requirements.

## Scoring Contract

Every scorer declares:

- a local score name
- semantic type
- optional unit
- optimization direction
- role: objective, constraint, or diagnostic
- whether scorer failure is optional

Scores are stored as `ScoreRecord` values and projected into observations with score-source
precedence. The original task observations remain available.

## Output Metric

Project an output value directly:

```yaml
score:
  coverage:
    value: output.coverage
    semantic: coverage.ratio
    goal: maximize
    role: objective
```

Use this when the task already computes a trustworthy metric.

## Pass/Fail

```yaml
score:
  success:
    pass: output.ok
    semantic: result.success
    role: constraint
```

The path must resolve to a boolean-like success value.

## Exact Match

```yaml
score:
  route_correctness:
    exact:
      actual: output.queue
      expected: case.expected.queue
    semantic: quality.correctness
    goal: maximize
```

Paths can address `output`, `case.input`, `case.expected`, factors, and structured values.

## Schema Validation

`SchemaScorer` validates a selected output path against a JSON Schema mapping. It is appropriate
for contracts where structural validity is separate from domain correctness.

```python
from autobench import SchemaScorer, Semantic

scorer = SchemaScorer(
    name="output_schema",
    path="output",
    schema={
        "type": "object",
        "required": ["customer_name", "id"],
        "properties": {
            "customer_name": {"type": "string"},
            "id": {"type": "string"},
        },
    },
    semantic_type=Semantic.AGENT_OUTPUT_STRUCTURE_VALIDITY,
)
```

## Python Scorers

Custom scorers receive `ScoringCall`, not loose callback dictionaries:

```python
from autobench import ScoreRecord, ScoringCall


def field_accuracy(call: ScoringCall) -> ScoreRecord:
    expected = call.case.expected
    output = call.output
    fields = ("name", "id", "pocket_id")
    matches = sum(output[field] == expected[field] for field in fields)
    return ScoreRecord(
        name="field_accuracy",
        semantic_type="quality.field_accuracy",
        value=matches / len(fields),
    )
```

`ScoringCall` exposes the case, variant, task output/result, observations, spans, and selected spans.
Python scorers may be sync or async. Optional scorers record errors without failing the run.

## Expected Actions

`ExpectedActionScorer` deterministically evaluates action/tool selection, arguments, or ordered
sequence from spans. See [Agentic Evaluation](agentic-evaluation.md).

## Dotted Paths

`resolve_dotted_path` is the shared structured-path resolver used by built-in scorers. Missing
paths produce explicit scorer errors instead of silently returning `None`.

## Per-Run Derivation

`derive` runs after task observations and scores are available for one run. `TokenCostDeriver` is
the built-in per-run deriver.

```yaml
derive:
  - kind: token_cost
    pricing: file://pricing/models.yaml
    output:
      name: request_cost
      semantic_type: money.cost
      unit: usd
      direction: minimize
      role: constraint
```

By default it reads:

- `llm.tokens.input`
- `llm.tokens.output`
- `llm.provider`
- `llm.model.name`

Input semantics and output metadata can be overridden through `TokenCostInputs` and
`DerivedMetricOutput` in the Python API.

Unknown models, missing usage, missing rates, and ambiguous inputs produce diagnostics; Autobench
does not invent a zero cost.

## Pricing DSL

Pricing is normalized into a `PricingTable` keyed by stable model IDs. Provider-specific aliases
allow input forms such as `provider:model`, `provider/model`, or application-specific model slugs
to resolve to the same entry.

```yaml
pricing:
  version: 1
  provider: openai
  models:
    openai/gpt-demo:
      aliases: [openai:gpt-demo, gpt-demo]
      input:
        unit: mtok
        price_per_million_tokens: 1.0
      output:
        unit: mtok
        tiers:
          - up_to_tokens: 100000
            price_per_million_tokens: 4.0
          - price_per_million_tokens: 6.0
      cache_read:
        unit: mtok
        price_per_million_tokens: 0.1
```

Supported fields include input, output, cache-read, and cache-write prices plus token-count tiers.
`StaticPriceSource`, `LLMPricesSource`, and `GenAIPricesSource` only import external price data into
this model. They do not make an external catalog authoritative at runtime.

## Paired Baseline Post-Derivation

`post_derive` has access to the full experiment:

```yaml
post_derive:
  - kind: paired_baseline
    baseline_variant: baseline
    match_on:
      - kind: case_id
      - kind: factor
        name: workload.size
    metric: time.latency
    formula: baseline_over_candidate
    include_baseline: true
    output:
      name: speedup
      semantic_type: performance.speedup
      unit: ratio
      direction: maximize
      role: objective
```

Formulas:

- `baseline_over_candidate`
- `candidate_over_baseline`
- `candidate_minus_baseline`
- `baseline_minus_candidate`
- `percent_change_from_baseline`

Matching supports case IDs and factor keys. Missing matches, nonnumeric metrics, absent metrics, and
zero division can be skipped or recorded as diagnostics.

Relative-noise thresholds and `ComparisonVerdictSpec` can emit improved, regressed, unchanged, or
inconclusive verdicts. These are controlled comparisons, not automatic causal claims.

## Policies

Policies evaluate projected semantic values and append `PolicyResult` evidence:

```yaml
policies:
  - name: request-must-succeed
    metric: result.success
    must_equal: true
  - name: cost-cap
    metric: money.cost
    must_less_equal: 0.001
  - name: acceptable-latency
    metric: time.latency
    must_between:
      min: 0
      max: 500
      inclusive: true
```

Each policy declares exactly one requirement:

- `must_equal` / `must_not_equal`
- `must_greater` / `must_greater_equal`
- `must_less` / `must_less_equal`
- `must_in` / `must_not_in`
- `must_between`

A failed constraint can change final run status while preserving the successful task output and
all evidence that explains the decision.

## Repeated Measurement

`measure_callable` avoids repeating warmup and sampling loops in benchmark tasks:

```python
from autobench import MeasurementBudget, measure_callable

measurement = measure_callable(
    lambda: search(case.input["items"], case.input["query"]),
    budget=MeasurementBudget(warmup=3, repetitions=20, max_seconds=2.0),
)
ctx.record_measurement("search", measurement)
```

`Measurement` includes samples, count, min, max, mean, median, p95, standard deviation, and relative
noise. A custom timer can measure accelerators or remote systems without adding domain-specific
logic to Autobench.

---

## Agentic Evaluation

Canonical page: https://vcoderun.github.io/autobench/agentic-evaluation/

# Agentic Evaluation

Autobench evaluates agents as traced systems rather than treating only the final text as evidence.
The same primitives also work for workflow engines, retrievers, and tool-using applications.

## Record Agent Behavior

```python
from autobench import Semantic, SpanKind


def run_case(ctx, case):
    with ctx.span("support_agent", kind=SpanKind.AGENT, input=case.input) as agent:
        with ctx.span(
            "lookup_user",
            kind=SpanKind.TOOL,
            input={"user_id": case.input["user_id"]},
        ) as tool:
            profile = lookup_user(case.input["user_id"])
            tool.set_output(profile)

        answer = compose_answer(profile, case.input["message"])
        agent.set_output(answer)
        agent.metric(
            "task_completed",
            True,
            semantic_type=Semantic.AGENT_TASK_COMPLETION,
        )
        return answer
```

Spans preserve selection, arguments, output, order, duration, errors, tags, and hierarchy.

## Declare Expected Actions

Cases can use generic `actions` or the tool-oriented `tool_calls` compatibility shape:

```yaml
cases:
  - id: refund
    input:
      user_id: u1
      message: Refund order 42
    expected:
      actions:
        - id: lookup
          kind: tool
          target: lookup_user
          input:
            user_id: u1
          order: 1
          required: true
```

Expected input matching is subset-based, so a tool may receive additional nonessential arguments.
Actions may also declare expected output, tolerance metadata, optional status, and explicit order.

## Score Selection, Arguments, And Sequence

```yaml
score:
  tool_selection:
    expected_action:
      metric: selection
      observed_kind: tool
      span:
        kind: tool
    semantic: agent.tool.selection.correctness
    goal: maximize

  tool_arguments:
    expected_action:
      metric: arguments
      observed_kind: tool
      span:
        kind: tool
    semantic: agent.tool.argument.correctness
    goal: maximize

  tool_sequence:
    expected_action:
      metric: sequence
      observed_kind: tool
      span:
        kind: tool
    semantic: agent.tool.sequence.correctness
    goal: maximize
```

These scorers are deterministic and do not require an LLM judge. They produce normal scores and
semantic observations, so policies and reports consume them like any other metric.

## Span Selection

`SpanSelector` filters spans by:

- kind
- name
- tags
- nested path
- emitted semantic type

Selectors can be composed with positive and negative report/evaluation filters. A scorer receives
the selected spans through `ScoringCall`, allowing custom component-level evaluators without
parsing raw traces.

## Agentic Semantic Types

Built-in semantics include:

- task completion and goal accuracy
- plan quality and plan adherence
- step efficiency and orchestration quality
- tool name and version
- tool selection, argument, and sequence correctness
- tool-call quality
- output correctness and structure validity
- agent version and serving volume

Applications may add more specific child semantics through the registry.

## Metric Packs

The `agentic` metric pack contributes standard semantic definitions and report defaults. Metric
packs are optional: they provide conventions, not a required agent SDK. A custom agent runtime can
emit the same evidence through spans or a trace adapter.

## Optimization Feedback

`build_feedback_records` compacts run evidence into one record per case. It captures:

- score and evaluator reasons
- task, scorer, policy, and span errors
- `failure_category` only when a failure exists
- factor values and tracked asset versions
- selected observations and trace context

`build_optimization_feedback_input` packages those records with benchmark identity and semantic
context. pydantic-gepa or autoptimize can consume this structured evidence without scraping Rich
tables or replay YAML.

Autobench reports association and comparison evidence; it does not claim causal attribution when
multiple factors changed together. Controlled experiment planning belongs to the optimizer layer.

---

## Recording And Reporting

Canonical page: https://vcoderun.github.io/autobench/recording-and-reporting/

# Recording And Reporting

Recording turns an experiment into portable evidence while it is running. Each completed matrix
item is committed to a mutable staging directory; only a validated complete or explicitly partial
experiment is published as an immutable record. Replay and analysis use the immutable record
without executing the application again.

## Record Layout

```bash
autobench run autobench.yaml --record runs/support-routing
```

The directory contains:

```text
runs/support-routing/
  experiment.yaml
  summary.yaml
  manifest.yaml
  cases/<case-id>/<variant-id>/run.yaml
  assets/index.yaml
  assets/<safe-asset-id>.yaml
  artifacts/asset-content.sqlite3
  artifacts/<other-payloads>...
```

Paths are stable and artifact references are relative so the directory can be moved or archived.
Recording is append-only: an existing run payload is never silently replaced.

The CLI and `FileRecorder` stage each completed run before the next serial run, or independently as
concurrent runs complete. Final publication preserves matrix-plan order rather than wall-clock
completion order. `record_experiment()` remains the compatible one-shot API for an already
in-memory result.

Both finalization paths build the complete record in a temporary sibling directory, write
`experiment.yaml` and `summary.yaml`, validate `manifest.yaml`, and only then publish the final
directory with one atomic rename. A normal process failure therefore leaves either no new final
directory or a complete one; readers never observe a half-written final experiment.

Each manifest entry records the relative path, SHA-256 hash, byte count, file kind, and logical
identity of one file. The manifest excludes itself to avoid a recursive hash. Replay validates the
manifest before loading runs, so changed, missing, or unexpected payloads fail explicitly.

Asset manifests contain identity, lineage, hashes, changed fields, and content references. The
versioned prompt/tool/schema snapshots and readable diffs live in the single experiment-local,
content-addressed `artifacts/asset-content.sqlite3` registry. Resolve them with
`load_asset_content(...)` and `load_asset_diff(...)`.

## Incremental Durable Recording

Use `FileRecorder` when completed runs must survive a later task, scorer, policy, recorder, or
process failure:

```python
import asyncio
from pathlib import Path

from autobench import FileRecorder, run_benchmark_spec

output = Path("runs/routing-42")
result = asyncio.run(
    run_benchmark_spec(
        spec,
        experiment_id="routing-42",
        concurrency_limit=4,
        recorder=FileRecorder(output, durability="atomic"),
    )
)
```

The pipeline owns the recording session lifecycle. It opens the recorder after constructing the
fixed experiment plan, stages each `ExecutionSnapshot`, finalizes after cross-run derivation and
policies, aborts on failure, and closes under success, failure, or cancellation. A recorder failure
is a required persistence failure: Autobench does not report a run as durably recorded when its
snapshot was not committed.

Without `recorder=`, `run_benchmark_spec()` stays purely in memory and creates no staging files.
The CLI constructs a `FileRecorder` for `--record` and for its default `.autobench/...` destination;
`--no-record` selects the in-memory path explicitly.

During execution, a sibling staging directory is used:

```text
runs/.routing-42.staging/
  staging.yaml
  staging-manifest.yaml
  cases/<case-id>/<variant-id>/run.yaml
  checkpoints/<run-id>/<name>.yaml
  artifacts/<run-id>/...
  assets/<run-id>/...
```

`staging.yaml` owns experiment identity, the immutable plan, environment, semantic registry,
source hashes, post-processing requirements, and session state. `staging-manifest.yaml` is the
commit index for run and checkpoint payloads. Both use versioned JSON Schema headers. A run is
recoverable only after all of its files and hashes appear in the manifest.

Staging is intentionally not accepted by `replay_experiment()`. Mutable execution state and
immutable experiment evidence are different formats.

## Explicit Checkpoints And Cancellation

An async task can commit the evidence collected so far without ending its run:

```python
from autobench import Case, RunContext


async def evaluate_route(ctx: RunContext, case: Case) -> dict[str, str]:
    route = await choose_route(case.input)
    ctx.metric("route_confidence", route.confidence)
    ctx.artifact("route_preview", route.model_dump(mode="json"))
    await ctx.checkpoint("route-selected")

    response = await execute_route(route)
    return {"response": response}
```

`checkpoint()` snapshots the current run phase, available task output, observations, errors,
legacy spans, canonical ABP trace, artifacts, tracked asset versions and uses, source snapshots,
and signal-sequence watermark. It returns only after the staging manifest commits the checkpoint.
Checkpoint names beginning with `autobench.` are reserved for runtime lifecycle records.

Explicit checkpoints require durable recording. Calling `ctx.checkpoint()` in an in-memory run
raises before pretending evidence was persisted. A task that does not need interruption recovery
does not pay checkpoint or staging overhead.

On cooperative cancellation Autobench:

1. records the original `CancelledError`;
2. finalizes open ABP spans as partial;
3. commits an `autobench.cancelled` checkpoint with the phase where cancellation occurred;
4. cancels active concurrent siblings and gives each bounded time to perform the same cleanup;
5. marks staging cancelled, closes the record session, and re-raises the original cancellation.

Recorder commits are cancellation-safe boundaries. Once a run has produced its complete
`RunResult`, cancellation does not discard it while `stage()` is publishing the payload or
manifest revision. Autobench keeps ownership of that commit, waits for its terminal state within
the cleanup bound, and computes `recorded_run_ids` only after outstanding recorder operations have
settled. `abort()` and `close()` are ordered after stage and checkpoint work, so they never race a
surviving file-system worker.

The same rule applies to final publication. If cancellation arrives after `finish()` has begun,
the finalization remains owned until it either commits or fails; Autobench does not report an
untracked publication and let a final directory appear later. A non-cooperative third-party
cleanup may outlive the public wait bound because Python cannot forcibly terminate an arbitrary
coroutine. Such a task is retained, its eventual exception is delivered to the event-loop exception
handler, and the cancellation receives a diagnostic that cleanup is still active.

If cancellation happens during scoring or derivation, a task output already produced by the
application remains in the partial snapshot. Recorder failures are attached to the cancellation as
notes and never replace its identity. A task timeout remains a timeout and is not reclassified as
cooperative cancellation.

The CLI maps `SIGTERM` to this cooperative path where event-loop signal handlers are supported.
`KeyboardInterrupt` also reaches pipeline cancellation before the CLI exits. `SIGKILL`, abrupt
runtime termination, and power loss cannot run cleanup; they preserve only payloads already
committed to `staging-manifest.yaml`. With `durability="synced"`, those commits receive the
documented filesystem-sync guarantee. Autobench does not claim that uncommitted in-memory evidence
survives a hard kill.

Autobench does not write on every signal and does not resume application code from a checkpoint.
The application owns executable workflow state; Autobench owns durable evidence. Automatic
periodic checkpoint policy is intentionally outside this release.

Async file artifacts use the same ownership model. `artifact_file_async()` returns cancellation
within a bounded interval even when a filesystem read is blocked. The run retains a partial
artifact reference, while the session owns the underlying transfer and settles it before
checkpoint, abort, final publication, or close. A late transfer cannot be mistaken for a complete
artifact before its payload is available.

## Inspect And Recover Staging

```bash
autobench recording inspect runs/.routing-42.staging
autobench recording finalize runs/.routing-42.staging \
  --output runs/routing-42-recovered \
  --allow-partial
autobench recording archive runs/.routing-42.staging \
  --output archives/routing-42-staging
autobench recording discard runs/.routing-42.staging --yes
```

Inspection reports health, recoverability, complete and checkpointed runs, missing identities,
corrupt or conflicting runs, orphaned files, and diagnostics. The health states are:

| Health | Meaning |
| --- | --- |
| `complete` | Every planned run has committed, valid evidence |
| `partial` | Some committed run or checkpoint evidence exists |
| `missing` | No planned run has committed evidence yet |
| `corrupt` | A manifest-committed file is missing, malformed, or has the wrong hash |
| `conflicting` | Identity, revision, or uncommitted-file state needs an explicit decision |

Not every conflict destroys committed evidence. Files left by a process failure before manifest
commit and a state/manifest revision mismatch are reported as recoverable; recovery ignores those
uncommitted files and trusts the last manifest revision. Plan, experiment, run, checkpoint, or
payload identity conflicts are not recoverable because choosing one side would silently rewrite
lineage.

Python callers can inspect or load the committed subset without application imports:

```python
from pathlib import Path

from autobench import inspect_staging, recover_staging

staging = Path("runs/.routing-42.staging")
inspection = inspect_staging(staging)
if inspection.recoverable:
    recovered = recover_staging(staging)
    print([run.run_id for run in recovered.runs])
```

`finalize_staging(..., allow_partial=False)` refuses an incomplete matrix. With
`allow_partial=True`, committed runs and the newest checkpoint per missing run become one immutable
partial experiment. Its terminal metadata lists planned, recorded, and missing run IDs and marks
incomplete cross-run derivation or policies. The source staging directory remains until it is
archived or explicitly discarded.

A finalized cancellation uses normal `RunRecord`, `ExperimentRecord`, replay, and report paths.
Recovered checkpoint runs have `RunStatus.CANCELLED`, `TaskStatus.CANCELLED`,
`EvaluationStatus.NOT_EVALUATED`, `partial=true`, and `end_reason=cancelled`; they are not rewritten
as failed evaluations.

## RunRecord

One `RunRecord` represents one case x variant execution:

- record, run, experiment, benchmark, case, and variant IDs
- final, task, and evaluation statuses
- explicit `partial` state and ABP `end_reason`, including cancelled runs
- complete case snapshot and task output
- observations and scores
- canonical ABP trace, including signals, span graph, measurements, events, links, references,
  diagnostics, and instrumentation scope provenance
- ABP protocol and semantic registry versions
- legacy span tree for records created before canonical trace storage
- materialized artifacts
- factors and tracked asset versions
- extraction and source-map replay lineage
- immutable invocation correlation: group, attempt, phase, external experiment associations, and
  scalar labels
- structured errors

The YAML view groups the data for people rather than dumping internal Pydantic fields. A schema
header points editors to the versioned Autobench JSON schema.

Small traces remain inline in `run.yaml`. Larger traces are written to
`artifacts/<run-id>/trace.yaml`; the RunRecord keeps a relative `ArtifactRef` and a compact trace
summary. Trace artifacts have their own versioned JSON Schema header and load back into the same
typed `Trace` model.

## ExperimentRecord

The experiment-level record stores:

- benchmark plan and counts
- captured environment metadata
- semantic registry
- report configuration
- normalized benchmark snapshot and hash
- hashes of resolved specs, datasets, pricing files, tasks, and scorer modules
- relative run paths and status counts
- terminal experiment state: completed, cancelled, or aborted
- planned, recorded, and missing run identities
- whether cross-run derivation and experiment policies completed
- the relative integrity-manifest path
- the same resolved execution correlation copied to every run

This is enough to explain what was planned, which files defined it, and where every run record
lives.

An experiment may be terminal and still partial. Cancelled or recovered evidence uses the same
immutable record models as complete evidence, while `EvaluationStatus.NOT_EVALUATED` distinguishes
a cancelled task from a scored failure. Records written before format version 5 load as completed,
non-partial experiments with complete post-processing. Format version 6 adds optional execution
correlation; older records load it as `None`.

Execution correlation is not replay lineage. `parent_run_id` and `RecordLineage` describe derived or
replayed evidence, while `ExecutionCorrelation` groups separately invoked experiments. Reports,
Rich tables, YAML/CSV exports, staging inspection, finalized partial records, and replay preserve
that distinction.

## Atomicity And Durability

Python callers can select the publication guarantee:

```python
record_experiment(result, Path("runs/latest"), durability="atomic")
record_experiment(result, Path("runs/durable"), durability="synced")
```

- `atomic` uses sibling temporary files/directories and `os.replace`. It protects readers from
  partially published records after an ordinary process failure.
- `synced` adds file and directory `fsync` calls before and after publication on supported POSIX
  filesystems. It is intended for callers that also require the strongest available power-loss
  durability.

`synced` fails explicitly when directory syncing is unsupported. Autobench does not silently call
an atomic record power-loss durable. Neither mode overwrites a non-empty experiment directory.

Atomicity applies to each staged payload and final-directory publication. The staging manifest is
the commit boundary: a file that exists but is absent from the manifest is uncommitted evidence,
not a completed run. `synced` also syncs staged state and manifest files where supported.

## Environment And Source Identity

`capture_environment` records reproducibility metadata such as Python, platform, package, and
working-environment details. `collect_benchmark_source_files` resolves benchmark dependencies and
records content hashes.

Source paths are stored portably when possible. Missing optional source files do not erase a run;
recording captures what was resolvable at execution time.

## Artifacts

`ctx.artifact(name, value)` adds an `ArtifactRef`. During recording, supported values are
materialized under `artifacts/` and the RunRecord keeps the relative path, media type, and tags.

Use artifacts for:

- generated specs and prompts
- traces too large for `run.yaml`
- measurement samples
- model responses and structured debug payloads
- Markdown or text reports produced by the subject

Artifact path collisions and attempts to overwrite existing payloads are recording errors.

## Replay

```bash
autobench replay runs/support-routing
```

Replay loads `ExperimentRecord` and every `RunRecord` into an `ExperimentResult`. It deliberately
does not import task or scorer modules, call models, or mutate the original directory.

This enables:

- offline report regeneration
- new exports from old evidence
- baseline/candidate comparison after execution
- future rescoring into a separate derived experiment
- optimization systems consuming stable records

Autobench distinguishes three replay modes:

- **report replay** reads stored observations without re-extracting evidence
- **extraction replay** runs a typed `TraceExtractor` against the immutable ABP trace and creates a
  derived RunRecord
- **canonicalization replay** applies newer source maps to retained source snapshots and creates a
  separate derived RunRecord

Derived records point to the original `run_id`, identify the extractor or source-map versions, and
retain the source protocol and semantic registry versions. The original record and trace bytes are
never rewritten. Replay resolves trace artifacts only inside the experiment directory and imports
neither application task modules nor optional SDK integrations.

The default `SignalExtractor` reconstructs canonical observations from stored ABP measurements and
events. `SpanExtractor` derives generic topology and workflow evidence, while `UsageExtractor`
owns LLM request/token/model accounting. `CompositeExtractor` can run them as one versioned replay
processor. Custom extractors implement the typed `TraceExtractor` interface and return
observations, diagnostics, and evidence references without mutating the trace.

When a newer version of the same extractor is replayed, its observations replace the older
version's observations in the new derived record. The previous derived record remains the lineage
parent, so extractor evolution is auditable without mixing two versions of one derived metric.

## Rich Reports

```bash
autobench report runs/support-routing
```

The terminal report can include:

- experiment overview and status counts
- variant configuration table with factor values
- semantic leaderboards
- per-run metric tables grouped by semantic family
- case x variant matrices
- baseline/candidate factor and metric deltas
- metric distributions

Reports use projected semantic metrics. They do not depend on application-specific local names.

## Report Configuration

```yaml
report:
  leaderboard:
    show:
      accuracy:
        metric: quality.correctness
        aggregate: ratio_true
      total_cost:
        metric: money.cost
        aggregate: sum
      p95_latency:
        metric: time.latency
        aggregate: p95
  matrix:
    metric: quality.correctness
  compare:
    baseline -> candidate:
      show:
        accuracy:
          metric: quality.correctness
          aggregate: ratio_true
  distributions:
    - name: request_latency
      semantic_type: time.latency
      summaries: [min, median, p95, max]
  markdown:
    profile: full
    layout: auto
    output: reports/benchmark.md
```

Aggregation functions include count, mean, sum, min, max, median, p95, standard deviation,
geometric mean, and boolean true ratio.

The default Markdown projection is decision-facing: quality gate, score range, purposeful inline
SVG, case outcomes, paired deltas, issue totals, and priority evaluator feedback. The `audit`
profile adds run health, metric coverage, ABP traces, asset lineage, artifact inventory, optimizer
evidence, hashes, and provenance. A configured `output` is staged before the record manifest is
sealed. See [Markdown Reports](markdown-reports.md) for the task-output evaluation convention,
profiles, bundle layout, audit safety, and publication.

## Comparison Semantics

```bash
autobench compare runs/support-routing --baseline baseline --candidate candidate
```

Comparison pairs runs by case, displays changed factors, aggregates requested semantic metrics, and
sets `confounded=true` when multiple relevant factors changed. It reports association and deltas;
it does not claim which factor caused the result.

Use paired-baseline post-derivation when a per-run derived metric such as speedup must be written
back into candidate evidence.

## Exports

```bash
autobench export runs/support-routing --format yaml --path report.yaml
autobench export runs/support-routing --format csv --path runs.csv
autobench export runs/support-routing --format markdown --path report.md
```

- YAML is a human-readable summary projection.
- CSV is a flat run-and-metric table for analysis tools.
- Markdown is a portable rendered report.

The CLI always writes the requested file and then renders a Rich preview. Machine exports never
replace immutable source RunRecords.

---

## Markdown Reports

Canonical page: https://vcoderun.github.io/autobench/markdown-reports/

# Markdown Reports

Autobench Markdown reports explain benchmark outcomes to the people deciding whether a system is
good enough to ship. They are designed for release reviews, pull requests, model or pipeline
comparisons, incident follow-up, and long-lived experiment archives. The default report starts with
the quality gate, score range, case-level failures, meaningful charts, and evaluator explanations.
Run IDs, hashes, traces, asset internals, and other engineering evidence remain available in the
explicit `audit` profile instead of overwhelming the normal report.

Markdown is not the source of truth. Immutable experiment and run records remain authoritative;
the report is a deterministic projection that can be regenerated without running the application.

## What The Report Contains

A full report can include:

- a plain-language executive summary and benchmark verdict;
- quality-gate pass rate, average, median, best, and lowest case score;
- deterministic inline charts for quality-gate composition, case ranking, and normalized dimensions;
- a case table that separates benchmark quality from execution success;
- issue totals such as omissions, leaks, failures, and policy violations;
- priority evaluator feedback for the lowest-scoring failed cases;
- variant configuration, direction-aware leaderboards, and paired comparisons;
- absolute and relative deltas, win/tie/loss counts, and explicit confounding warnings;
- configured matrices and distributions when they add decision value;
- a concise benchmark setup and optimization outcome.

The `audit` profile adds lifecycle status, metric coverage, run IDs, failures, ABP traces, asset
lineage, artifacts, hashes, provenance, and explicitly permitted captured content.

Autobench uses fixed deterministic finding rules. It does not call an LLM to write the summary,
claim statistical significance, or infer causality from a confounded comparison.

## Terminal Versus Document

The default command remains the interactive Rich report:

```bash
autobench report runs/support-routing
```

Generate a Markdown document explicitly:

```bash
autobench report runs/support-routing \
  --format markdown \
  --profile full \
  --layout single \
  --output analysis/support-routing.md
```

The CLI prints a Rich publication summary containing the selected layout, file count, byte count,
run count, section count, and notices. It never dumps the Markdown document into the terminal.

The compatibility export command uses the same single-file writer:

```bash
autobench export runs/support-routing \
  --format markdown \
  --path analysis/support-routing.md
```

Both commands replay stored evidence. They do not import the task module, contact a model provider,
or execute the benchmark subject.

## Profiles

| Profile | Purpose | Included detail |
| --- | --- | --- |
| `summary` | fast decision review | verdict, quality KPIs, charts, bounded case outcomes, comparisons, policies, and material limitations |
| `full` | stakeholder benchmark report | summary plus benchmark setup, configured analysis, and priority evaluator feedback |
| `audit` | engineering evidence inspection | full report plus lifecycle, metric coverage, runs, failures, traces, assets, artifacts, hashes, provenance, and permission-gated captured content |

`full` is the default. A run whose task executed successfully can still fail the benchmark quality
gate; reports present these as separate facts. `audit` does not automatically reveal captured content. The command must
also receive `--include-captured-content`, and the original capture/redaction policy must have
retained that content. Reporting can never recover data that was not recorded or weaken a sensitive
asset policy.

`limits.value_excerpt_chars` bounds each captured value or diff excerpt independently. Autobench
serializes JSON-compatible values deterministically and marks unsupported values unavailable rather
than calling application code. `assets.diffs` controls asset history independently:

- `none`: omit changed fields and diff detail;
- `summary`: show changed field paths and whether stored diff evidence exists;
- `full`: allow a bounded stored diff excerpt, but only in an explicitly content-enabled audit.

Sensitive assets remain omitted in every profile, including an audit with content permission.

```bash
autobench report runs/support-routing \
  --format markdown \
  --profile audit \
  --include-captured-content \
  --output analysis/support-routing-audit.md
```

## Quality Outcome Convention

Autobench does not equate a completed task call with a successful benchmark outcome. A task may
return a mapping or Pydantic model with a report-ready evaluation envelope:

```python
return {
    "hard_pass": result.meets_release_gate,
    "score": result.overall_score,
    "metrics": {
        "semantic_score": result.semantic_score,
        "critical_omissions": result.critical_omissions,
        "forbidden_leaks": result.forbidden_leaks,
    },
    "feedback": result.feedback,
}
```

The same fields may live under an `evaluation` key. `hard_pass`, `passed`, and `pass` are accepted
quality-gate names. `feedback` may be one string or a list of strings. Numeric score records remain
valid evidence and provide a fallback objective score when the output has no `score` field.

Normalized score-like metrics become quality dimensions. Boolean quality metrics are summarized as
rates. Counts are shown as quality issues only when their names describe failures, omissions, leaks,
violations, or similar problems; neutral measurements such as output length remain out of the issue
table. If no recognizable quality evidence exists, Autobench does not invent a verdict or chart.

## Layouts

### Single

`single` writes one portable `.md` file. Use it for small and medium experiments, pull requests,
and attachments.

### Bundle

`bundle` publishes a directory containing:

```text
benchmark-report/
  index.md
  cases/<stable-id>.md
  variants/<stable-id>.md
  runs/<stable-id>.md       # audit only
  assets/<stable-id>.md     # audit only
```

The index contains the decision summary and links to user-facing case and variant pages. Technical
run and asset pages are created only for `audit`. Page names use normalized identifiers plus a
stable hash suffix, so human-readable IDs do not silently collide.

```bash
autobench report runs/support-routing \
  --format markdown \
  --layout bundle \
  --output analysis/support-routing-report
```

### Auto

`auto` chooses from report size, never terminal width. For normal reports, case evidence determines
detail size. For audits, run, failure, asset-version, and artifact inventories are included. It
selects a bundle when any of these deterministic limits is exceeded:

- more than 50 runs;
- more than 100 relevant detail rows;
- more than 1,000 case-matrix cells.

The selected layout is returned in `MarkdownReportPublication` and shown by the CLI.

## YAML Configuration

Configure Markdown alongside leaderboard, matrix, comparison, and distribution views:

```yaml
report:
  leaderboard:
    show:
      quality:
        metric: quality.score
        aggregate: mean
      total_cost:
        metric: money.cost
        aggregate: sum
      p95_latency:
        metric: time.latency
        aggregate: p95
  compare:
    baseline -> candidate:
      show:
        quality:
          metric: quality.score
          aggregate: mean
        p95_latency:
          metric: time.latency
          aggregate: p95
  markdown:
    profile: full
    layout: single
    output: reports/benchmark.md
    limits:
      table_rows: 200
      run_details: 100
      failure_details: 100
      value_excerpt_chars: 2000
    traces:
      top_slowest: 20
    assets:
      diffs: summary
    content:
      include_captured: false
```

`output` must be a portable relative path without `..`. When `autobench run` records through the
CLI, a configured output is generated in the staging record before final metadata and manifest
hashes are sealed. A failed publication prevents the record from being presented as successfully
finalized.

Omit `output` when reports should only be generated post-hoc.

## Python API

Build, render, or publish the same typed report model:

```python
from pathlib import Path

from autobench import (
    build_report,
    load_experiment_record,
    replay_experiment,
    render_markdown_report,
    write_markdown_report,
)

record_dir = Path("runs/support-routing")
result = replay_experiment(record_dir)
record = load_experiment_record(record_dir)
report = build_report(
    result,
    experiment_record=record,
    experiment_root=record_dir,
)

markdown = render_markdown_report(report)
publication = write_markdown_report(
    report,
    Path("analysis/support-routing"),
    layout="bundle",
    immutable_root=record_dir,
)

print(publication.layout)
print([(file.path, file.byte_count, file.sha256) for file in publication.files])
```

`render_markdown_report()` returns text and performs no write. `write_markdown_report()` owns
atomic publication and returns `MarkdownReportPublication` with every output path, byte count, and
SHA-256 hash. Existing `export_markdown_report(result, path)` remains supported and delegates to the
same single-file writer.

When `immutable_root` is supplied, the writer rebases every validated artifact link relative to the
Markdown document that contains it. Index and nested run pages therefore resolve to the same
recorded payload even when a bundle is published beside the immutable record. Direct rendering
accepts `record_link_prefix` for callers that own publication themselves.

For a custom recorder-level experiment publication, pass an `ExperimentPublisher` to
`FileRecorder`. The generic seam receives the current result, the not-yet-sealed experiment record,
and the staging root. Treat the root as read-only and return typed `ExperimentFile` values:

```python
from collections.abc import Sequence
from pathlib import Path

from autobench import ExperimentFile, ExperimentRecord, ExperimentResult


def publish_analysis(
    result: ExperimentResult,
    record: ExperimentRecord,
    experiment_root: Path,
) -> Sequence[ExperimentFile]:
    # Inspect staged record evidence through experiment_root; do not write to it directly.
    return (
        ExperimentFile(
            path="reports/custom.txt",
            content=f"runs={record.run_count}\n".encode(),
            identity=f"custom-report:{result.experiment_id}",
        ),
    )
```

The recorder validates collisions and writes returned files before sealing final metadata. The CLI
uses `MarkdownExperimentPublisher` to honor configured report output. Supplying the staging root is
important: configured reports can inspect the same recorded asset content and materialized artifact
paths as post-hoc replay reports.

## Comparison Integrity

Comparisons are paired by case. Each metric reports:

- baseline and candidate aggregate;
- absolute and relative delta;
- metric direction and resulting `improved`, `regressed`, `unchanged`, or `indeterminate` outcome;
- paired and missing pair counts;
- paired win/tie/loss count.

A lower latency or cost can therefore improve while a higher quality score improves. If direction
metadata conflicts across observations, Autobench does not mark a global winner. If multiple
factors changed, the report remains explicitly confounded and uses association language.

## Evidence And Missingness

Reports keep missing values distinct from numeric zero. Every leaderboard metric includes observed
and missing sample counts, and case matrices leave absent cells empty. Non-zero micro-costs retain
enough precision to avoid appearing as zero.

Findings link to typed experiment, run, metric, comparison, trace, asset, artifact, or optimization
identities. They describe the evidence available; they do not manufacture data for a cleaner story.

Artifact links come from the materialized path stored in the recorded `ArtifactRef.value`, never
from a user-facing filename. Autobench rejects absolute, escaping, missing, symlink-escaped, or
hash/size-mismatched payloads and reports a notice instead of emitting a broken or unsafe link.
Malformed optional run or asset detail cannot erase valid aggregate evidence.

## Safety And Immutability

Markdown values are treated as untrusted application data:

- HTML is escaped;
- table separators and line breaks are neutralized;
- links are relative and containment-checked;
- binary artifacts are referenced rather than embedded;
- remote images and active HTML are never injected;
- captured content is bounded and permission-gated;
- report generation performs no network I/O.

Post-hoc publication refuses destinations inside a finalized experiment root. Write the report next
to the record or into a separate analysis directory. Configured in-run publication is safe because
the files are staged before `experiment.yaml` and `manifest.yaml` are finalized.

Single files use a sibling temporary file and atomic replacement. Bundles are fully rendered in a
sibling staging directory before publication; an overwrite replaces the old bundle only after the
new bundle is complete.

## Charts And Optional PDF

Autobench emits deterministic inline SVG only when the evidence has a defined visual meaning:

- a stacked quality-gate pass/fail bar;
- a case-score ranking with pass/fail color encoding;
- normalized quality-dimension averages with sample counts.

Charts escape labels, use no remote assets, perform no network I/O, and remain readable as vector
graphics in Markdown viewers and PDF. Autobench intentionally does not guess arbitrary plots from
unclassified metrics. Matplotlib, Mermaid, dashboarding, and custom chart DSLs remain outside the
core report dependency graph.

PDF is an optional downstream projection rather than an Autobench runtime dependency. One supported
conversion uses `markdown-pdf>=1.13.2`:

```python
from pathlib import Path

from markdown_pdf import MarkdownPdf, Section

source = Path("analysis/support-routing.md")
pdf = MarkdownPdf(toc_level=2, optimize=True)
pdf.add_section(
    Section(
        source.read_text(encoding="utf-8"),
        root=str(source.parent),
        paper_size="A4",
    )
)
pdf.save("analysis/support-routing.pdf")
```

Keep PDF conversion outside the immutable benchmark record unless it is installed as an explicit
recorder publisher. The Markdown report remains the portable, deterministic document source.

---

## Asset Tracking

Canonical page: https://vcoderun.github.io/autobench/asset-tracking/

# Asset Tracking

Benchmarks need to know which prompt, tool, schema, or configuration produced each result.
Autobench tracking assigns content-derived versions, captures structured metadata, persists history,
and binds exact asset versions to RunRecords.

Supported SDK instrumentors can discover these assets without decorators. Use explicit tracking on
this page when the application owns a better logical identity or when the component never crosses an
instrumented boundary. See [Automatic Asset Discovery](automatic-asset-discovery.md) for unannotated
Pydantic AI, OpenAI, OpenAI Agents, capability scopes, privacy, and custom SDK extraction.

## Prompts And Text Assets

Track inline text:

```python
from autobench import track

SYSTEM_PROMPT = track.prompt(
    name="support_system_prompt",
    text="Route the request to billing, account, or technical support.",
)
```

Or load it from a file:

```python
SYSTEM_PROMPT = track.prompt(
    name="support_system_prompt",
    source="prompts/support.md",
)
```

`TrackedPrompt.raw` returns the text, and `str(SYSTEM_PROMPT)` provides the same value for APIs that
expect a string. File-backed prompts retain their source path and source hash.

## Tools

`@track.tool` preserves the callable's exact signature and return type while collecting tool
metadata:

```python
from typing import Literal

from autobench import track


@track.tool
def route_ticket(
    queue: Literal["billing", "account", "technical"],
    priority: int = 1,
) -> bool:
    """Route a ticket to a support queue."""
    return priority > 0
```

The resulting `ToolAsset` records:

- qualified name and docstring
- parameter names, kinds, annotations, defaults, and requirements
- return annotation
- source path and source hash
- structured parameter schema
- semantic type and version lineage

Annotations are normalized by structure rather than alias spelling. If the contents of a
`Literal`, union, generic, model, or referenced type change, the asset hash changes even when the
alias name stays the same.

## Pydantic Models, Dataclasses, And Classes

```python
from dataclasses import dataclass
from typing import Literal

from autobench import track
from pydantic import BaseModel, Field


@track.type
class Car(BaseModel):
    make: Literal["audi", "bmw", "mercedes"]
    model: str = Field(examples=["a3", "320i"])
    year: int = Field(gt=0)


@track.dataclass(frozen=True, slots=True)
class CarRequest:
    make: Literal["audi", "bmw", "mercedes"]
    model: str
    year: int
```

Pydantic models are hashed from normalized JSON Schema plus source identity. Standard dataclasses
use dataclass field definitions and resolved annotations. Other typed classes use resolved class
annotations, inspectable signatures, and source hashes.

`TypeAsset` and `FieldAsset` preserve field names, resolved annotations, descriptions, examples,
aliases, defaults, required state, and relevant constraints.

## Composing Another Class Decorator

When `@track.type` above a class-transforming decorator gives poor type-checker inference, use
`track.decorate_type`:

```python
from dataclasses import dataclass

from autobench import track


@track.decorate_type(dataclass, frozen=True, slots=True)
class Request:
    value: str
```

The decorator and its normalized arguments are stored as asset metadata. `track.dataclass(...)` is
the typed convenience form for the standard dataclass decorator.

## Arbitrary Assets

Use `track.asset` for configurations, policies, routing tables, or other application components:

```python
@track.asset(kind="routing_policy", name="enterprise_routing")
def route_policy(ticket):
    return "priority" if ticket["enterprise"] else "standard"
```

The decorator returns the original object unchanged. Callables use source and signature metadata;
manual `version`, `hash`, `source_path`, `parent_version`, and metadata values are available when
automatic identity is not enough.

## Versions, Diffs, And Persistence

`TrackingRegistry` keeps current assets and version history in memory during execution. Persist it
with:

```python
from pathlib import Path

from autobench import track

track.write_assets(Path(".autobench/assets"))
```

The directory contains an index, one lightweight manifest per asset, and `content.sqlite3`, a local
content-addressed registry keyed by asset ID and version. Every new manifest version links to its
parent and stores changed paths plus typed content and diff references. Prompt text, tool bodies,
schema definitions, and readable diffs never appear in the manifests. SQLite transactions avoid
rewriting the complete history as it grows, while content hashes deduplicate identical snapshots.

```python
from autobench import load_asset_content

version = track.asset_version_of(SYSTEM_PROMPT)
snapshot = load_asset_content(
    Path(".autobench/assets/content.sqlite3"),
    asset_id=version.asset_id,
    version=version.version,
)
assert snapshot["raw"] == SYSTEM_PROMPT.raw
```

For a version with a parent, resolve its readable diff with
`load_asset_diff(path, asset_id=..., version=..., parent_version=...)`.

## Binding Assets To Runs

```python
def run_case(ctx, case):
    ctx.attach_tracked_asset(SYSTEM_PROMPT)
    ctx.attach_tracked_asset(route_ticket)
    ctx.attach_tracked_asset(Car)
    return execute(case.input)
```

The exact `AssetVersion` values are copied into the RunRecord. Reports and optimization feedback
can then relate metric changes to prompt, tool, or output-schema versions without guessing from
source control state.

---

## Automatic Asset Discovery

Canonical page: https://vcoderun.github.io/autobench/automatic-asset-discovery/

# Automatic Asset Discovery

Automatic asset discovery turns SDK-visible prompts, tools, output schemas, capabilities, agents,
guardrails, handoffs, and policies into versioned benchmark evidence. It is part of the common ABP
instrumentation runtime, not a Pydantic AI-specific tracking mode.

The application does not need `@track.prompt`, `@track.tool`, or `@track.type` when a supported
instrumentor can already see the behavioral component at a stable SDK boundary. Explicit tracking
still composes with discovery when the application owns a better identity, semantic type, source
path, or parent version.

## Quick Start

Install the SDK integrations used by the application:

```bash
pip install 'autobench[instrumentation]'
```

Then enable compatible integrations for the benchmark:

```python
from autobench import Benchmark

benchmark = Benchmark("support-agent").instrument_all()
```

While a benchmark run is active, Autobench now:

1. observes definitions at supported framework and client surfaces;
2. normalizes them into SDK-independent asset candidates;
3. resolves stable logical identity and aliases;
4. computes behavioral content versions;
5. attaches exact asset uses to the owning run and span;
6. persists referenced histories when the experiment is recorded.

Calls made outside an active Autobench run remain unchanged and produce no discovery evidence.

## What Counts As An Asset

An asset is a versionable component whose content or behavior can change benchmark outcomes.

| Observed value | Autobench treatment |
| --- | --- |
| static or callable instructions | prompt definition |
| rendered system/developer instructions | effective prompt |
| function, hosted, native, or MCP tool | tool definition or effective tool schema |
| Pydantic type, dataclass, or JSON Schema | output schema |
| Pydantic AI capability | scoped composite asset |
| OpenAI Agents guardrail or handoff | guardrail or handoff asset |
| routing, retry, output, or tool-use configuration | policy asset |
| composed agent or toolset | composite asset with child locators |
| model/provider/settings | factor or configuration evidence |
| user input, output, message history, tool arguments/results | evidence, not an asset |
| tokens, cost, latency, quality | metric |

Kinds are open strings. Autobench does not require every domain to fit an AI-only taxonomy.

## Definitions And Effective Representations

Definitions and model-facing values answer different questions:

```text
definition
  Which source component did the application declare?

effective
  Which resolved representation did this operation actually use?
```

A dynamic instruction callable is a definition asset. Calling it during the SDK's normal lifecycle
may produce an effective prompt for one run. Autobench records both and links the effective
`AssetUse.definition_asset_id` and `definition_version` to the source definition. Case interpolation
therefore does not create a fake source edit for every input.

Autobench discovery never invokes instruction callbacks, tools, validators, or guardrails merely to
inspect them. It observes declaration values or values already resolved by the actual SDK lifecycle.

## Pydantic AI Without Tracking Decorators

This agent has no explicit Autobench tracking:

```python
from pydantic import BaseModel, Field
from pydantic_ai import Agent


class SupportAnswer(BaseModel):
    answer: str = Field(description="A grounded support answer.")
    queue: str


def lookup_policy(topic: str) -> str:
    """Return the active support policy."""
    return f"policy for {topic}"


agent = Agent(
    model,
    name="support-router",
    output_type=SupportAnswer,
    instructions="Use the policy tool before routing.",
    tools=[lookup_policy],
)
```

With `instrument_all()`, a run discovers the agent composite, prompt, function tool, toolset, output
schema, final request instructions, effective tool definitions, and validated output schema. The
plain Python function and Pydantic type keep their source-aware identities; no wrapper replaces
their signatures or results.

The complete offline example uses Pydantic AI's real `Agent` lifecycle and `TestModel`:

```bash
uv run python examples/automatic_assets/pydantic_ai_discovery.py \
  --record /tmp/autobench-pydantic-assets
```

## Capability Scopes

Pydantic AI capabilities are first-class scopes because multiple capabilities can expose an
`instructions` or `search` component with the same local name:

```python
from pydantic_ai.capabilities import AbstractCapability


class RetrievalCapability(AbstractCapability[None]):
    id = "retrieval"

    def get_instructions(self) -> str:
        return "Ground answers in retrieved evidence."
```

The local and global locators are both retained:

```text
retrieval:prompt:instructions
pydantic_ai:retrieval:prompt:instructions
pydantic_ai:retrieval:capability:self
```

Capability identity uses a non-empty `id`, then `get_serialization_name()`, then the stable module
and qualified class name. A shared explicitly tracked component keeps one logical version while its
uses retain each capability alias and provenance.

## OpenAI Client And OpenAI Agents

The OpenAI client instrumentor discovers provider-facing assets:

- Chat Completions system/developer messages, tools, legacy functions, and response format;
- Responses instructions, managed prompt references, tools, text format, and output schema;
- Pydantic output types passed through structured parsing surfaces.

These are effective representations because the client sees the final request rather than the
framework's source declaration. Embedding inputs remain evidence and produce no asset.

The OpenAI Agents instrumentor discovers public Agent and Runner definitions:

- instructions and prompt references;
- function, hosted, MCP, computer, and agent tools;
- output schema;
- input/output/tool guardrails;
- handoffs and routing/tool-use policy;
- the composed agent.

It does not run callback instructions, guardrails, handoffs, or tools for discovery. Native trace
processing and public Runner resolution are combined without changing callback counts.

HTTPX intentionally discovers no semantic assets. A transport request lacks ownership context and
may contain secrets. It remains transport evidence while framework/client layers provide the
authoritative asset projections.

## Selecting Asset Families

Definition and effective discovery are enabled by default for semantic instrumentors. Restrict the
surface from Python:

```python
from autobench import AssetDiscoverySettings, AssetRepresentation, Benchmark

benchmark = Benchmark("support-agent").instrument_all(
    assets=AssetDiscoverySettings(
        representations=(
            AssetRepresentation.DEFINITION,
            AssetRepresentation.EFFECTIVE,
        ),
        include=("prompt", "tool", "output_schema", "capability"),
    )
)
```

Or use the YAML DSL:

```yaml
# yaml-language-server: $schema=schemas/0.3.0/benchmark_schema.json
benchmark:
  support-agent:
    instrumentation:
      all:
        assets:
          discover: true
          representations: [definition, effective]
          include: [prompt, tool, output_schema, capability]
```

An explicit integration can use a different filter and overrides automatic selection:

```yaml
instrumentation:
  all:
    exclude: [openai]
  openai:
    assets:
      representations: [effective]
      include: [prompt, tool, output_schema]
```

Set `discover: false` to keep spans and metrics from that instrumentor while disabling only its
asset discovery.

## Privacy And Capture Policy

Asset definitions and runtime evidence share one `CapturePolicy`, but they have separate fallback
levels because they serve different purposes:

- `default_level` defaults to `metadata` for messages, model payloads, HTTP data, and other runtime
  evidence;
- `asset_default_level` defaults to `full` so a successful prompt, tool, or output schema can be
  reconstructed, diffed, replayed, and optimized later.

Configure a stricter asset policy when the experiment directory cannot retain behavioral content:

```python
from autobench import Benchmark, CaptureLevel, CapturePolicy

benchmark = Benchmark("private-agent").capture(
    CapturePolicy.hashed(
        semantic_overrides={
            "tool": CaptureLevel.FULL,
            "output_schema": CaptureLevel.FULL,
        },
        deny_paths=("assets.*:prompt:private_notes",),
    )
)
```

The equivalent YAML is validated and completed by the versioned schema:

```yaml
benchmark:
  private-agent:
    capture:
      default_level: hash
      asset_default_level: hash
      use_semantic_defaults: false
      semantic_overrides:
        tool: full
        output_schema: full
      deny_paths:
        - assets.*:prompt:private_notes
```

Capture levels are `none`, `metadata`, `hash`, `redacted`, and `full`. The preset constructors
`none()`, `metadata()`, `hashed()`, `redacted()`, and `full()` set both fallbacks together. A direct
`CapturePolicy()` keeps runtime evidence metadata-first while retaining captured asset definitions.
Semantic overrides apply to both paths, and secret field names are filtered by the capture
normalizer.

The private content fingerprint still drives behavioral versioning. Changing capture from hash to
full does not create a false asset version. Omitted values retain an omission marker and digest;
large allowed values can be stored as bounded artifact references.

## Explicit Tracking Composition

Explicit tracking wins when Autobench observes the exact Python target:

```python
from autobench import track


@track.tool(name="knowledge_search")
def search(query: str) -> list[str]:
    """Search the approved knowledge base."""
    return backend.search(query)
```

If Pydantic AI or another instrumented SDK receives `search`, automatic discovery reuses the
explicit `ToolAsset` identity and version. SDK and capability locators become aliases; Autobench
does not create a second source version. The final provider schema remains a linked effective asset.

Use explicit tracking when the application needs:

- a domain-owned ID or name;
- a manually supplied source path or parent version;
- custom semantic classification;
- lineage before an instrumented run exists;
- a component that no SDK boundary exposes.

## Custom SDK Discovery

`InstrumentAssetSpec` adds the same lineage to arbitrary methods without modifying the target SDK:

```python
from autobench import InstrumentAssetSpec, SpanKind, instrument_method

handle = instrument_method(
    WorkflowClient,
    "execute",
    span="workflow_client.execute",
    span_kind=SpanKind.WORKFLOW,
    assets=[
        InstrumentAssetSpec(
            kind="prompt",
            local_id="instructions",
            value_path="kwargs.instructions",
        ),
        InstrumentAssetSpec(
            kind="tool",
            local_id="tools",
            value_path="kwargs.tools",
            many=True,
        ),
        InstrumentAssetSpec(
            kind="output_schema",
            local_id="output",
            value_path="kwargs.output_type",
        ),
    ],
)

try:
    benchmark.run()
finally:
    handle.close()
```

`value_path` traverses trusted call arguments, mappings, attributes, results, and zero-argument
accessors. Python integrations can use a typed `value_factory`. Serializable integration settings
can use `extractor_target="package.extractors:extract_assets"`; Autobench imports the callable and
passes its `InstrumentCall`. It never evaluates expression strings.

Set `many=True` when the extracted value is a sequence or mapping of independent assets. Each use
is attached to the method span. Extraction failures become run errors or instrumentation
diagnostics without replacing the SDK method's result or exception.

Run the complete offline custom SDK example:

```bash
uv run python examples/automatic_assets/custom_sdk_discovery.py \
  --record /tmp/autobench-custom-assets
```

## Identity, Aliases, And Cross-Layer Correlation

Autobench resolves identity conservatively:

1. the exact explicitly tracked Python target;
2. an explicit asset ID;
3. the target's stable module and qualified name;
4. a previously registered source locator or alias;
5. the SDK, scope, kind, and local ID locator.

Content equality alone does not merge implementations. Two tools can expose the same schema while
running different code. When one framework definition has a unique matching client projection,
Autobench links source and effective forms. Multiple possible definitions produce an
`asset_correlation_ambiguous` diagnostic instead of a silent merge.

Repeated observations of the same asset/version/source/span are deduplicated. Observations on
different spans remain separate `AssetUse` evidence because they prove separate participation.

## Persistence And Replay

`record_experiment(...)` persists only assets referenced by the experiment:

```text
recording/
  experiment.yaml
  cases/
    <case>/<variant>/run.yaml
  assets/
    index.yaml
    <safe-asset-id>.yaml
  artifacts/
    asset-content.sqlite3
```

Each asset history is a lightweight manifest containing identity, immutable versions, parent
links, changed paths, and typed `content_ref` and `diff_ref` values. Captured snapshots and readable
diffs are stored in the experiment-local `artifacts/asset-content.sqlite3` registry; prompt, tool,
and schema bodies never appear in manifests. The content-addressed store deduplicates identical
payloads, performs indexed lookups, and updates transactionally without rewriting the complete
history. A file lock coordinates manifest updates while SQLite protects registry writes.

Resolve any historical snapshot directly:

```python
from pathlib import Path

from autobench import load_asset_content, load_asset_diff

snapshot = load_asset_content(
    Path("recording/artifacts/asset-content.sqlite3"),
    asset_id="pydantic_ai:agent:support-router:prompt:instructions",
    version="595012541db0",
)
prompt = snapshot["content"]

diff = load_asset_diff(
    Path("recording/artifacts/asset-content.sqlite3"),
    asset_id="pydantic_ai:agent:support-router:prompt:instructions",
    version="595012541db0",
    parent_version="21477c4a101a",
)
```

Every run record contains:

- `assets`: exact `AssetVersion` references used by the run;
- `asset_uses`: representation, source locator, scope, span, provenance, aliases, and source link;
- ABP references on the participating spans;
- capture and conflict diagnostics in the materialized trace.

Replay loads those values without importing Pydantic AI, OpenAI, OpenAI Agents, or the application
task module. Rescoring and report replay never mutate the original asset history.

## Compatibility And Failure Behavior

| Integration | Discovery | Representations | Default asset families |
| --- | --- | --- | --- |
| Pydantic AI `>=2.22,<2.24` | yes | definition + effective | agent, capability, prompt, tool, toolset, output schema, policy |
| OpenAI Python `>=2.52,<2.55` | yes | effective | prompt, tool, output schema |
| OpenAI Agents `>=0.19.2,<0.20` | yes | definition | agent, prompt, tool, output schema, guardrail, handoff, policy, toolset |
| HTTPX `>=0.28,<0.29` | no | none | transport evidence only |

Run `autobench instrumentation doctor` to inspect installed compatibility and declared asset
families. Unsupported versions are not patched silently.

Discovery failure is non-fatal by default. Autobench emits typed diagnostics for normalization,
correlation, callback, capture, or persistence problems while preserving the host call's return
value, exception identity, streaming lifecycle, and callback count.

## Choosing The Right Surface

Use automatic discovery for SDK-visible behavioral components. Use explicit tracking for
application-owned identity and components that never cross an SDK boundary. Use
`InstrumentAssetSpec` for a custom SDK or framework. Keep user inputs and outputs as evidence,
models as factors, and measured outcomes as metrics.

That separation makes the resulting RunRecords suitable for reporting today and controlled
candidate optimization later without coupling Autobench core to one AI framework.

---

## Protocol And Traces

Canonical page: https://vcoderun.github.io/autobench/instrumentation-and-traces/

# Instrumentation And Traces

Autobench supports four collection styles that can be mixed in one run:

1. Explicit `RunContext` and `Span` calls inside a task.
2. Lightweight method instrumentation for existing application classes.
3. Trace-envelope adapters for an external agent or workflow runtime.
4. Native Pydantic AI, pydantic-gepa, OpenAI, OpenAI Agents, and HTTPX instrumentors configured
   from Python or YAML.

OpenTelemetry is not a core dependency. The optional outbound
[OTLP exporter](otlp-export.md) can replay immutable Autobench evidence to compatible backends,
while ABP remains the canonical evidence model.

See [Native Instrumentation](native-instrumentation.md) for the typed fluent API, YAML DSL,
compatibility doctor, privacy defaults, layered traces, and provider examples.

Optimizer runs use the same ABP collector through the typed pydantic-gepa event subscription. See
[Pydantic-GEPA Instrumentation](pydantic-gepa-instrumentation.md) for keyed detached lifecycle,
candidate lineage, budgets, and Optimize Anything composition evidence.

ABP is the native collection protocol underneath these APIs. It owns signal ordering, task-local
context, capture policy, instrumentation scope, trace materialization, and compatibility
diagnostics. Instrumentors emit ABP evidence directly; they do not create OpenTelemetry spans and
then convert them back into Autobench records.

## Manual Spans

```python
from autobench import DurationMetricSpec, Semantic, SpanKind


def run_case(ctx, case):
    with ctx.span(
        "support_agent",
        kind=SpanKind.AGENT,
        input=case.input,
        duration_metric=DurationMetricSpec(
            name="agent_latency",
            semantic_type=Semantic.TIME_LATENCY,
            unit="ms",
        ),
    ) as agent:
        result = call_agent(case.input)
        agent.set_output(result)
        agent.outcome(result["ok"])
        return result
```

Span duration is calculated when the context manager closes. Nested spans preserve parent-child
relationships and retain evidence emitted before an exception.

## Span Kinds

`SpanKind` includes:

- agent
- LLM
- tool
- retriever
- parser
- workflow
- custom

Kinds are semantic selectors, not restrictions. A domain can use custom kinds and tags while
generic agentic scorers continue selecting standard spans.

## Method Instrumentation

`instrument_method` is the high-level helper for one class method. It records evidence only while
a `RunContext` is active:

```python
from autobench import InstrumentMetricSpec, Semantic, instrument_method

handle = instrument_method(
    SearchClient,
    "search",
    span="search.request",
    metrics=[
        InstrumentMetricSpec(
            name="result_count",
            semantic_type="retrieval.result_count",
            value_factory=lambda call: len(call.result),
        ),
        InstrumentMetricSpec(
            name="request_count",
            semantic_type="llm.requests",
            value_path="result.usage.requests",
        ),
    ],
)

try:
    run_benchmark()
finally:
    handle.close()
```

Instrumentation supports:

- instance, static, class, and inherited methods;
- synchronous and asynchronous calls;
- iterators and generators, including `send`, `throw`, and early close;
- asynchronous iterators and generators, including `asend`, `athrow`, and `aclose`;
- synchronous and asynchronous context managers.

The wrapper preserves the original descriptor, callable signature, return value, exception
identity, and lazy streaming behavior. A stream span ends when the stream actually completes,
fails, times out, or closes, so its duration is not merely the time required to construct an
iterator.

`value_factory` is the typed Python extraction seam. It receives an `InstrumentCall` containing
the bound instance, arguments, result, error, stream item count, and last stream item. `value_path`
is the declarative alternative for trusted attribute, mapping, and zero-argument accessor paths.
Autobench does not execute arbitrary YAML expressions.

Extraction and lifecycle callback errors are recorded as evidence or compatibility diagnostics.
They do not replace the application's result or exception.

The returned `InstrumentationHandle` is also a context manager and restores the original method on
close.

## Scoped Suppression

Instrumentation can be suppressed for the current task without changing global process state:

```python
from autobench import suppress_instrumentation

with suppress_instrumentation("search.client"):
    result = client.search("internal health check")
```

Suppression keys can identify an instrumentor or an operation family. Unrelated instrumentors stay
active, nested scopes compose, and context tokens are reset even when application code raises. An
empty `suppress_instrumentation()` scope suppresses all ABP instrumentation in the current task.

## Native Instrumentors

Reusable SDK integrations implement the `Instrumentor` contract:

```python
from autobench import (
    Compatibility,
    InstrumentationHandle,
    InstrumentationRuntime,
    InstrumentorInfo,
)
from autobench.protocol import AbstractionLayer, CaptureMechanism


class ClientInstrumentor:
    info = InstrumentorInfo(
        id="example.client",
        version="1.0.0",
        target_distribution="example-client",
        supported_versions=">=2,<3",
        mechanism=CaptureMechanism.HOOK,
        layer=AbstractionLayer.CLIENT,
        span_kinds=("client.request",),
        semantic_families=("request", "response"),
    )

    def check(self) -> Compatibility:
        return Compatibility.compatible()

    def install(self, runtime: InstrumentationRuntime) -> InstrumentationHandle:
        unsubscribe = register_native_callback(...)
        return InstrumentationHandle(unsubscribe, info=self.info)
```

Install instrumentors directly through one manager when building a custom integration:

```python
from autobench import InstrumentationManager

with InstrumentationManager() as manager:
    compatibility = manager.check(ClientInstrumentor())
    if compatibility.installable:
        manager.install(ClientInstrumentor())
        run_benchmark()
```

`InstrumentorInfo` declares stable identity, target package and version range, mechanism, layer,
semantic families, source convention, optional dependencies, and sync/async/streaming/native-hook
capabilities. `Compatibility` distinguishes compatible, degraded, unavailable, unsupported, and
conflicting installations. Missing or incompatible optional dependencies degrade only the feature
that needs them; a missing required target package prevents installation.

Installing the same instrumentor version twice increments an owner reference count instead of
installing duplicate hooks. Closing the final handle unregisters native callbacks or restores the
exact patched descriptor. Competing owners can instrument the same method independently, while an
external wrapper replacement produces a conflict diagnostic instead of being overwritten.

### External Backend Integration

An SDK that already exposes its own telemetry backend does not need a second method-patching layer.
Its Autobench adapter can install a backend that asks `InstrumentationRuntime` for a span:

```python
from collections.abc import Generator
from contextlib import contextmanager

from autobench import InstrumentationRuntime, InstrumentorInfo, Span


class AutobenchBackend:
    def __init__(self, runtime: InstrumentationRuntime, info: InstrumentorInfo) -> None:
        self.runtime = runtime
        self.info = info

    @contextmanager
    def start_span(
        self,
        name: str,
        attributes: dict[str, bool | str | int | float],
    ) -> Generator[Span | None, None, None]:
        span = self.runtime.span(
            self.info,
            name,
            kind="workflow",
            attributes=attributes,
            target_version="1.4.0",
            suppression_keys=("example-sdk",),
        )
        if span is None:
            yield None
            return
        with span:
            yield span
```

`runtime.span()` returns an unentered span only during an active Autobench run. It returns `None`
outside a run, in a matching suppression scope, or when the active protocol context is not owned by
a `RunContext`. Entering the span makes it the parent of nested built-in instrumentation, including
Pydantic AI, OpenAI, and HTTPX calls. Exiting it records normal completion, failure, timeout, or
cancellation without changing the host SDK result or exception.

External counters and histograms can use the matching observation seam:

```python
runtime.metric(
    info,
    "framework.jobs",
    1,
    semantic_type="workflow.jobs",
    unit="job",
    suppression_keys=("example-sdk",),
)
```

The observation is attached to the currently entered Autobench span when one exists and is marked
with `source=instrumentation`. Like `runtime.span()`, this method returns `None` outside an owned,
unsuppressed benchmark context.

Backends that enrich a span owned by another installed instrumentor can use the active-span view:

```python
current = runtime.current_span(info, suppression_keys=("example-sdk",))
if current is not None:
    current.set_attribute("framework.profile", "reviewer")
    current.set_usage("input_tokens", 420)
    current.event("framework cache hit", semantic_type="event.occurrence")
```

`CurrentSpan` can read and set attributes, set usage, attach instrumentation events, and record an
exception. It never renames the ABP operation after `span_start`; an external backend that supports
display-name updates should store that value as an attribute. Calls return `None` outside an active
legacy-backed span, including a raw ABP context that is not owned by a `RunContext`.

`runtime.is_installed(instrumentor_id)` and `runtime.installed_ids` let a composite backend choose
the native owner dynamically. For example, a framework adapter can suppress its model/tool spans
only while `autobench.pydantic_ai` is actually installed, avoiding both missing evidence and double
accounting.

The adapter remains responsible for its host's backend ownership. If a backend is already installed,
compose or multiplex it with the Autobench backend; do not replace it silently. The close callback
must restore the previous backend only when the adapter still owns the active slot. This preserves
an existing OTel, Logfire, or application backend and avoids clobbering a newer installation.

Mechanisms should be selected in this order:

1. stable native processor or callback;
2. stable native wrapper/decorator extension point;
3. public method patch;
4. explicitly version-pinned private method patch;
5. unsupported with a compatibility diagnostic.

Application benchmarks normally use the higher-level lifecycle owner instead:

```python
from autobench import Benchmark, HTTPXInstrumentation, OpenAIInstrumentation

benchmark = Benchmark("chat").instrument(
    OpenAIInstrumentation(),
    HTTPXInstrumentation(),
)
result = benchmark.run()
```

`Benchmark.instrument(...)` installs configured and custom instrumentors before any matrix item,
keeps them active through concurrent runs and streams, and closes them after execution.

## Trace Envelopes

Adapters can normalize a completed external trace into `TraceEnvelope`:

```python
from autobench import TraceEnvelope, attach_trace

trace = TraceEnvelope(
    trace_id="trace-42",
    name="checkout-agent",
    input={"cart_id": "c1"},
    output={"status": "complete"},
    spans=tuple(converted_spans),
    attributes={"framework": "custom-agent-runtime"},
)

attach_trace(ctx, trace)
```

`attach_trace` preserves spans and errors and projects known usage, model, provider, duration, and
outcome fields into semantic observations. Large native trace payloads should be written as an
artifact and referenced by `raw_artifact`.

## Pydantic AI Usage

Install the optional native instrumentor when the application uses Pydantic AI:

```bash
pip install 'autobench[pydantic-ai]'
```

The integration uses Pydantic AI's public capability hooks and only injects its
capability while an Autobench run is active:

```python
from autobench import Benchmark, PydanticAIInstrumentation

experiment = benchmark.instrument(PydanticAIInstrumentation()).run()
```

No manual span or metric calls are required. The instrumentor captures:

- agent runs and streamed execution;
- model requests, requested and response model identities, providers, and direct usage;
- tool argument validation, execution, retry, failure, approval, and deferred control flow;
- structured-output validation;
- first-chunk latency, partial streams, failures, and normal completion;
- tracked prompt, tool, and output-schema versions;
- multimodal metadata, with binary references only when the capture policy requests full content.

The instrumentor composes with user event handlers and Pydantic AI's own
`Instrumentation` capability. It does not configure, replace, or require
OpenTelemetry. Autobench supports the audited public integration seam across
Pydantic AI 2.22.x and 2.23.x; `InstrumentationManager.check()` reports incompatible
versions before installing hooks.

Application outputs and exceptions are passed through unchanged. Aggregate agent
usage and direct model usage retain distinct accounting scopes, and cost remains a
downstream derivation. Replaying recorded ABP evidence does not require Pydantic AI
to be installed.

See the [live Pydantic AI example](examples.md#pydantic-ai-live-layered-instrumentation) for a tool-using,
structured-output, streaming benchmark with a retry path.

### Usage Bridge

Pydantic AI usage can be normalized without importing Pydantic AI into core:

```python
from autobench import PydanticAIUsage, record_pydantic_ai_usage

record_pydantic_ai_usage(
    ctx,
    PydanticAIUsage(
        requests=1,
        input_tokens=420,
        output_tokens=83,
        model_name="gemini-3-flash-preview",
        provider="openrouter",
    ),
)
```

The bridge emits canonical LLM token, model, and provider observations that pricing derivation and
reports can consume.

## Trace Extraction And Accounting

Instrumentors record immutable facts. Extractors turn a completed ABP trace into semantic
observations without mutating that trace:

```python
from autobench import (
    CompositeExtractor,
    SignalExtractor,
    SpanExtractor,
    UsageExtractor,
    replay_extraction,
)

extractor = CompositeExtractor(
    SignalExtractor(),
    SpanExtractor(),
    UsageExtractor(),
)
derived_record = replay_extraction(record, extractor)
```

The extractors have separate ownership:

- `SignalExtractor` reconstructs measurements and events and preserves their accounting scope,
  abstraction layer, logical operation ID, and instrumentor identity.
- `SpanExtractor` derives generic operation counts, direct durations, maximum depth and fan-out,
  critical-path makespan, parallelism, incomplete work, retry/recovery, validation, approval,
  tool-call, message-growth, and reference evidence.
- `UsageExtractor` derives LLM request, token, requested-model, response-model, and provider
  evidence. It never derives cost.

Every extractor has a stable name and version. Replay records both in extraction evidence and
RunRecord lineage. Replaying a newer version replaces observations owned by the older version in
the derived record; the parent record remains unchanged.

### Direct And Aggregate Evidence

ABP keeps all raw measurements but prevents framework/client nesting from inflating totals:

1. Aggregate parent measurements are never added to direct child measurements.
2. Usage totals select one abstraction layer per semantic, preferring client evidence before
   framework, application, and transport evidence.
3. Equivalent direct operations with a shared logical operation ID are counted once.
4. Equal equivalent values are deduplicated. Conflicting values require a unique explicit
   authority; unresolved conflicts produce `ambiguous_direct_measurement` and are excluded from
   the derived total.
5. Aggregate values are retained as validation evidence. A disagreement with the direct total
   produces `aggregate_measurement_mismatch`.
6. Requested and response model identities remain separate factors.

Reports and `ObservationQuery.first_exact()` prefer an accounting-safe aggregate summary over
same-source per-operation direct evidence. Raw and projected queries can still inspect every
underlying observation.

Graph timing uses monotonic timestamps only. `time.critical_path` is the observed trace makespan,
and `operation.parallelism` is completed leaf work divided by that makespan. Invalid or partial
clock evidence is retained through diagnostics rather than repaired with wall-clock subtraction.

## Adapter Boundary

Core instrumentation intentionally does not know Pydantic AI, OpenAI Agents, LangChain, DSPy, or
OpenTelemetry internals. An integration should:

1. Collect from the framework's stable hooks.
2. Convert native calls or traces into Autobench spans and observations.
3. Store large raw payloads as artifacts.
4. Keep native dependencies optional.

This boundary lets applications use existing instrumentation while RunRecords remain portable.

---

## Native Instrumentation

Canonical page: https://vcoderun.github.io/autobench/native-instrumentation/

# Native Instrumentation

Autobench native instrumentors collect ABP traces from supported SDKs without task-level
`ctx.span()` or `ctx.metric()` calls. They are optional adapters around public hooks or pinned,
reviewed patch points. Core benchmark, record, replay, and report imports do not require any of the
instrumented SDKs.

## Install

Install one integration or the complete set:

```bash
pip install 'autobench[pydantic-ai]'
pip install 'autobench[openai]'
pip install 'autobench[openai-agents]'
pip install 'autobench[httpx]'
pip install 'autobench[pydantic-gepa]'
pip install 'autobench[instrumentation]'
```

The integration registry is lazy. Loading a YAML spec, replaying evidence, or running
`autobench instrumentation doctor` does not import an SDK that is not installed.

## Automatic Discovery

Use `instrument_all()` when the benchmark should activate every built-in integration that is
installed and compatible in the current environment:

```python
from autobench import Benchmark

benchmark = Benchmark("support-agent").instrument_all()
```

Semantic instrumentors also discover SDK-visible behavioral assets by default. The application does
not need tracking decorators for prompts, tools, output schemas, capabilities, agents, guardrails,
handoffs, or policies already visible at those boundaries. See
[Automatic Asset Discovery](automatic-asset-discovery.md) for identity, source/effective links,
privacy, persistence, and custom SDK extraction.

Unavailable or unsupported integrations are skipped by default and recorded on each run as
`instrumentation.skipped` diagnostic evidence. Use `strict=True` when the environment must support
the complete selected set:

```python
benchmark = Benchmark("support-agent").instrument_all(
    exclude={"httpx"},
    strict=True,
)
```

Explicit settings take precedence over discovery, including an explicit `false`. A custom runtime
instrumentor with the same instrumentor ID also takes precedence, so automatic discovery does not
install a duplicate. Calling `instrument_all()` again replaces the previous automatic settings.

### Live OpenRouter Trace

The live Pydantic AI example exercises automatic discovery across all three active layers:

```bash
uv sync --extra instrumentation
export OPENROUTER_API_KEY=...
export OPENROUTER_MODEL=openrouter:openai/gpt-5.6-luna
uv run python examples/pydantic_ai/openrouter_instrument_all.py \
  --record /tmp/autobench-openrouter
```

The benchmark itself only opts in once:

```python
benchmark = Benchmark("openrouter-shopping-agent").instrument_all()
```

The example uses plain instructions, a plain tool function, and an undecorated Pydantic output type.
Automatic discovery adds Pydantic AI, OpenAI client, and HTTPX transport instrumentors. One real
request therefore records agent, model, tool, output
validation, stream, client request, and transport spans with their native parentage. It also records
model identity, token usage, durations, HTTP method/host/path/status, score observations, asset
versions, capture diagnostics, and replayable source provenance. HTTP bodies and credentials remain
redacted by the default capture policy.

The full source is `examples/pydantic_ai/openrouter_instrument_all.py`. It deliberately contains no
manual `ctx.span()` or `ctx.metric()` calls so the resulting record demonstrates native collection
rather than hand-authored benchmark telemetry.

## YAML

Instrumentation belongs to the named benchmark:

```yaml
# yaml-language-server: $schema=schemas/0.3.0/benchmark_schema.json
benchmark:
  support-agent:
    dataset:
      source: file://datasets/cases.yaml
    run:
      python: support_benchmark:run
    variants:
      baseline:
        factors:
          model.name: openrouter:openai/gpt-5.6-luna
    instrumentation:
      all:
        exclude: [httpx]
        strict: false
        assets:
          representations: [definition, effective]
          include: [prompt, tool, output_schema, capability]
      pydantic_ai: {}
      openai: {}
      httpx:
        capture:
          path: hash
          request_headers: [x-request-id]
          response_headers: [x-request-id]
          request_body: false
          response_body: false
          max_body_bytes: 65536
```

Use `false` to retain a known integration in a shared spec without installing it:

```yaml
instrumentation:
  openai_agents: false
```

Unknown integration names and unknown settings fail validation. YAML never evaluates Python
expressions.

The `all` block follows the same precedence rules as the Python builder. In this example HTTPX is
excluded from discovery but its explicit capture settings still install it; all explicit entries
remain authoritative.

## Python

The fluent API accepts typed, serializable settings:

```python
from autobench import (
    Benchmark,
    HTTPXCaptureSettings,
    HTTPXInstrumentation,
    OpenAIInstrumentation,
)

benchmark = Benchmark("streaming-chat").instrument(
    OpenAIInstrumentation(),
    HTTPXInstrumentation(
        capture=HTTPXCaptureSettings(
            path="hash",
            response_headers=("x-request-id",),
        )
    ),
)
experiment = benchmark.run()
```

It also accepts a custom `Instrumentor` instance. Runtime instances are installed for the whole
benchmark matrix and closed even when execution fails. They are intentionally not serialized into
the YAML spec:

```python
benchmark.instrument(MyNativeInstrumentor(settings))
```

Duplicate instrumentor IDs are rejected before hooks are installed. This avoids ambiguous
ownership when a typed setting and a custom instance configure the same integration.

### SDK-Owned Telemetry Backends

Some frameworks already route lifecycle spans through a process-wide or context-local backend.
Implement the Autobench integration against that stable backend contract instead of patching every
framework method. `InstrumentationRuntime.span()` creates an unentered ABP span in the active
benchmark run and returns `None` when capture is inactive or suppressed:

```python
span = runtime.span(
    instrumentor.info,
    "framework.workflow",
    kind="workflow",
    attributes={"framework.operation": "run"},
    target_version=installed_framework_version,
    suppression_keys=("framework",),
)

if span is None:
    return call_subject()

with span:
    return call_subject()
```

Use `runtime.metric()` for host counters and histograms that should become semantic observations.
It follows the same active-context and suppression rules and attaches the observation to the
currently entered ABP span.

Use `runtime.current_span()` when the host enriches a span owned by a nested native instrumentor.
The returned `CurrentSpan` updates attributes, usage, events, and errors on that active ABP span.
Check `runtime.is_installed()` before claiming native ownership so disabling a child instrumentor
does not silently remove model or tool evidence.

The entered span becomes the parent of nested Autobench instrumentors automatically. This is the
preferred integration for an agent runtime whose workflow span should contain native Pydantic AI
agent/model/tool spans. Keep provider usage on the native provider or framework span; a higher-level
workflow span should not copy the same token totals.

Backend installation must preserve existing observers:

1. read the current host backend;
2. install a host-owned composite containing the previous backend and the ABP backend;
3. retain the exact installed composite identity;
4. on close, restore the previous backend only if the composite is still current.

This policy allows existing OTel or Logfire telemetry and Autobench evidence to coexist. Autobench
does not own the host backend protocol and core does not import the external framework.

## Built-In Integrations

| Integration | Layer | Collection seam | Evidence |
| --- | --- | --- | --- |
| Pydantic AI | framework | public agent capability | spans plus agent/capability/prompt/tool/toolset/output-schema lineage |
| Pydantic-GEPA | optimizer | typed event subscription | optimization/composition/engine/evaluation spans, budgets, candidate lineage, and component asset versions |
| OpenAI Python | client | reviewed public client methods and stream types | spans plus effective prompt/tool/output-schema lineage |
| OpenAI Agents | framework | native trace processor and public Runner surface | spans plus agent/prompt/tool/output-schema/guardrail/handoff/policy lineage |
| HTTPX | transport | public transport methods | request method/host/path policy, status, selected headers, body metadata, stream lifecycle |

The pydantic-gepa integration has its own complete guide, including Optimize Anything composition,
detail modes, semantic accounting, assets, replay, and reports: [Pydantic-GEPA
Instrumentation](pydantic-gepa-instrumentation.md).

Run compatibility diagnostics before a benchmark:

```bash
autobench instrumentation doctor
```

The Rich output shows availability, installed version, supported range, abstraction layer,
mechanism, sync/async/streaming capabilities, span kinds, semantic families, capture defaults, and
degradation diagnostics.

## Layered Traces

Instrumentors compose instead of flattening one another. A Pydantic AI request using the OpenAI
client over HTTPX can produce this parent chain:

```text
task
  agent
    llm framework operation
      OpenAI client operation
        HTTP request
```

Transport spans do not emit token or cost usage. Framework aggregate usage and client direct usage
retain different accounting scopes. Trace extraction selects one authoritative direct layer and
keeps aggregate values as validation evidence, so enabling HTTPX cannot inflate LLM totals.

## Streaming Lifecycle

A stream span does not end when an SDK returns an iterator. It remains open until the stream:

- completes normally;
- raises;
- is cancelled;
- is explicitly closed early;
- is abandoned when the instrumentor manager closes.

ABP records first-chunk evidence, item/chunk counts, partial state, and the final end reason. Native
items, exceptions, iterator methods, and context-manager behavior pass through unchanged.

## HTTP Privacy Defaults

HTTPX capture defaults are deliberately conservative:

- query-free path hash, not the raw path;
- no request or response headers unless named;
- authorization, cookies, API keys, tokens, passwords, and secrets always redacted;
- no request or response body capture;
- bounded capture when bodies are explicitly enabled;
- binary bodies represented by metadata and a digest, not embedded bytes.

`path: full` is an explicit opt-in. Query strings and URL user information are not recorded by the
path setting. Capture policies apply before evidence reaches a RunRecord.

## Trace Diagnostics

Every native span records an `InstrumentationScope`: instrumentor and target versions, mechanism,
abstraction layer, and source convention. Source facts can be retained alongside canonical
Autobench semantic attributes. Unsupported library versions fail installation instead of silently
patching an unknown lifecycle.

Inspect recorded trace shape without importing task modules or optional SDKs:

```bash
autobench instrumentation trace runs/support-agent/exp_...
```

The command reports per-case span/root counts, partial traces, diagnostics, span-kind totals, and
instrumentor composition.

## Replay Without SDKs

RunRecords contain materialized ABP traces, not live provider objects. A reporting or optimization
worker can replay and re-extract evidence with only Autobench installed:

```python
from autobench import CompositeExtractor, SignalExtractor, SpanExtractor, UsageExtractor
from autobench.records.replay import load_run_record, replay_extraction

record = load_run_record(path, root_dir=run_dir)
derived = replay_extraction(
    record,
    CompositeExtractor(SignalExtractor(), SpanExtractor(), UsageExtractor()),
)
```

Extraction creates a derived record with lineage; it never mutates the original record.

## ABP And OpenTelemetry

ABP is not an OpenTelemetry wrapper and has no OTel dependency. It is Autobench's evidence protocol
for benchmark execution, semantic measurements, accounting scope, partial streams, replay, and
optimization lineage. Native instrumentors use the same kinds of stable SDK hooks that mature OTel
instrumentations validate, but emit ABP directly.

The optional [OTLP exporter](otlp-export.md) can replay immutable ABP evidence to systems such as
Logfire, Datadog, or a vendor-neutral collector. It remains an outbound adapter: ABP is the source
evidence model, and importing the base Autobench package does not require an OTel SDK or collector.

## Protocol Stability

ABP protocol version `1` is the initial public serialized contract. Autobench `0.2.x` preserves the
meaning of its signal, trace, scope, provenance, and accounting fields. Readers retain unknown
additive data through extension maps, while a breaking wire-format change requires a new protocol
version. Instrumentor patch points are compatibility-gated separately because provider SDK
lifecycles can change independently of ABP.

## Examples

- `examples/abp_manual`: explicit workflow spans plus method instrumentation.
- `examples/abp_concurrent`: concurrent sibling operations with task-local parentage.
- `examples/pydantic_ai`: tool use, retry, streaming, and structured output; OpenAI models add
  OpenAI and HTTPX layers.
- `examples/abp_openai`: offline official OpenAI streaming over an HTTPX mock transport.
- `examples/abp_openai_agents`: offline native OpenAI Agents trace-processor workflow.
- `examples/abp_replay`: trace extraction from recorded evidence without importing provider SDKs.

---

## Pydantic-GEPA Instrumentation

Canonical page: https://vcoderun.github.io/autobench/pydantic-gepa-instrumentation/

# Pydantic-GEPA Instrumentation

Autobench can record a pydantic-gepa optimization as structured experiment evidence without
handwritten optimizer spans, metric calls, asset decorators, or an Autobench-specific observer.
The optimization continues to use the normal pydantic-gepa API. Autobench subscribes to its typed
event contract only while a benchmark run is active.

This integration preserves:

- optimization, composition, engine, stage, iteration, reflection, evaluation, candidate, and
  final-rescore lifecycles;
- train, validation, and test dataset declarations;
- objective identity, direction, role, semantic type, and unit;
- evaluation-call and optimizer-cost limits, usage, and remaining budget;
- every candidate lifecycle, parent relationship, component value, and tracked asset version;
- BestOf, Vote, AdaptiveSequential, Sequential, Parallel, Single, and Pipeline structure;
- selection method, all contenders, their scores, and the selected engine execution;
- errors, cancellation, checkpoints, diagnostics, and partial evidence;
- a compact, versioned projection that remains readable after replay without pydantic-gepa.

Autobench does not run GEPA itself, replace pydantic-gepa checkpoints, or copy arbitrary live
optimizer state into a record.

## Install

Install the dedicated extra when only this integration is needed:

```bash
uv add 'autobench[pydantic-gepa]'
```

The combined native instrumentation environment also includes it on supported Python versions:

```bash
uv add 'autobench[instrumentation]'
```

Check the installed event contract before a benchmark:

```bash
autobench instrumentation doctor
```

The pydantic-gepa instrumentor requires event contract version `1`. An unavailable or unsupported
package is reported independently; base Autobench, record replay, and report rendering continue to
work without importing pydantic-gepa.

## Run The Offline Examples

The repository includes standard GEPA, Optimize Anything Omni, multi-component, and checkpoint
resume benchmarks. They use deterministic local tasks and real optimizer runtimes.

```bash
uv run autobench run examples/pydantic_gepa/autobench.yaml \
  --record /tmp/autobench-pydantic-gepa

uv run autobench instrumentation trace /tmp/autobench-pydantic-gepa
uv run autobench report /tmp/autobench-pydantic-gepa
uv run autobench replay /tmp/autobench-pydantic-gepa
```

Run the remaining contracts with `standard.yaml`, `multi_component.yaml`, and `resume.yaml`. Each
spec records, replays, reports, and exports independently in `make examples`.

The task in `examples/pydantic_gepa/optimizer_benchmark.py` contains no Autobench span or metric
calls. Its relevant shape is:

```python
from pydantic_gepa import Optimization
from pydantic_gepa.experimental.optimize_anything import (
    BestOf,
    OptimizeAnythingConfig,
    Pipeline,
    Single,
)


def run(ctx, case):
    del ctx
    optimization = Optimization.from_examples(...)
    return optimization.optimize(
        config=OptimizeAnythingConfig(
            composition=Pipeline(
                steps=(
                    BestOf(engines=(weak_engine, strong_engine)),
                    Single(engine=continuation_engine),
                )
            )
        )
    )
```

Instrumentation is declared once in YAML:

```yaml
instrumentation:
  pydantic_gepa:
    detail: full
    assets:
      discover: true
      representations: [definition, effective]
      include: [prompt]
```

## Python Configuration

Use the typed configuration with a fluent benchmark:

```python
from autobench import (
    AssetDiscoverySettings,
    Benchmark,
    PydanticGEPAInstrumentation,
)

benchmark = Benchmark("optimizer-evaluation").instrument(
    PydanticGEPAInstrumentation(
        detail="evaluations",
        assets=AssetDiscoverySettings(include=("prompt", "tool_schema")),
    )
)
```

Use automatic discovery when every compatible installed SDK integration should compose:

```python
benchmark = Benchmark("optimizer-evaluation").instrument_all()
```

`instrument_all()` installs the pydantic-gepa observer once. Explicit `pydantic_gepa` settings
override automatic selection, including an explicit disabled state.

## Detail Modes

`detail` controls high-cardinality spans, not durable summary correctness.

| Mode | Always retained | Additional evidence |
| --- | --- | --- |
| `summary` | optimization, composition/engine lifecycle, objective, budgets, selections, terminal result, projection | no per-case, candidate, iteration, or reflection spans |
| `evaluations` | everything in `summary` | candidate, evaluation, case, and metric spans and observations |
| `full` | everything in `evaluations` | GEPA iterations, reflection/proposal detail, Pareto and backend progress events |

All modes preserve the typed optimization projection and candidate summaries. Select `summary` for
large production optimization jobs, `evaluations` for score debugging, and `full` when reflection
and proposal behavior matters.

## ABP Trace Shape

A full Pipeline can produce this hierarchy:

```text
task
  pydantic_gepa.optimization                 kind=optimization
    pydantic_gepa.composition_step           kind=workflow
      pydantic_gepa.engine                   kind=workflow
      pydantic_gepa.engine                   kind=workflow
      pydantic_gepa.candidate                kind=candidate
      pydantic_gepa.evaluation               kind=evaluation
        pydantic_gepa.case                   kind=evaluation
          pydantic_gepa.metric               kind=scorer
    pydantic_gepa.composition_step           kind=workflow
      pydantic_gepa.engine                   kind=workflow
    pydantic_gepa.final_rescore              kind=evaluation
```

GEPA-backed engines may add iteration and reflection spans. AutoResearch, Meta-Harness, Best-of-N,
and custom engines still produce the common engine/evaluation/budget/selection lifecycle even when
they do not expose GEPA-specific reflection callbacks.

Parallel engine events can arrive on worker threads. pydantic-gepa assigns stable pipeline, step,
branch, engine-execution, candidate, and parent IDs before dispatch. Autobench correlates from those
IDs and reuses the run context captured at optimization start; callback arrival order is not used
as parentage.

## Metrics And Accounting

The integration classifies optimizer evidence with semantic observations:

| Evidence | Semantic type |
| --- | --- |
| raw metric and candidate/final score | `evaluation.score` or the metric's declared semantic type |
| candidate status | `evaluation.label` |
| evaluator feedback | `evaluation.explanation` |
| calls used, limit, remaining | `optimization.evaluations.used`, `.limit`, `.remaining` |
| optimizer cost used, limit, remaining | `optimization.optimizer_cost.used`, `.limit`, `.remaining` |
| evaluator cost used | `optimization.evaluation_cost.used` |
| optimizer plus evaluator aggregate cost | `optimization.cost.used` |

Objective scores retain their direction and role. A minimizing domain metric is not silently
replaced by an optimizer's transformed selection scalar.

Native Pydantic AI, OpenAI, OpenAI Agents, and HTTPX instrumentors remain authoritative for direct
model, token, request, transport, and serving-cost evidence. The pydantic-gepa layer records
optimizer semantics and does not duplicate child SDK token totals. This setup is valid:

```yaml
instrumentation:
  pydantic_gepa:
    detail: full
  pydantic_ai: {}
  openai: {}
  httpx: {}
```

Native model calls retain their Agent -> provider -> HTTP transport hierarchy in the same benchmark
trace as the optimizer lifecycle. Accounting stays separated by source layer: pydantic-gepa owns
optimization evidence, while the native SDK instrumentors own direct model, token, request, and
transport evidence.

## Candidate Assets And Lineage

pydantic-gepa declares optimization components independently of Autobench. The native adapter maps
those components to logical tracked assets:

| Component kind | Tracked asset family |
| --- | --- |
| `instructions`, `system_prompt` | prompt/instruction asset |
| `tool_schema` | tool asset |
| `input_schema`, `output_schema` | schema asset |
| `field_description`, `schema_description` | schema-description asset |
| custom semantic component | optimization component with its declared semantic type |

Definition assets describe the initial component. Effective assets describe each candidate value.
Candidate IDs remain optimizer lineage IDs; asset versions remain content/version identities. They
are linked, not conflated.

With full asset capture, content is stored in the experiment's asset-content registry rather than
duplicated in every run YAML. Run evidence stores `asset_id@version`, parent version, use records,
and diffs. Capture policy still controls raw candidate text, evaluator output, feedback, and trace
payloads.

## Durable Projection

Detailed ABP signals remain the source of truth. Each run also contains a replay-friendly extension
under:

```text
autobench.pydantic_gepa/v1
```

Load it through the public typed model:

```python
from autobench import PydanticGEPAEvidence, replay_experiment

experiment = replay_experiment("/tmp/autobench-pydantic-gepa")
payload = experiment.runs[0].extensions["autobench.pydantic_gepa/v1"]
evidence = PydanticGEPAEvidence.model_validate(payload)

execution = evidence.executions[0]
print(execution.final_score)
print(execution.candidates)
print(execution.engines)
print(execution.selections)
```

The projection contains execution identity, backend/engine/composition, datasets, objective,
budgets, candidate lifecycle and component versions, engine summaries, selections, checkpoints,
stop reason, event count, and diagnostic count. Unknown or invalid future projection versions
produce report warnings rather than breaking normal record replay.

## Reports, Replay, And Export

Rich reports add dedicated sections for:

- optimization outcome and resources;
- engine executions and branch identity;
- candidate lifecycle and parent lineage;
- component asset versions;
- selection method, all contenders, winner, score, and reason;
- partial status or diagnostics.

Replay and export operate from the immutable record and do not rerun an optimizer:

```bash
autobench replay /tmp/autobench-pydantic-gepa
autobench report /tmp/autobench-pydantic-gepa
autobench export /tmp/autobench-pydantic-gepa \
  --format yaml \
  --path /tmp/pydantic-gepa-report.yaml
```

YAML and Markdown report views include the same typed optimization projection. CSV remains the
flat run-metric view.

## Failure And Cancellation

Instrumentation observer failures are isolated and become bounded diagnostics. They do not alter
the optimizer result, exception, cancellation, or checkpoint behavior.

- completed optimizations close normalized-but-unselected candidate spans normally;
- explicit candidate accepted/rejected events retain their lifecycle labels;
- a failed optimization closes the root with error evidence;
- cancellation closes the root as cancelled and marks unfinished operations partial;
- genuinely unmatched starts/ends produce ABP diagnostics;
- duplicate event delivery is bounded and cannot overwrite another run's state;
- calls outside an active Autobench run are no-ops.

## Ownership Boundary

pydantic-gepa owns typed events, upstream GEPA compatibility, engine/composition correlation,
candidate values, and normalized evaluation evidence. Autobench owns ABP conversion, semantic
classification, capture, tracked assets, immutable records, replay, reports, and exports.

Autoptimize can later consume these records for planning and promotion. This integration does not
make causal claims, choose the next optimization strategy, or promote a candidate.

---

## Compatibility Contract

Canonical page: https://vcoderun.github.io/autobench/abp-compatibility/

# ABP Compatibility Contract

This page freezes the observable behavior preserved while the Autobench
Instrumentation Protocol (ABP) replaces legacy span and instrumentation
internals. Phases 1 through 8 now satisfy this contract. The complete public
instrumentation guide is in [Instrumentation And Traces](instrumentation-and-traces.md).

## Compatibility Boundary

The following top-level imports remain available while ABP is introduced:

```python
from autobench import (
    ArtifactRef,
    AssetVersion,
    DurationMetricSpec,
    ErrorRecord,
    InstrumentationHandle,
    InstrumentCall,
    InstrumentFactorSpec,
    InstrumentMetricSpec,
    Observation,
    RunContext,
    RunRecord,
    Span,
    SpanKind,
    SpanRecord,
    TraceEnvelope,
    attach_trace,
    get_active_run_context,
    instrument_method,
    trace_to_observations,
)
```

ABP may move implementations into new packages, but these imports and their
current behavior remain compatibility facades until a separately announced
deprecation cycle.

### Manual spans

Existing manual spans preserve these guarantees:

- `ctx.span(...)` is a synchronous context manager;
- nested spans receive the active span as `parent_id`;
- every completed span has UTC start/end timestamps and a non-negative
  monotonic duration;
- a configured duration metric is linked to the span;
- metrics, factors, events, errors, and artifacts retain their span link;
- exceptions are recorded and then propagated;
- `Span.set_output`, `Span.set_attribute`, and `Span.set_usage` continue to
  update the recorded span;
- entering instrumentation without an active run context remains a no-op;
- closing the final `InstrumentationHandle` restores the original descriptor.

The tests in `tests/test_abp_compatibility.py` are the executable form of this
contract.

### Stored evidence

Legacy model-shaped RunRecord and TraceEnvelope YAML remains loadable. ABP
will add protocol data additively and preserve the existing `spans` input
during migration. Replay must not require the task module or an optional
instrumented SDK.

The frozen legacy examples are:

- `tests/fixtures/abp/legacy_run_record.yaml`
- `tests/fixtures/abp/legacy_trace_envelope.yaml`

## Concurrency Regression Contract

`RunContext` now uses task-local ABP context. The concurrency migration is
covered by passing regression tests for all of these cases:

1. concurrent sibling spans under one parent both point to that parent;
2. a nested task inherits the parent active at task creation;
3. completing one sibling does not change the other sibling's active parent;
4. out-of-order completion does not corrupt later parent selection;
5. cancellation closes only the cancelled branch and restores its context;
6. separate RunContexts never share active spans.

The old mutable-stack reproduction is retained only in design history; it is
not the current runtime behavior.

## Canonical Trace Decision

ABP will have one canonical immutable `Trace` model. `TraceEnvelope` does not
have behavior that justifies a second trace representation, so it will become
a compatibility name for `Trace` rather than a parallel model. Existing
`TraceEnvelope(...)`, `attach_trace(...)`, and `trace_to_observations(...)`
callers continue to work.

This avoids conversion drift between manually attached traces and traces
materialized from native ABP signals.

## Package Shape

ABP code is introduced only when its phase needs it. Empty placeholder modules
are not created.

```text
autobench/
  protocol/
    ids.py
    values.py
    signals.py
    traces.py
    context.py
    capture.py
    collector.py
  instrumentation/
    models.py
    manager.py
    patching.py
    streaming.py
    pydantic_ai.py
    openai.py
    openai_agents.py
    httpx.py
```

Small modules are combined when separation would only create navigation cost.
Existing unrelated modules are not moved as part of ABP.

## Optional Integration Targets

The initial integration extras are reserved as follows:

| Extra | Research baseline | First implementation phase |
| --- | ---: | ---: |
| `autobench[pydantic-ai]` | Pydantic AI 2.22.0 and 2.23.0 | 10 |
| `autobench[openai]` | OpenAI Python 2.52.0 through 2.54.0 | 11 |
| `autobench[openai-agents]` | OpenAI Agents 0.19.2 | 11 |
| `autobench[httpx]` | HTTPX 0.28.1 | 12 |
| `autobench[instrumentation]` | all integrations above | 13 |

These versions are the public-API research baseline captured on 2026-08-03,
not a compatibility claim. Dependency metadata is added only when each native
instrumentor and its version matrix exist. Autobench core remains free of
these dependencies.

## Manual Span Performance Baseline

The baseline measures a minimal completed manual span with no observations,
artifacts, or errors. Each repeat creates one RunContext and records 10,000
spans. Timing uses `timeit.repeat`; duration comes from the host monotonic
clock. The benchmark does not enforce a CI latency threshold because shared CI
timing is not stable.

Reproduce it with:

```bash
uv run python scripts/benchmark_spans.py --iterations 10000 --repeats 7
```

Baseline captured before ABP runtime changes:

| Field | Value |
| --- | ---: |
| Date | 2026-08-03 |
| Python | 3.11.13 |
| Platform | macOS 26.1 arm64 |
| Minimum | 3,227.4 ns/span |
| Median | 3,324.8 ns/span |

The raw per-repeat values were `3467.0`, `3308.4`, `3299.3`, `3394.4`,
`3394.3`, `3324.8`, and `3227.4` ns/span. Later phases compare using the same
script and workload; they do not compare unrelated machine results.

Release measurements captured after ABP materialization on the same host:

| Workload | Measurement |
| --- | ---: |
| Manual ABP span | 28,836.7 ns/span median |
| HTTPX baseline request | 35,346.2 ns/request median |
| Instrumented HTTPX request | 224,263.6 ns/request median |
| HTTPX instrumentation overhead | 188,917.4 ns/request median |
| 10,000 x 32-byte HTTP stream | 2,618.0 ns/chunk median |
| Long-stream peak allocation | 30,918 bytes |

The long-stream result is about `3.1` peak allocated bytes per emitted chunk, which confirms the
instrumentor does not retain chunk payloads as the stream grows. These numbers characterize this
machine and are not release thresholds. Reproduce transport and stream measurements with:

```bash
uv run python scripts/benchmark_spans.py --httpx --iterations 1000 --repeats 7
uv run python scripts/benchmark_spans.py --httpx-stream --chunks 10000 --chunk-size 32 --repeats 7
```

## Outbound OTLP Compatibility

The optional ABP-to-OTLP adapter is a projection of immutable evidence, not part of the ABP wire
format. OTel SDK/exporter version changes therefore do not change `PROTOCOL_VERSION`. The adapter
must preserve Autobench identities, partial state, source provenance, and semantic event payloads
while leaving the input records byte-for-byte unchanged. The base package and replay remain usable
without the `otlp` extra.

---

## OTLP Export

Canonical page: https://vcoderun.github.io/autobench/otlp-export/

# OTLP Export

Autobench can replay an immutable experiment record into OTLP HTTP/protobuf traces for systems
such as Logfire, Datadog, or a vendor-neutral OpenTelemetry Collector. This is an outbound adapter,
not Autobench's collection protocol:

```text
application -> ABP instrumentation -> immutable Autobench record -> optional OTLP export
```

ABP remains the complete source of truth for semantic observations, benchmark identity, partial
execution, behavioral assets, and replay. Export never converts a record back into application
calls and never mutates the source directory.

## Install The Exporter

The base package does not import or install OpenTelemetry. Add the dedicated extra only on a worker
that exports records:

=== "uv"

    ```bash
    uv add 'autobench[otlp]'
    ```

=== "pip"

    ```bash
    python -m pip install 'autobench[otlp]'
    ```

This installs the OpenTelemetry SDK and HTTP/protobuf trace exporter. It does not replace ABP
instrumentors and does not enable OTel-to-ABP ingestion.

## Export From The CLI

Point the command at a completed or explicitly partial Autobench record:

```bash
autobench telemetry export runs/routing-42 \
  --endpoint https://collector.example/v1/traces \
  --header authorization 'Bearer ...' \
  --service-name routing-benchmark \
  --service-namespace evaluation
```

When `--endpoint` and `--header` are omitted, the underlying exporter can use standard
`OTEL_EXPORTER_OTLP_TRACES_ENDPOINT` and `OTEL_EXPORTER_OTLP_HEADERS` configuration. Credentials
belong in environment variables or CLI secret injection, never in a benchmark YAML file.

Useful options:

| Option | Meaning |
| --- | --- |
| `--endpoint` | OTLP HTTP/protobuf trace endpoint |
| `--header NAME VALUE` | Repeatable request header |
| `--timeout SECONDS` | Positive exporter timeout |
| `--service-name` | OTLP `service.name`; default is `autobench` |
| `--service-namespace` | Optional OTLP `service.namespace` |
| `--certificate-file` | Custom CA certificate path |
| `--include-captured-content` | Export captured bodies and other content-bearing evidence |

The command prints a Rich summary containing experiment and benchmark identity, record version,
run/trace/span counts, partial counts, and the selected endpoint. Export failures return a nonzero
exit status and leave the record unchanged.

## Export From Python

```python
from pathlib import Path

from autobench import OTLPSettings, export_record_otlp

result = export_record_otlp(
    Path("runs/routing-42"),
    settings=OTLPSettings(
        endpoint="https://collector.example/v1/traces",
        headers={"authorization": "Bearer ..."},
        timeout_seconds=10,
        service_name="routing-benchmark",
        service_namespace="evaluation",
        resource_attributes={"deployment.environment": "staging"},
    ),
)

print(result.exported_span_count)
```

`export_otlp(experiment, runs, ...)` accepts already loaded `ExperimentRecord` and `RunRecord`
models. It validates run count, unique run IDs, benchmark and experiment ownership, and execution
correlation before mapping anything. Tests and custom delivery layers may inject an OTel
`SpanExporter`; Autobench does not shut down an exporter it does not own.

## Evidence Mapping

One export creates this hierarchy:

```text
autobench.experiment <benchmark_id>
  autobench.run <case_id> / <variant_id>
    autobench.trace <trace_id-prefix>
      <ABP root span>
        <ABP child span>
```

| Autobench evidence | OTLP representation |
| --- | --- |
| Experiment lifecycle | Root span, status, identity attributes, termination event |
| Run lifecycle | Child span with run/case/variant, record path, status, and partial state |
| ABP trace envelope | Child span with protocol version, trace ID, partial state, diagnostics |
| ABP span | Nested span with original operation, timestamps, scope, source convention, usage, and stream attributes |
| Factor, observation, score | Timestamped span event on the run |
| Measurement | Timestamped event on its ABP span |
| Asset use and source snapshot | Run events preserving version/provenance identity |
| ABP diagnostic/reference | Trace or span events |
| Trace-and-span link target | Native OTLP span link plus the lossless ABP link event |
| Error | Error status and structured exception/termination event |

Historical benchmark measurements are exported as span events, not fabricated live OTLP metric
streams. Their canonical semantic type, unit, direction, role, and source remain in the event, so a
consumer can project them deliberately without confusing replay time with measurement time.

The exporter preserves experiment, benchmark, run, case, variant, record-version, ABP trace/span,
dataset, spec hash, manifest, correlation, instrumentation scope, package version, source-map, and
source-convention identities where available. Reserved Autobench resource identities cannot be
overridden through custom resource attributes.

## Partial And Cancelled Records

Completed, cancelled, aborted, and recovered partial records are all exportable. Experiment and run
spans retain terminal status, `partial`, end reason, planned/recorded/missing run IDs, and whether
cross-run derivation and policy evaluation completed. An open or partial ABP trace is marked as
partial rather than being represented as a successful complete trace.

OTLP export is reporting, not recovery. Use `autobench recording inspect` and
`autobench recording finalize --allow-partial` to publish recoverable staging evidence first.

## Content And Privacy

By default, export keeps structural and semantic evidence while omitting captured content:

- event bodies and ABP span outputs;
- score actual/expected values;
- non-metric observation values;
- retained source-fact values;
- tracebacks.

References, hashes, counts, semantic types, status, usage, asset IDs/versions, and provenance remain
available. `--include-captured-content` or `OTLPSettings(include_captured_content=True)` includes
content that ABP already captured. It cannot recover content excluded or redacted by the original
capture policy. Treat this switch as a deliberate data-export decision.

## Failure Boundary

Mapping, network export, exporter rejection, and owned-exporter shutdown failures raise
`OTLPExportError`. The original Pydantic record models and files are never changed. An injected
exporter remains caller-owned and is not shut down. This allows export retries or delivery to
several telemetry systems without changing benchmark evidence or replay lineage.

## What This Adapter Does Not Do

- It does not make OTel an Autobench runtime dependency.
- It does not ingest arbitrary OTel traces into ABP.
- It does not patch application SDKs through OTel instrumentations.
- It does not use OTLP as benchmark persistence.
- It does not add vendor configuration to `BenchmarkSpec`.

Use [Native Instrumentation](native-instrumentation.md) to collect automatic application evidence
into ABP, then use this adapter when the resulting record also needs to appear in an OTLP backend.

---

## YAML Spec

Canonical page: https://vcoderun.github.io/autobench/yaml-spec/

# YAML Spec

Autobench is YAML-first. Python builders compile to the same internal `BenchmarkSpec`.
Every YAML file written by Autobench includes a `yaml-language-server` schema header that points
to the versioned schema cache under `~/.autobench/<version>/schemas/`.

## Authoring Sections

The authoring DSL places the benchmark ID under `benchmark` and keeps all behavior inside that
named benchmark:

| Section | Required | Purpose |
| --- | --- | --- |
| `description` | No | Human-readable benchmark intent |
| `execution` | No | Static cross-invocation correlation metadata |
| `dataset` | Yes | Inline or file-backed cases, defaults, version, and metadata |
| `run` | For execution | Python task target |
| `variants` | Yes | Named factor combinations |
| `score` | No | Built-in or Python scorers |
| `derive` | No | Per-run semantic derivation such as token cost |
| `post_derive` | No | Cross-run derivation such as paired baseline |
| `policies` | No | Semantic metric constraints |
| `report` | No | Leaderboard, matrix, comparisons, and distributions |
| `semantic_registry` | No | Custom semantic definitions and aliases |

## Complete Authoring Example

```yaml
# yaml-language-server: $schema=./schemas/0.3.0/benchmark_schema.json
benchmark:
  support-routing:
    description: Compare current and candidate routing behavior.
    execution:
      correlation:
        group_id: routing-proposal-42
        attempt: 1
        phase: validation
        parent_experiment_id: routing-baseline-17
        labels:
          owner: evaluation
          seed: 7
    dataset:
      source: file://datasets/cases.yaml
      version: v2
      defaults:
        tags: [regression]
    run:
      python: benchmark_tasks:run_case
    variants:
      baseline:
        factors:
          model:
            value: openrouter:openai/gpt-5.6-luna
            semantic: llm.model.name
          prompt_version:
            value: route-v3
            semantic: prompt.version
            optimize: true
      candidate:
        factors:
          model:
            value: openrouter:openai/gpt-5.6-luna
            semantic: llm.model.name
          prompt_version:
            value: route-v4
            semantic: prompt.version
            optimize: true
    score:
      route_correctness:
        exact:
          actual: output.route
          expected: case.expected.route
        semantic: quality.correctness
        goal: maximize
        role: objective
      success:
        pass: output.ok
        semantic: result.success
        role: constraint
    derive:
      - kind: token_cost
        pricing: file://pricing/models.yaml
        output:
          name: request_cost
          semantic_type: money.cost
          unit: usd
          direction: minimize
          role: constraint
    policies:
      - name: must-succeed
        metric: result.success
        must_equal: true
    report:
      leaderboard:
        show:
          accuracy:
            metric: quality.correctness
            aggregate: ratio_true
          total_cost:
            metric: money.cost
            aggregate: sum
      matrix:
        metric: quality.correctness
      compare:
        baseline -> candidate:
          show:
            accuracy:
              metric: quality.correctness
              aggregate: ratio_true
      markdown:
        profile: full
        layout: auto
        output: reports/benchmark.md
```

`report.markdown` also accepts bounded `limits`, `traces.top_slowest`, `assets.diffs`, and
`content.include_captured`. A configured output is staged before a CLI-recorded experiment is
sealed. See [Markdown Reports](markdown-reports.md) for the complete document, audit, and
publication contracts.

`execution.correlation` groups separate benchmark invocations without changing matrix identity.
`attempt` must be positive, and labels accept only stable string, integer, finite float, or boolean
values. Correlation is copied unchanged to the experiment and every run record. It is not replay
lineage and does not claim that Autobench can resume application workflow state.

Python or CLI overrides merge field by field. An omitted override keeps the YAML value; supplied
labels replace matching keys and preserve the other YAML labels. Per-case or per-variant correlation
resolvers are intentionally not part of this surface.

## Resolution Rules

- File references resolve relative to the benchmark YAML.
- Python targets use `module:callable` and receive inferred search paths from the spec directory.
- Duplicate case and variant IDs are validation errors.
- A nonempty runnable matrix requires a task.
- Scorer definitions must select exactly one scoring action.
- Remote file references are rejected; price-source URL loading is an explicit integration API.
- Custom semantics should be declared in the semantic registry.

## Shape

```yaml
benchmark:
  support-routing:
    description: Deterministic support routing benchmark.
    dataset:
      source: file://datasets/cases.yaml
      defaults:
        metadata:
          owner: docs
    run:
      python: app.benchmarks.support:run_ticket_case
    variants:
      route_v1:
        factors:
          prompt_version:
            value: route-v1
            semantic: prompt.version
            optimize: true
          routing_profile: baseline
    score:
      routing_correctness:
        exact:
          actual: output.queue
          expected: case.expected.queue
        semantic: quality.correctness
      tool_arguments:
        expected_action:
          metric: arguments
          observed_kind: tool
        span:
          kind: tool
        semantic: agent.tool.argument.correctness
    report:
      leaderboard:
        show:
          pass_rate:
            metric: result.success
            aggregate: ratio_true
```

## Exported Benchmark YAML

When Autobench renders a benchmark spec back to YAML, it uses a DSL-like shape instead of a raw
model dump:

```yaml
benchmark:
  support-routing:
    description: Route support tickets.
    dataset:
      source: datasets/cases.yaml
      cases:
        - id: ticket_1
          input:
            subject: Refund
    run:
      python: app.benchmarks.support:run_ticket_case
    variants:
      route_v1:
        factors:
          prompt_version:
            value: route-v1
            semantic: prompt.version
            optimize: true
          routing_profile: baseline
    score:
      success:
        pass: output.matched
        semantic: result.success
        goal: maximize
    report:
      leaderboard:
        show:
          pass_rate:
            metric: result.success
            aggregate: ratio_true
```

## Notes

- `dataset.source` supports local `file://` references and globs.
- task targets use `module:function`.
- variant factors accept either mapping or list form.
- YAML does not execute inline expressions.
- importable code hooks such as Python scorers remain explicit dotted targets.
- `score.<name>.span` can target component spans by kind, name, tag, path, or semantic type.
- `expected_action` scores compare `case.expected.actions` or `case.expected.tool_calls` with observed spans.

## Native Instrumentation

The optional `instrumentation` section installs ABP SDK integrations for the complete benchmark
matrix:

```yaml
benchmark:
  support-agent:
    instrumentation:
      all:
        exclude: [httpx]
        strict: false
        assets:
          discover: true
          representations: [definition, effective]
          include: [prompt, tool, output_schema, capability]
      pydantic_ai: {}
      openai: {}
      openai_agents: false
      httpx:
        capture:
          path: hash
          request_headers: [x-request-id]
          response_headers: [x-request-id]
          request_body: false
          response_body: false
          max_body_bytes: 65536
```

`all` discovers every installed, compatible built-in integration. Missing integrations are skipped
and recorded as run diagnostics unless `strict: true` is set. `exclude` accepts `pydantic_ai`,
`openai`, `openai_agents`, and `httpx`. An explicit entry, including `false`, overrides discovery;
the explicit HTTPX block above therefore remains enabled despite the discovery exclusion.

`{}` selects privacy-safe defaults. `false` disables a known integration. Unknown integration
names, settings, exclusions, or HTTP capture modes are validation errors. Optional SDKs are
imported only when their enabled integration is resolved for execution. Replay never resolves this
section.

The versioned `benchmark_schema.json` describes this surface, so YAML language servers complete
integration names and capture settings. See [Native Instrumentation](native-instrumentation.md) for
the lifecycle and privacy contract.

## Capture Policy

The benchmark-level `capture` section controls ABP evidence and discovered asset content for every
case/variant run:

```yaml
benchmark:
  private-agent:
    capture:
      default_level: hash
      asset_default_level: hash
      use_semantic_defaults: false
      semantic_overrides:
        tool: full
        output_schema: full
      deny_paths:
        - assets.*:prompt:private_notes
```

Supported levels are `none`, `metadata`, `hash`, `redacted`, and `full`. `default_level` controls
runtime evidence and defaults to `metadata`; `asset_default_level` controls versioned behavioral
assets and defaults to `full`. Other fields include
semantic/path allow and deny lists, secret names, inline/artifact limits, collection/string/depth
limits, binary retention, and source-attribute retention. Unknown fields or levels fail validation.
See [Automatic Asset Discovery](automatic-asset-discovery.md#privacy-and-capture-policy) for the
content/version behavior.

## Safe Extensibility

YAML is intended to be shareable and replayable. For that reason:

- file references are resolved relative to the spec path
- remote URLs are rejected
- inline Python expressions are not part of the YAML surface

## Exported Run Record YAML

Run records are the immutable per-case/per-variant evidence files used by replay. The trace signal
objects below are abridged; recorded files retain their timestamps, sequence IDs, execution
references, scope provenance, and captured attributes:

```yaml
record:
  type: run
  version: 4

protocol:
  name: abp
  version: 1
  semantic_registry: 1

run:
  id: run_ticket_1_route_v1
  experiment: exp_support_routing_20260507T120000Z
  benchmark: support-routing
  case: ticket_1
  variant: route_v1
  status: passed
  outcome:
    evaluation: passed
    task: passed

case:
  id: ticket_1
  input:
    subject: Refund
  expected:
    queue: billing

variant:
  id: route_v1
  factors:
    prompt_version:
      value: route-v1
      semantic: prompt.version
      optimize: true

scores:
  routing_correctness:
    value: true
    semantic: quality.correctness
    role: objective

metrics:
  measurements:
    routing_correctness:
      id: observation_1
      name: routing_correctness
      kind: metric
      value: true
      semantic: quality.correctness
  diagnostics:
    latency_ms:
      value: 12.4
      semantic: time.latency
      unit: ms

trace:
  protocol: abp
  protocol_version: 1
  trace_id: 70d8f4b6742d412a85cb7a198db07fe1
  execution:
    benchmark_id: support-routing
    experiment_id: exp_support_routing_20260507T120000Z
    run_id: run_ticket_1_route_v1
    case_id: ticket_1
    variant_id: route_v1
  root_span_ids: [3f2f6c57b9f56a11]
  spans:
    - span_id: 3f2f6c57b9f56a11
      operation: benchmark.run
      kind: task
      scope:
        instrumentor_name: autobench.manual
        instrumentor_version: 0.3.0
        package_name: autobench
        package_version: 0.3.0
        mechanism: manual
        layer: application
      status: ok
      end_reason: completed
      measurements: []
      events: []
      links: []
      references: []
      partial: false
  links: []
  references: []
  diagnostics: []
  signals:
    - type: span_start
      protocol: abp
      protocol_version: 1
      span_id: 3f2f6c57b9f56a11
      operation: benchmark.run
      kind: task
    - type: span_end
      protocol: abp
      protocol_version: 1
      span_id: 3f2f6c57b9f56a11
      status: ok
      reason: completed
  partial: false

spans:
  call_router:
    kind: workflow
    started_at: "2026-05-07T12:00:00Z"
    duration: 0.0124
    attributes:
      component: router
  lookup_user:
    kind: tool
    parent: call_router
    input:
      user_id: u1
    output:
      tier: gold
    duration: 0.004

artifacts:
  generated_spec:
    media: application/x-yaml
    path: artifacts/run_ticket_1_route_v1/generated_spec.yaml

assets:
  prompt.router:
    version: 7c91d4d7b1af

output:
  queue: billing
```

When the serialized ABP trace exceeds the inline limit, the same section becomes a compact summary
and artifact reference:

```yaml
trace:
  id: 70d8f4b6742d412a85cb7a198db07fe1
  partial: false
  spans: 7
  signals: 31
  artifact:
    id: abp_trace
    name: ABP trace
    media: application/vnd.autobench.abp-trace+yaml
    path: artifacts/run_ticket_1_route_v1/trace.yaml
```

## Exported Dataset YAML

Dataset exports use a DSL-like shape instead of raw model dumps:

```yaml
record:
  type: dataset
  version: 1

dataset:
  id: tickets
  version: v1
  metadata:
    owner: support
  defaults:
    tags: [smoke]
  cases:
    - id: ticket_1
      input:
        subject: Refund
```

Generated datasets use this exact dataset format. Generation itself has two additional schemas. A
request is a portable input to an application-owned generator:

```yaml
# yaml-language-server: $schema=schemas/0.3.0/generation_request_schema.json
generation:
  request:
    seed: 17
    prompt:
      content: Generate privacy-safe routing cases.
      asset_version: prompt.routing-generator@v1
    settings:
      count: 20
    seed_cases:
      - id: reviewed-refund
        input: {message: Refund a duplicate charge}
        expected: {route: billing}
```

The resulting `.generation.yaml` or `.incomplete.yaml` manifest uses
`generation_schema.json`. It records generator/provider/model identity, determinism, request and
case hashes, usage, cost, review states, rejected-case reasons, output counts, and the published
dataset reference. Prompt content is represented by its hash and tracked asset version rather than
duplicated into the manifest. See [Generated Datasets](generated-datasets.md).

## Exported Semantic Registry YAML

Semantic registry exports use stable type ids with compact metadata:

```yaml
record:
  type: semantic_registry
  version: 1

semantic_registry:
  version: 1
  aliases:
    quality.answer: quality.score
  types:
    money.cost:
      unit: usd
      shape: number
    serving.cost:
      parent: money.cost
      unit: usd
      shape: number
```

## Exported Pricing YAML

Pricing tables are helper data, not a required runtime dependency. They keep provider/model aliases
and tiered token prices readable:

```yaml
record:
  type: pricing
  version: 1

pricing:
  provider: openrouter
  source: genai-prices
  updated_at: "2026-05-07"
  models:
    google/gemini-3-flash-preview:
      name: Gemini 3 Flash Preview
      aliases:
        - google:gemini-3-flash-preview
        - openrouter/google/gemini-3-flash-preview
      input:
        unit: mtok
        price: 0.3
        tiers:
          - up_to: 1000000
            price: 0.3
          - price: 0.6
      output:
        unit: mtok
        price: 2.5
      cache_read:
        unit: mtok
        price: 0.03
```

## Exported Report YAML

Report exports keep the summary under a single `report:` body:

```yaml
record:
  type: report
  version: 1

report:
  benchmark: support-routing
  experiment: exp_support_routing_20260507T120000Z
  runs: 6
  status:
    passed: 5
    failed: 1
  variants:
    baseline:
      factors:
        model.name: openrouter:openai/gpt-5.6-luna
  leaderboard:
    baseline:
      runs: 2
      metrics:
        avg_coverage: 0.82
  cases:
    ticket_1:
      baseline:
        status: passed
        metrics:
          coverage (coverage.ratio): 0.8
  matrix:
    metric: coverage.ratio
    cases:
      ticket_1:
        baseline: 0.8
  compare:
    baseline -> candidate:
      runs: 2
      confounded: true
  distributions:
    cost_distribution:
      semantic: money.cost
      variants:
        baseline: [0.01, 0.02]
```

## Exported Experiment YAML

Experiment records keep replay data structured, but the outer shape stays readable:

```yaml
record:
  type: experiment
  version: 5

experiment:
  id: exp_support_routing_20260507T120000Z
  benchmark: support-routing
  termination:
    status: completed
    partial: false
    post_processing:
      cross_run_derivation: true
      policies: true
    planned_runs: [run_ticket_1_route_v1]
    recorded_runs: [run_ticket_1_route_v1]
    missing_runs: []

benchmark:
  id: support-routing
  dataset:
    id: tickets
    version: v1
    hash: 9b5d...
  cases:
    - ticket_1
    - ticket_2
  counts:
    cases: 2
    variants: 3
    runs: 6
  warnings: []
  spec:
    hash: a13c...
    snapshot:
      benchmark:
        id: support-routing

runs:
  count: 6
  passed: 5
  failed: 1
  errored: 0
  skipped: 0
  cancelled: 0
  paths:
    - cases/ticket_1/route_v1/run.yaml

manifest: manifest.yaml

files:
  /abs/path/autobench.yaml: 3c4d...

environment:
  python: "3.11.13"
  platform: macOS-15.5-arm64-arm-64bit
  cwd: /workspace/autobench

semantic_registry:
  version: 1
  aliases:
    quality.answer: quality.score
  types:
    money.cost:
      unit: usd
      shape: number
```

The corresponding integrity manifest is also human-readable and schema-backed:

```yaml
record:
  type: manifest
  version: 1

experiment:
  id: exp_support_routing_20260507T120000Z

files:
  - path: experiment.yaml
    sha256: 58f4...
    bytes: 2148
    kind: experiment
    identity: exp_support_routing_20260507T120000Z
  - path: cases/ticket_1/route_v1/run.yaml
    sha256: 2c91...
    bytes: 4892
    kind: run
    identity: run_ticket_1_route_v1
```

Run records in format version 5 add `run.partial` and `run.end_reason`. Format version 6 adds
optional execution correlation to experiment, summary, run, staging, and checkpoint documents.
Existing records without
these fields remain loadable; Autobench infers completed, failed, deferred, or cancelled lifecycle
state from their legacy status.

## Exported Artifact YAML

Artifacts are split into metadata and payload files. Text payloads stay as text. Structured payloads
are wrapped so they remain recognizable YAML records:

```yaml
record:
  type: artifact
  version: 1

artifact:
  id: trace
  name: trace
  media_type: application/x-yaml
  span_id: call_router
  payload: artifacts/run_ticket_1_route_v1/trace.yaml
```

```yaml
record:
  type: artifact_payload
  version: 1

artifact:
  id: trace
  name: trace
  media_type: application/x-yaml

payload:
  steps:
    - tool: route_ticket
      arguments:
        queue: billing
```

## Exported Asset YAML

Tracked assets are stored as a readable index plus per-asset history files:

```yaml
record:
  type: asset_index
  version: 1

assets:
  tool.create_car:
    kind: tool
    name: create_car
    semantic: agent.tool
    current_version: 7c91d4d7b1af
    file: tool_create_car.yaml
```

```yaml
record:
  type: asset
  version: 2

asset:
  id: tool.create_car
  kind: tool
  name: create_car
  semantic: agent.tool
  current_version: 7c91d4d7b1af
  content_ref:
    asset_id: tool.create_car
    version: 7c91d4d7b1af
    path: artifacts/asset-content.sqlite3

versions:
  - version: 15aa0dbceb02
    content_ref:
      asset_id: tool.create_car
      version: 15aa0dbceb02
      path: artifacts/asset-content.sqlite3
    hashes:
      content: ...
    changes:
      fields: [initial]
  - version: 7c91d4d7b1af
    parent: 15aa0dbceb02
    content_ref:
      asset_id: tool.create_car
      version: 7c91d4d7b1af
      path: artifacts/asset-content.sqlite3
    hashes:
      content: ...
      source: ...
    source:
      path: ./vsh.py
    changes:
      fields:
        - params.year.type
      diff_ref:
        asset_id: tool.create_car
        version: 7c91d4d7b1af
        parent_version: 15aa0dbceb02
        path: artifacts/asset-content.sqlite3
```

The referenced SQLite artifact is an internal, transaction-safe content store rather than an
authoring DSL. It keeps content-addressed snapshots and readable diffs out of manifests while
supporting indexed lookup by `asset_id` and `version`. Use `load_asset_content(...)` or
`load_asset_diff(...)`; application code does not query its tables directly.

---

## Python API

Canonical page: https://vcoderun.github.io/autobench/python-api/

# Python API

Autobench exposes the same runtime through a fluent builder, typed specification models, and
lower-level extension seams. Use the highest-level surface that can express the benchmark clearly.

## Surface Selection

| Surface | Use it when |
| --- | --- |
| `Benchmark` | Application code composes a benchmark dynamically |
| `BenchmarkSpec` | You need the complete typed configuration surface |
| YAML + `load_benchmark_spec` | Humans or agents author portable benchmark definitions |
| Runtime/evaluation functions | You are building an adapter, service, or custom runner |

All three authoring paths execute through `run_benchmark_spec()`.

## Fluent Builder

```python
from autobench import (
    Benchmark,
    Case,
    Direction,
    ExactScorer,
    FactorValue,
    ObservationRole,
    PassFailScorer,
    Semantic,
    Variant,
)

benchmark = (
    Benchmark("builder-demo")
    .description("Compare current and candidate behavior.")
    .correlation(
        group_id="routing-proposal-42",
        attempt=1,
        phase="validation",
        labels={"owner": "evaluation"},
    )
    .dataset(
        [
            Case(
                id="refund",
                input={"message": "Refund order 42"},
                expected={"route": "billing"},
            )
        ],
        dataset_id="routing-regressions",
        version="v3",
    )
    .variants(
        [
            Variant(
                id="current",
                factors=[FactorValue(name="routing_profile", value="v3")],
            ),
            {
                "id": "candidate",
                "factors": {
                    "routing_profile": {
                        "value": "v4",
                        "optimize": True,
                    }
                },
            },
        ]
    )
    .task("my_app.benchmarks:run_case")
    .scoring(
        [
            ExactScorer(
                name="route",
                actual="output.route",
                expected="case.expected.route",
                semantic_type=Semantic.QUALITY_CORRECTNESS,
                direction=Direction.MAXIMIZE,
                role=ObservationRole.OBJECTIVE,
            ),
            PassFailScorer(
                name="success",
                path="output.success",
                semantic_type=Semantic.RESULT_SUCCESS,
                role=ObservationRole.CONSTRAINT,
            ),
        ]
    )
)

result = benchmark.run(experiment_id="routing-candidate-42", concurrency_limit=4)
```

Attach live lifecycle observers without changing the benchmark definition:

```python
from autobench import ProgressEvent


def observe(event: ProgressEvent) -> None:
    print(event.sequence, event.kind, event.run_id, event.run_status)


result = benchmark.run(progress_handlers=(observe,))
```

### Builder Methods

| Method | Configures |
| --- | --- |
| `description(value)` | Benchmark description |
| `correlation(...)` | Static invocation group, attempt, phase, ancestry hints, and scalar labels |
| `capture(policy)` | ABP and asset capture policy |
| `dataset(...)` | Inline cases or a typed dataset source |
| `variants(items)` | Typed variants or normalized dictionaries |
| `task(target, kind="python")` | Task target |
| `scoring(items)` | Built-in or Python scorer specs |
| `derive(items)` | Per-run derivers |
| `instrument(*items)` | Typed built-ins or runtime custom instrumentors |
| `instrument_all(...)` | Compatible built-in discovery |
| `to_spec()` | Canonical `BenchmarkSpec` |
| `run(...)` / `run_async(...)` | Sync or async execution |

Post-derivation, policies, report views, and custom semantic registries currently live on the full
`BenchmarkSpec`. Extend the compiled spec rather than inventing builder-only state:

```python
import asyncio

from autobench import PolicySpec, run_benchmark_spec

spec = benchmark.to_spec().model_copy(
    update={
        "policies": [
            PolicySpec(
                name="quality-floor",
                metric=Semantic.QUALITY_CORRECTNESS,
                must_greater_equal=0.9,
            )
        ]
    }
)
result = asyncio.run(run_benchmark_spec(spec, concurrency_limit=4))
```

## Task Contract

```python
from autobench import Case, RunContext


def run_case(ctx: RunContext, case: Case) -> Result:
    ...
```

`ctx` is always first and `case` is always second. A task may be sync or async and may return any
serializable result. A Pydantic model is useful because scorers can resolve output fields reliably.

The runtime resolves `module:function` targets relative to the benchmark file before falling back to
normal Python import paths.

## RunContext

`RunContext` owns evidence for one case x variant run:

| Method | Purpose |
| --- | --- |
| `factor(name)` | Read a configured factor value |
| `span(...)` | Time and nest an operation |
| `metric(...)` / `metrics(...)` | Record numeric, boolean, or structured metrics |
| `factor_observation(...)` | Record a factor discovered at runtime |
| `event(...)` | Record a discrete event |
| `diagnostic(...)` | Record non-objective evidence |
| `outcome(...)` | Record semantic success |
| `check(...)` | Record a correctness constraint and reason |
| `record_measurement(...)` | Record summaries plus optional raw samples |
| `artifact(...)` | Attach a payload |
| `error(...)` | Preserve a structured error |
| `attach_tracked_asset(...)` | Bind an explicit tracked asset version |
| `await checkpoint(name)` | Commit all currently available evidence to durable staging |

Evidence emitted before an exception remains in the failed run.

`phase` reports whether the run is resolving, executing, scoring, deriving, post-processing, or
finalizing. Checkpoints preserve that phase automatically. Application tasks normally only call
`checkpoint()`:

```python
async def run_case(ctx: RunContext, case: Case) -> Result:
    draft = await build_draft(case.input)
    ctx.artifact("draft", draft)
    await ctx.checkpoint("draft-built")
    return await validate_draft(draft)
```

The method requires an active `Recorder`; without one it raises `RuntimeError`. It is cancellation
aware: if the caller is cancelled while persistence is active, Autobench gives the atomic write
bounded time to finish and then propagates cancellation. Names beginning with `autobench.` belong
to framework lifecycle checkpoints and are rejected for application calls.

## Load And Run YAML

```python
import asyncio
from pathlib import Path

from autobench import (
    ExecutionCorrelation,
    load_benchmark_spec,
    run_benchmark_path,
    run_benchmark_spec,
)

path = Path("benchmarks/routing.yaml")
spec = load_benchmark_spec(path)

sync_result = run_benchmark_path(
    path,
    experiment_id="routing-42",
    concurrency_limit=4,
)

async_result = asyncio.run(
    run_benchmark_spec(
        spec,
        experiment_id="routing-43",
        concurrency_limit=4,
        correlation=ExecutionCorrelation(attempt=2, labels={"review": "holdout"}),
        progress_handlers=(observe,),
    )
)
```

Import `ExecutionCorrelation` for invocation-level overrides. Explicit fields merge with
`spec.execution.correlation`; omitted fields remain unchanged and label maps merge by key. The
resolved value is immutable and identical on `ExperimentResult`, every `RunResult`, durable
records, and replayed results. `parent_experiment_id` and `resumed_from_experiment_id` are grouping
metadata only. Replay ancestry remains `RecordLineage` / `parent_run_id`, and Autobench does not
infer workflow resume from either correlation field.

`run_benchmark_path()`, `run_benchmark_spec()`, `Benchmark.run()`, and `Benchmark.run_async()` share
the same `progress_handlers`, `progress_error_policy`, and `progress_error_handler` contract. Bare
handlers are strict by default; see [Tasks and Runtime](tasks-and-runtime.md#progress-events) for
ordering, terminal status, and backpressure guarantees.

Loading resolves dataset, pricing, task, and Python scorer references relative to the YAML file.

## Record And Replay

```python
from pathlib import Path

from autobench import (
    collect_benchmark_source_files,
    record_experiment,
    replay_experiment,
)

record_dir = Path("runs/routing-42")
record = record_experiment(
    async_result,
    record_dir,
    source_files=list(collect_benchmark_source_files(path)),
    path_root=Path.cwd(),
    durability="atomic",
)
replayed = replay_experiment(record_dir)
```

`record_experiment()` refuses to overwrite a non-empty experiment directory. It assembles the
record in a temporary sibling, validates its integrity manifest, and publishes it atomically.
Referenced tracked-asset histories and large trace artifacts are persisted automatically. Use
`durability="synced"` when supported POSIX file and directory `fsync` calls are also required.

For durable execution, pass a recorder to the pipeline instead of waiting for the full result:

```python
from autobench import FileRecorder

durable_result = asyncio.run(
    run_benchmark_spec(
        spec,
        experiment_id="routing-durable-44",
        concurrency_limit=4,
        recorder=FileRecorder(
            Path("runs/routing-durable-44"),
            source_files=collect_benchmark_source_files(path),
            path_root=Path.cwd(),
            durability="atomic",
        ),
    )
)
```

`FileRecorder` commits complete run snapshots during execution and atomically publishes the final
directory after post-processing. Its sibling `.routing-durable-44.staging` directory survives an
interrupted process. Use `inspect_staging`, `recover_staging`, and `finalize_staging` to recover it;
use `archive_staging` or `discard_staging` for explicit lifecycle decisions.

Cooperative task cancellation, `KeyboardInterrupt`, and supported CLI `SIGTERM` handling commit a
terminal partial checkpoint before propagating. A hard process kill can preserve only the last
explicit checkpoint that had already returned; it cannot execute a final write.

```python
from autobench import finalize_staging, inspect_staging

staging = Path("runs/.routing-durable-44.staging")
inspection = inspect_staging(staging)
if inspection.recoverable and inspection.missing_run_ids:
    partial_record = finalize_staging(
        staging,
        Path("runs/routing-durable-44-partial"),
        allow_partial=True,
    )
```

`Recorder` and `RecordSession` are the typed extension contracts for another persistence backend.
The pipeline owns `open`, `stage`, `finish` or `abort`, and `close`; application code should not
open a live session just to run a normal benchmark. `ExperimentStart`, `ExecutionSnapshot`, and
`PartialRunSnapshot` are frozen transfer models. Recovery never imports the task or optional SDKs.

A custom session exposes immediate file and stream storage explicitly through
`artifact_sink: ArtifactSink | None`. The session may return itself, as `FileRecordSession` does,
or delegate to a separate local, remote, or composite sink. Returning `None` is valid, but a task
that calls `ctx.artifact_file()` or `ctx.artifact_stream()` then receives
`ArtifactSinkRequiredError` before Autobench consumes the source. The lifecycle session does not
need to proxy the complete sink protocol merely to delegate storage.

`ArtifactSink` provides synchronous file/stream methods, `prepare_file_async()`, and
`prepare_stream_async()`. Async implementations own in-flight transfers after caller cancellation
and must settle them before the associated session closes or publishes its final record.

Load one exact record when building an audit or optimizer adapter:

```python
from autobench import load_experiment_record, load_run_record

experiment = load_experiment_record(record_dir)
run = load_run_record(record_dir / experiment.run_paths[0], root_dir=record_dir)
```

## Reports And Exports

```python
from pathlib import Path

from autobench import (
    build_report,
    compare_variants,
    export_markdown_report,
    export_runs_csv,
    export_summary_yaml,
    load_experiment_record,
    write_markdown_report,
)

record = load_experiment_record(record_dir)
report = build_report(
    replayed,
    experiment_record=record,
    experiment_root=record_dir,
)
comparison = compare_variants(
    replayed,
    baseline="current",
    candidate="candidate",
)

export_summary_yaml(replayed, Path("analysis/summary.yaml"))
export_runs_csv(replayed, Path("analysis/runs.csv"))
export_markdown_report(replayed, Path("analysis/report.md"))
publication = write_markdown_report(
    report,
    Path("analysis/report-bundle"),
    layout="bundle",
    immutable_root=record_dir,
)
```

`build_leaderboard`, `build_case_matrix`, `build_metric_distribution`, and
`build_run_metric_rows` expose individual projections. `write_markdown_report()` returns profile,
selected layout, output paths, byte counts, and SHA-256 hashes. See
[Markdown Reports](markdown-reports.md) for configuration and safety boundaries.

For a case-level benchmark verdict, return a mapping or Pydantic model containing `hard_pass`,
`score`, `metrics`, and `feedback` from the task. `build_report()` keeps that quality outcome
separate from task execution status and projects it into KPIs, case tables, and purposeful inline
SVG. Technical run, trace, asset, artifact, hash, and provenance detail belongs to `audit`.

Group or select several invocation results without changing their records:

```python
from autobench import ExecutionCorrelation, build_grouped_reports, filter_experiments

validation = filter_experiments(
    results,
    correlation=ExecutionCorrelation(group_id="routing-proposal-42", phase="validation"),
)
groups = build_grouped_reports(results)
```

Filters match only explicitly supplied fields. Grouped reports retain each experiment report and
summarize the attempts and phases present under each `group_id`.

## Optional OTLP Export

```python
from pathlib import Path

from autobench import OTLPSettings, export_record_otlp

delivery = export_record_otlp(
    Path("runs/routing-42"),
    settings=OTLPSettings(
        endpoint="https://collector.example/v1/traces",
        service_name="routing-benchmark",
    ),
)
```

Install `autobench[otlp]` on the exporting process. `export_record_otlp()` loads immutable record
models; `export_otlp()` accepts already-loaded `ExperimentRecord` and `RunRecord` values. Both
return `OTLPExportResult` and raise `OTLPExportError` without modifying evidence. Vendor settings
remain separate from `BenchmarkSpec`. See [OTLP Export](otlp-export.md).

## Native Instrumentation

```python
from autobench import Benchmark

benchmark = Benchmark("agent").instrument_all(
    exclude={"httpx"},
    strict=False,
    assets={
        "representations": ["definition", "effective"],
        "include": ["prompt", "tool", "output_schema"],
    },
)
```

Unavailable integrations become diagnostic observations. `strict=True` instead requires every
selected integration to be compatible.

Use typed settings for explicit control:

```python
from autobench import HTTPXCaptureSettings, HTTPXInstrumentation, OpenAIInstrumentation

benchmark.instrument(
    OpenAIInstrumentation(),
    HTTPXInstrumentation(
        capture=HTTPXCaptureSettings(
            path="hash",
            response_headers=("x-request-id",),
        )
    ),
)
```

Explicit settings override automatic discovery, including `enabled=False`. A custom runtime
`Instrumentor` can also be passed to `instrument()` and remains Python-only.

## Explicit Tracking

```python
from autobench import track

SYSTEM_PROMPT = track.prompt(
    name="support_system",
    source="prompts/support.md",
)


@track.tool
def lookup_order(order_id: str) -> dict[str, str]:
    """Return the current order status."""
    ...
```

`track.prompt`, `track.tool`, `track.type`, `track.dataclass`, and `track.asset` register exact
versions. `track.write_assets(path)` writes DSL-shaped manifests plus one `content.sqlite3`
registry. `load_asset_content(...)` resolves an exact historical snapshot and
`load_asset_diff(...)` resolves the corresponding readable diff. Native discovery can attach
unadorned SDK-visible components to runs. Experiment recording uses the same contract at
`artifacts/asset-content.sqlite3`.

## Production And Generated Cases

```python
from pathlib import Path

from autobench import (
    CaseGeneratorInput,
    SamplingPolicy,
    generate_dataset_sync,
    generated_batch_from_cases,
    samples_to_cases,
    write_generation_result,
)

review_cases = samples_to_cases(production_samples, policy=SamplingPolicy(max_samples=50))
generated = generated_batch_from_cases(
    synthetic_cases,
    generator_asset_version="prompt.generator@v4",
    model_provider="openrouter",
    model_name="openai/gpt-5.6-luna",
)

result = generate_dataset_sync(
    generate_cases,
    CaseGeneratorInput(seed=17, settings={"count": 20}),
    generator_id="generation:generate_cases",
    dataset_id="generated-routing",
    version="v1",
)
write_generation_result(result, Path("datasets/generated-routing.yaml"))
```

Production helpers normalize reviewed samples. Generated-dataset APIs own the typed preparation,
hashing, review projection, manifest, and safe publication boundary, while the application owns the
actual generator/provider logic. Generation finishes before normal benchmark planning. See
[Generated Datasets](generated-datasets.md).

## Extension Rules

- Put subject execution in a task.
- Put domain judgment in a Python scorer.
- Use a deriver for same-run computations and a post-deriver for matched runs.
- Use a policy for acceptance boundaries.
- Use an instrumentor for a stable SDK boundary.
- Use source maps and extractors for external field normalization.
- Use metric packs for reusable domain defaults.
- Never mutate recorded evidence; create a derived record or a new experiment.

See [API Reference](api-reference.md) for generated signatures and model fields.

---

## CLI

Canonical page: https://vcoderun.github.io/autobench/cli/

# CLI

The CLI is human-first. Commands render Rich panels and tables; YAML, CSV, and Markdown are explicit
file exports instead of raw terminal output.

## Command Map

| Command | Executes subject? | Input | Purpose |
| --- | --- | --- | --- |
| `validate` | No | benchmark YAML | Resolve and validate the planned matrix |
| `dataset generate` | Yes | generator target + request YAML | Prepare and freeze generated cases before planning |
| `run` | Yes | benchmark YAML | Execute, record, and render an experiment |
| `replay` | No | record directory | Reconstruct recorded results |
| `report` | No | record directory | Render Rich analysis or publish Markdown |
| `compare` | No | record directory | Compare two variants without causal claims |
| `export` | No | record directory | Write a YAML, CSV, or Markdown projection |
| `recording inspect` | No | staging directory | Diagnose committed, missing, corrupt, and conflicting evidence |
| `recording finalize` | No | staging directory | Publish complete or explicitly partial immutable evidence |
| `recording archive` | No | staging directory | Copy mutable staging for retention or investigation |
| `recording discard` | No | staging directory | Permanently remove a validated staging directory |
| `instrumentation doctor` | No | environment | Inspect integration compatibility |
| `instrumentation trace` | No | record directory | Summarize ABP traces and diagnostics |
| `telemetry export` | No | record directory | Replay immutable ABP evidence to OTLP traces |

## Validate

```bash
autobench validate benchmarks/routing.yaml
```

Validation parses the DSL, loads file or glob datasets, resolves pricing and task sources relative
to the spec, checks duplicate IDs and runnable requirements, and displays case, variant, and run
counts. It does not invoke the task.

## Generate A Dataset

```bash
autobench dataset generate generator:generate_cases \
  --request generation-request.yaml \
  --output datasets/generated.yaml \
  --id generated-routing \
  --version v1
```

The target is an importable sync or async callable that accepts `CaseGeneratorInput` and returns
`GeneratedCaseBatch`. Complete generation writes normal dataset YAML and a `.generation.yaml`
provenance manifest. Explicit incomplete generation writes only `.incomplete.yaml` and exits `2`;
it never publishes or replaces the requested dataset. `--force` replaces existing complete output
files but does not weaken that incomplete-result boundary. See
[Generated Datasets](generated-datasets.md).

## Run

```bash
autobench run benchmarks/routing.yaml \
  --concurrency 4 \
  --record runs/routing-42
```

Options:

| Option | Meaning |
| --- | --- |
| `--concurrency INTEGER` | Maximum active runs; default and minimum are `1` |
| `--record DIRECTORY` | Write immutable evidence to this new directory |
| `--no-record` | Execute and display without persistence |
| `--group-id TEXT` | Override the YAML correlation group |
| `--attempt INTEGER` | Override the positive invocation attempt |
| `--phase TEXT` | Override the invocation phase |
| `--parent-experiment-id TEXT` | Associate a prior experiment without creating replay lineage |
| `--resumed-from-experiment-id TEXT` | Record an external resume association |
| `--correlation-label KEY VALUE` | Merge one repeatable scalar label |

Without either recording flag, Autobench creates
`.autobench/<spec-stem>/<experiment-id>/`. `--record` and `--no-record` are mutually exclusive.

Correlation flags override `execution.correlation` field by field. Omitted flags preserve YAML
values, while repeated labels replace matching keys and retain the rest. These values group
independent invocations for analysis; they neither resume a task nor alter replay ancestry.

The CLI records the benchmark file and resolved referenced-source hashes so evidence can explain
what was executed. Recording is incremental: every completed run is committed before whole-matrix
post-processing. If execution stops, the Rich error output includes the sibling staging path.

`Ctrl-C` and `SIGTERM` use cooperative cancellation. Active runs finalize partial ABP state and
commit cancellation checkpoints before the command exits with status 130 or `128 + SIGTERM`.
Concurrent siblings receive cancellation and bounded cleanup time before the recorder is aborted.
The staging path printed by the CLI can then be inspected or finalized explicitly.

`SIGKILL` cannot run Python cleanup. After a hard kill, only completed runs and explicit
`await ctx.checkpoint(...)` calls already present in `staging-manifest.yaml` are recoverable. This
is a deliberate guarantee boundary, not an application-resume mechanism.

Interactive terminals also receive live Rich progress through the public `ProgressEvent` observer
API. The CLI explicitly selects best-effort delivery for this display: a renderer failure is written
to stderr and never replaces benchmark execution or recorder behavior. Python callers remain strict
by default so a lost lifecycle integration cannot pass silently.

## Recording Recovery

```bash
autobench recording inspect runs/.routing-42.staging

autobench recording finalize runs/.routing-42.staging \
  --output runs/routing-42-recovered \
  --allow-partial

autobench recording archive runs/.routing-42.staging \
  --output archives/routing-42

autobench recording discard runs/.routing-42.staging --yes
```

`inspect` displays whether the staging directory is recoverable as well as complete, checkpointed,
missing, corrupt, and conflicting identities. `finalize` is strict by default; `--allow-partial`
publishes explicit missing-run and incomplete-post-processing metadata rather than pretending the
matrix completed. Finalization does not delete the source staging directory.

`archive` and `discard` are separate operations. `discard` requires `--yes` and first verifies that
the path is an Autobench staging directory. A symlink or arbitrary directory is rejected.

## Replay

```bash
autobench replay runs/routing-42
```

Replay imports neither the task nor optional provider SDKs. It reconstructs normal result models
from `experiment.yaml`, per-run records, and referenced artifacts.

## Report

```bash
autobench report runs/routing-42
```

The Rich terminal report summarizes recorded execution evidence. The Markdown report is a separate
decision-facing projection: quality gate, scores, case outcomes, purposeful charts, comparisons,
and evaluator feedback. Missing metrics remain distinct from numeric zero.

Write the richer evidence-linked Markdown projection without printing Markdown to the terminal:

```bash
autobench report runs/routing-42 \
  --format markdown \
  --profile full \
  --layout auto \
  --output analysis/routing-report
```

Profiles are `summary`, `full`, and `audit`; layouts are `single`, `bundle`, and `auto`. Use `audit`
for technical runs, traces, assets, hashes, artifacts, and provenance. Captured audit detail
additionally requires `--include-captured-content`. See
[Markdown Reports](markdown-reports.md).

## Compare

```bash
autobench compare runs/routing-42 \
  --baseline current \
  --candidate candidate
```

Both IDs must exist. The view shows changed factors, aggregate metric values and deltas, paired run
count, and whether several factors changed. `confounded=true` is a warning against causal
attribution, not a failed comparison.

## Export

```bash
autobench export runs/routing-42 \
  --format yaml \
  --path analysis/routing-summary.yaml

autobench export runs/routing-42 \
  --format csv \
  --path analysis/routing-runs.csv

autobench export runs/routing-42 \
  --format markdown \
  --path analysis/routing-report.md
```

`--format` is required and accepts `yaml`, `csv`, or `markdown`. `--path` is also required. YAML
exports include a versioned schema header; CSV is a run-level projection; Markdown is a portable
single-file report written through the same atomic publisher. The complete evidence remains the
record directory.

## OTLP Telemetry Export

```bash
autobench telemetry export runs/routing-42 \
  --endpoint https://collector.example/v1/traces \
  --header authorization 'Bearer ...' \
  --service-name routing-benchmark
```

This command requires `autobench[otlp]`. It maps the experiment, runs, ABP trace hierarchy,
semantic evidence, source provenance, links, partial state, and record identities without changing
the source record. Captured content is omitted unless `--include-captured-content` is explicit.
See [OTLP Export](otlp-export.md) for the mapping and privacy contract.

## Instrumentation Doctor

```bash
autobench instrumentation doctor
```

The compatibility table shows distribution and version state, supported range, mechanism,
abstraction layer, sync/async/streaming support, asset discovery, capture defaults, optional extra,
and diagnostics for every built-in integration.

Use it before enabling `strict=True` or when an SDK upgrade stops producing evidence.

## Trace Inspection

```bash
autobench instrumentation trace runs/routing-42
```

This is replay-only. It summarizes span roots, kinds, instrumentors, partial state, and protocol
diagnostics without loading the original SDK.

## Exit And Failure Behavior

- Invalid YAML, unresolved sources, schema errors, recording collisions, and missing records exit
  nonzero.
- YAML failures include the file and source location when available.
- Task failures are isolated to their run; already collected evidence is preserved.
- An experiment can finish with passed, failed, errored, and skipped runs. Inspect status tables and
  policies rather than assuming process completion means every run passed.
- Replay and reporting never fall back to live execution.

## CI Workflow

```bash
set -e
autobench validate benchmarks/release.yaml
autobench run benchmarks/release.yaml \
  --concurrency 4 \
  --record artifacts/autobench
autobench report artifacts/autobench
autobench export artifacts/autobench \
  --format csv \
  --path artifacts/autobench-runs.csv
```

Upload the entire `artifacts/autobench` directory so replay, traces, asset histories, and source
lineage remain available.

---

## Capability Map

Canonical page: https://vcoderun.github.io/autobench/capabilities/

# Capability Map

This page is the inventory of what Autobench owns today. Every public feature belongs to one of
the layers below; application-specific behavior stays in tasks, scorers, adapters, and metric
packs.

## End-To-End Lifecycle

```text
BenchmarkSpec
  -> Dataset x Variants
  -> BenchmarkPlan
  -> Task(ctx, case)
  -> Observations + Spans + Artifacts + Errors
  -> Scores + Derived Metrics + Policies
  -> Cross-run Derivation
  -> Immutable RunRecord / ExperimentRecord
  -> Replay -> Report -> Compare -> Export -> Optimization Feedback
```

The same lifecycle is available through the YAML DSL, Python models, the `Benchmark` builder, and
the CLI. YAML is the portable authoring format; Python remains the extension surface for
application execution and custom evaluation logic.

## Definition And Data

| Capability | What it provides |
| --- | --- |
| `BenchmarkSpec` | Validated benchmark metadata, dataset, task, variants, scoring, derivation, policies, and reports |
| Dataset | Inline cases, file-backed datasets, glob-backed case files, defaults, tags, metadata, attachments, and versions |
| Cases | Arbitrary input and expected payloads with stable IDs and artifact references |
| Variants | Named factor combinations with labels, semantic types, and `optimize` hints |
| Generated datasets | Separate sync/async preparation API and CLI, typed requests/batches, review state, provenance, usage/cost, content hashes, complete publication, and incomplete sidecars |
| YAML schemas | Versioned JSON schemas and `yaml-language-server` headers for completion and validation |
| Source discovery | Hash collection for specs, datasets, pricing files, task modules, and scorer modules |

See [Datasets And Variants](datasets-and-variants.md) and [YAML Spec](yaml-spec.md).

## Planning And Execution

| Capability | What it provides |
| --- | --- |
| Matrix planning | Deterministic case x variant expansion and stable run IDs |
| Task runtime | Sync and async Python callables with `ctx` first and `case` second |
| Concurrency | Bounded async execution while preserving deterministic result ordering |
| Failure isolation | One task, scorer, derivation, or policy failure does not erase other runs |
| Progress events | Typed lifecycle events for runners and future UI integrations |
| Execution correlation | Immutable cross-invocation group, attempt, phase, association, and scalar-label metadata across Python, YAML, CLI, records, replay, and reports |
| Optional Pydantic Evals bridge | Internal conversion to Pydantic Evals-compatible case and dataset payloads |

See [Tasks And Runtime](tasks-and-runtime.md).

## Evidence Collection

| Capability | What it provides |
| --- | --- |
| Observations | Metrics, factors, events, diagnostics, artifacts, roles, units, directions, tags, and sources |
| Semantic registry | Canonical semantic types, aliases, parent relationships, and custom extensions |
| Projection | Source precedence and duplicate detection for one canonical metric view |
| Context spans | Nested agent, LLM, tool, retriever, parser, workflow, and custom spans |
| Automatic duration | Span timing and optional duration metrics owned by the runtime |
| Artifacts | Structured values and files materialized outside the main record payload |
| Errors | Structured task, scorer, trace, and policy errors with traceback capture |
| Measurement | Warmup, repetitions, time budgets, samples, median, p95, standard deviation, and noise |

See [Observations And Semantics](observations-and-semantics.md) and
[Instrumentation And Traces](instrumentation-and-traces.md).

Native Pydantic AI, pydantic-gepa, OpenAI, OpenAI Agents, and HTTPX integrations can be selected
through typed Python settings or the YAML `instrumentation` section. They emit ABP directly,
compose across optimizer/framework/client/transport layers, preserve lifecycle, and remain optional
for replay. See [Native Instrumentation](native-instrumentation.md) and
[Pydantic-GEPA Instrumentation](pydantic-gepa-instrumentation.md).

Semantic instrumentors automatically discover SDK-visible prompt, tool, output-schema, capability,
agent, guardrail, handoff, policy, and toolset versions. Definition/effective relationships,
capability scopes, aliases, privacy-controlled content, and span-local `AssetUse` evidence survive
recording and replay. HTTPX remains transport evidence and performs no semantic asset inference.
See [Automatic Asset Discovery](automatic-asset-discovery.md).

Immutable records can also be replayed through the optional outbound
[OTLP exporter](otlp-export.md). Experiment/run/ABP hierarchy, semantic events, record identity,
partial state, links, and source provenance are preserved without making OTel canonical or a base
dependency.

## Scoring And Constraints

| Scorer | Purpose |
| --- | --- |
| `output` | Project an output path into a semantic score |
| `pass_fail` | Turn a boolean output path into a pass/fail score |
| `exact` | Compare actual and expected paths |
| `schema` | Validate output against a schema/model |
| `python` | Run a sync or async custom scorer using `ScoringCall` |
| `expected_action` | Evaluate action/tool selection, arguments, or sequence from spans |

Scores declare semantic type, unit, direction, role, and optional failure behavior. Policies add
typed requirements including equality, membership, numeric bounds, and inclusive ranges.

See [Scoring And Derivation](scoring-and-derivation.md) and
[Agentic Evaluation](agentic-evaluation.md).

## Derivation And Cost

| Capability | What it provides |
| --- | --- |
| Token cost | Derive `money.cost` from input/output tokens and normalized model/provider factors |
| Pricing DSL | Static YAML pricing, aliases, provider maps, cache prices, token tiers, and model normalization |
| Price sources | Optional llm-prices and genai-prices importers that normalize external data into `PricingTable` |
| Paired baseline | Per-case or factor-matched speedup, delta, percent change, diagnostics, and verdicts |
| Comparison classifier | Improved, regressed, unchanged, or inconclusive outcomes with relative noise thresholds |

External price sources are convenience importers, not runtime dependencies or Autobench's source
of truth. A local pricing YAML remains fully supported.

## Agentic Evidence

Autobench records agent behavior without requiring OpenTelemetry:

- typed trace envelopes and nested span records
- expected tool/action selection, argument, and sequence checks
- span selectors by kind, name, tag, path, or semantic type
- Pydantic AI usage normalization
- metric packs for agentic, structured-output, LLM-usage, and performance defaults
- compact feedback records for optimization systems

See [Agentic Evaluation](agentic-evaluation.md).

## Asset Lineage

The tracking registry understands:

- text prompts from inline text or files
- arbitrary assets and configuration values
- callable tools, signatures, parameters, docs, and return types
- Pydantic models, standard dataclasses, and typed classes
- field names, annotations, descriptions, aliases, defaults, requirements, constraints, and examples
- source hashes, structured-schema hashes, versions, parent versions, and diffs
- persistent human-readable YAML asset histories
- automatic SDK-boundary discovery without tracking decorators
- source/effective representation links, capability scopes, provenance, and cross-layer aliases
- automatic experiment persistence and replayable span-local asset uses

Decorators preserve the original callable or class type so tracking does not degrade static
typing. See [Asset Tracking](asset-tracking.md) and
[Automatic Asset Discovery](automatic-asset-discovery.md).

## Records, Replay, And Analysis

| Capability | What it provides |
| --- | --- |
| `RunRecord` | Immutable case x variant evidence including output, scores, observations, spans, factors, assets, artifacts, and errors |
| `ExperimentRecord` | Plan, environment, semantic registry, report config, source hashes, and run paths |
| Correlated reports | Field filters and `group_id` report groups across independent experiment results |
| Replay | Load records without importing task or scorer modules |
| Rich terminal reports | Status, variant configuration, leaderboard, run metrics, case matrix, comparisons, and distributions |
| Markdown reports | Decision-facing quality gates, case outcomes, evaluator feedback, purposeful inline SVG, paired comparisons, summary/full/audit profiles, audit-only traces/assets/provenance, single/bundle/auto layouts, and atomic publication |
| Optimizer reports | pydantic-gepa outcome/resources, engine branches, candidate lineage, component versions, selections, and diagnostics |
| Exports | Human-readable YAML summary, CSV run projection, and Markdown report |
| Optimization feedback | Failure category, score, reasons, factors, asset versions, and selected evidence |

See [Recording And Reporting](recording-and-reporting.md) and
[Markdown Reports](markdown-reports.md).

## Ownership Boundaries

Autobench deliberately does not own:

- application or model execution
- hosted tracing or observability storage
- model-specific pricing as an always-current service
- causal claims from confounded comparisons
- optimizer search strategies or candidate promotion
- large catalogs of domain-specific LLM judges

Tasks and adapters own application execution. Optional integrations may import traces, pricing, or
evaluator results, but the core contract remains semantic, generic, and replayable.

---

## API Reference

Canonical page: https://vcoderun.github.io/autobench/api-reference/

# API Reference

This reference is generated from the installed public `autobench` package. The root package is the
supported import surface; subpackages organize implementation and extension areas.

## Public Package

::: autobench
    options:
      members: true
      members_order: source
      show_root_heading: true
      show_source: true
      show_signature_annotations: true
      separate_signature: true
      heading_level: 3

## Progress Runtime

- `ProgressEvent` carries a monotonic `sequence`, stable benchmark/run identity, event-specific
  `data`, and typed `run_status` or `experiment_status` terminal fields.
- `ProgressEventKind` defines benchmark start/finish, run start/finish, and actual policy violation
  events. Candidate decisions belong to Autoptimize rather than the benchmark lifecycle.
- `ProgressHandler` accepts synchronous and asynchronous observers.
- `ProgressErrorPolicy` selects strict library delivery or explicit best-effort delivery.
- `ProgressHandlerFailure` identifies the handler index, event sequence/kind, and original error.
- `ProgressDispatchError` is raised after strict delivery failure, terminal notification attempts,
  and durable recorder cleanup.

`run_benchmark_spec()`, `run_benchmark_path()`, `Benchmark.run()`, and `Benchmark.run_async()` all
accept `progress_handlers`, `progress_error_policy`, and `progress_error_handler`.

## Public Areas

| Area | Representative symbols |
| --- | --- |
| Definition | `Benchmark`, `BenchmarkSpec`, `TaskSpec`, `load_benchmark_spec` |
| Data | `Case`, `DatasetSpec`, `Variant`, `FactorValue`, production helpers, typed generated-dataset preparation and provenance |
| Runtime | `RunContext`, `RunPhase`, `Span`, `ExperimentResult`, `run_benchmark_spec`; durable `await ctx.checkpoint(name)` and cooperative cancellation |
| Semantics | `Observation`, `Semantic`, `SemanticRegistry`, queries and projection |
| Evaluation | scorers, derivers, policies, expected actions, measurement, feedback |
| Protocol | ABP signals, traces, capture, emitter, collector and context |
| Instrumentation | settings, manager, instrumentors, method instrumentation, diagnostics, `CurrentSpan`, keyed `InstrumentationRuntime.start_span()` / `end_span()`, and external backend composition |
| Tracking | `track`, asset models, discovery candidates, registry and history views |
| Records | `RunRecord`, `ExperimentRecord`, `ExperimentTermination`, `RecordManifest`, `FileRecorder`, frozen staging snapshots, inspection/recovery, atomic/synced publication, and replay helpers |
| Reports | report models, builders, Rich renderers and exporters |
| Telemetry export | `OTLPSettings`, `OTLPExportResult`, `export_otlp`, and `export_record_otlp` |

`ExecutionCorrelation` is the public invocation metadata model. It is accepted by the full spec,
fluent builder, run functions, and CLI; persisted on results and records; and consumed by
`correlation_matches()`, `filter_experiments()`, and `build_grouped_reports()`. It remains separate
from replay lineage and application workflow state.

Prefer root imports for application code:

```python
from autobench import Benchmark, Case, RunContext, Semantic
```

Native pydantic-gepa integration also exposes `PydanticGEPAInstrumentation`, `PydanticGEPA`, and
the replay-safe `PydanticGEPAEvidence` projection from the root package. Its full contract is
documented in [Pydantic-GEPA Instrumentation](pydantic-gepa-instrumentation.md).

Import a submodule when implementing an extension against that subsystem, such as a custom native
instrumentor or source-map adapter.

---

## Troubleshooting

Canonical page: https://vcoderun.github.io/autobench/troubleshooting/

# Troubleshooting

Start with the narrowest command that can identify the failing layer.

## Task Module Cannot Be Imported

```text
Could not import task module 'benchmarks.tasks'
```

Task targets must use `module:function`, not a file path:

```yaml
run:
  python: benchmark_task:run
```

Autobench first uses normal Python imports, then searches relative to the benchmark spec. Common
fixes:

- place `benchmark_task.py` next to `autobench.yaml` and use `benchmark_task:run`;
- for a package, ensure package directories have the expected Python import structure;
- do not include `.py` in the target;
- run `autobench validate path/to/autobench.yaml` from any directory to test resolution.

The function must accept `(ctx, case)` in that order.

## YAML Validates In One Editor But Not In Autobench

The schema directive improves editor completion; Autobench's installed Pydantic models remain the
runtime authority. Match the schema version to the installed package:

```bash
python -c "import autobench; print(autobench.__version__)"
```

```yaml
# yaml-language-server: $schema=./schemas/0.3.0/benchmark_schema.json
```

Run `autobench validate` and use its file/line diagnostics. Unknown scorer, policy,
instrumentation, and capture fields are rejected intentionally.

## Dataset File Is Not Found

`file://` references and glob patterns resolve relative to the benchmark YAML, not the shell's
current directory:

```yaml
dataset:
  source: file://datasets/cases.yaml
```

For a glob:

```yaml
dataset:
  source: file://datasets/cases/*.yaml
```

An unmatched glob is an error. Dataset files can contain a dataset DSL document, a case list, or a
single case mapping.

## Record Directory Already Exists

Autobench records are immutable. `record_experiment()` and `autobench run --record` refuse to
overwrite an existing experiment:

```text
Experiment record already exists
```

Use a new directory or remove/archive the old directory explicitly outside Autobench. Do not merge
unrelated experiments by copying run files together.

## A Metric Is Missing From Reports

Reports query semantic types, not only local names. Check:

1. the task/scorer/deriver emitted the observation;
2. `semantic_type` matches the report metric;
3. the value is numeric or boolean for the selected aggregate;
4. the selected span/query is not filtering it out;
5. source precedence did not intentionally select a score or derived observation instead.

Inspect the per-run YAML or use Python:

```python
from autobench import ObservationQuery

query = ObservationQuery(observations=run.task_result.observations)
matches = query.exact("money.cost")
```

Missing cost is not converted to zero. Ensure token, model, provider, and pricing inputs are all
available to the token-cost deriver.

## Paired Baseline Does Not Produce A Value

The baseline and candidate must match on the configured key, normally `case_id`, and both must have
the source metric. Check:

- `baseline_variant` exactly matches a variant ID;
- both runs emit the same semantic metric;
- the metric unit is compatible;
- `match_on` identifies a unique counterpart;
- the configured missing policy is appropriate.

Use a case matrix for the source metric before debugging the formula.

## `instrument_all()` Records Skipped Integrations

This is normal when optional SDKs are not installed. Automatic discovery records
`instrumentation.skipped` diagnostics and continues by default.

```bash
autobench instrumentation doctor
```

Install the relevant extra, remove the integration from `exclude`, or use `strict=True` when absence
must fail the benchmark.

Explicit `enabled: false` wins over automatic discovery. A custom runtime instrumentor with the
same ID also prevents a duplicate built-in installation.

## No Automatic Assets Appear

Automatic discovery only observes values that cross a supported instrumented SDK boundary while a
benchmark run is active. Check:

- the corresponding instrumentor is compatible and installed;
- asset discovery is enabled;
- `include` contains the expected family;
- `representations` includes `definition` or `effective` as needed;
- the SDK call occurs inside the task;
- capture policy does not reduce the asset below the expected content level.

Run the offline `examples/automatic_assets/` programs to separate environment issues from
application behavior.

## Pydantic-GEPA Optimization Evidence Is Missing

The observer records only while an Autobench run context is active. Check:

- `autobench[pydantic-gepa]` is installed on Python 3.11-3.13;
- `autobench instrumentation doctor` reports event contract version `1` as compatible;
- the YAML contains `instrumentation.pydantic_gepa` or `instrument_all()` selected it;
- the optimization call occurs inside `task(ctx, case)`;
- the instrumentor is not disabled or suppressed;
- the record contains the `autobench.pydantic_gepa/v1` extension.

Use `detail: evaluations` or `detail: full` when per-candidate and per-case spans are expected.
`summary` intentionally omits those spans but still records budgets, selections, candidates, and
the durable projection. The complete workflow is in
[Pydantic-GEPA Instrumentation](pydantic-gepa-instrumentation.md).

## Duplicate Or Conflicting Instrumentation

Autobench prevents unsafe double patching. Do not install two instrumentors with the same ID or
instrument the same owner/method with incompatible specs. Prefer one of:

- automatic discovery only;
- explicit typed settings only;
- a custom runtime instrumentor that owns the same ID.

`InstrumentationConflictError` and patch diagnostics identify the owner and method involved.

## Trace Is Partial

A partial trace can be valid evidence. It may result from cancellation, an interrupted stream, a
task exception, or unmatched start/end signals. Inspect:

```bash
autobench instrumentation trace runs/example
```

ABP materialization keeps completed spans and diagnostics instead of dropping the trace. Accounting
extractors avoid double counting aggregate and leaf usage even when evidence is incomplete.

## Captured Content Is Missing Or Hashed

Runtime evidence is metadata-first, while behavioral asset definitions are full by default so
historical candidates remain reconstructable. Both are controlled by the benchmark
`CapturePolicy`, path rules, semantic overrides, and SDK-specific HTTP settings.

To avoid retaining asset bodies, use a preset that changes both defaults or set the asset fallback
explicitly:

```python
from autobench import CaptureLevel, CapturePolicy

policy = CapturePolicy.hashed(
    semantic_overrides={"output_schema": CaptureLevel.FULL},
)
```

The experiment-local bodies and readable diffs are in `artifacts/asset-content.sqlite3`;
`assets/*.yaml` contains typed references and changed paths rather than copies. Secret names,
denied paths, truncation limits, and binary rules still apply.

## Replay Needs An Optional SDK

It should not. `replay`, `report`, `compare`, `export`, and `instrumentation trace` are designed to
load records without benchmark or provider imports. If replay fails, verify that:

- `experiment.yaml` and every path in `runs.paths` exist;
- trace artifact paths remain inside the experiment directory;
- referenced artifacts were copied with the records;
- the record version is supported.

Do not solve a missing artifact by re-executing the benchmark implicitly.

## Python Type Errors In Tasks

Autobench keeps case input and factor values generic because applications define their schemas.
Validate at the task boundary:

```python
from pydantic import BaseModel, TypeAdapter

request = Request.model_validate(case.input)
mode = TypeAdapter(Mode).validate_python(ctx.factor("mode"))
```

This gives application-specific errors without weakening Autobench's public types.

## Get A Reproducible Diagnostic Bundle

For a bug report, include:

```bash
autobench --help
autobench instrumentation doctor
autobench validate path/to/autobench.yaml
```

Also include the package version, Python version, failing record directory when it contains no
sensitive data, and the smallest benchmark/task that reproduces the problem. Review capture policy
before sharing records.

---

## Development

Canonical page: https://vcoderun.github.io/autobench/development/

# Development

Autobench uses `uv` for dependency management and exposes stable repository operations through
the Makefile.

## Environment

```bash
uv sync --extra dev
```

The committed lock file is the reproducible dependency contract used by CI.

## Quality Gates

```bash
make format
make prod
make pre-commit
```

`make prod` runs the test suite, enforces `100%` line and branch coverage, checks formatting and
typing, builds the documentation, validates Python 3.11 through 3.13, and executes the offline
examples end to end.

## Documentation

The site uses Zensical's modern theme while retaining `mkdocs.yml` as the supported migration
configuration format.

```bash
make docs
make docs-serve
```

Pushes to `main` build the site in strict mode. The workflow stores generated files in the
`gh-pages` branch and deploys the same artifact through GitHub Pages Actions.

## Release Artifacts

```bash
make build
```

The build produces a wheel and source distribution under `dist/`. Generated documentation,
benchmark runs, internal planning files, references, and agent instructions are excluded from the
published package.

---

## 0.2.0

Canonical page: https://vcoderun.github.io/autobench/release-notes/0.2.0/

# 0.2.0

Autobench `0.2.0` introduces the Autobench Instrumentation Protocol (ABP) and native evidence
collection for supported AI and HTTP SDKs.

## Included

- immutable ABP signals, traces, scopes, links, references, and protocol diagnostics
- task-local trace context with correct concurrent parentage and partial-run preservation
- privacy-first capture policy with redaction, truncation, hashing, and artifact references
- versioned semantic source maps and accounting-safe trace extraction
- native Pydantic AI, OpenAI Python, OpenAI Agents, and HTTPX instrumentors
- automatic prompt, tool, output-schema, capability, agent, guardrail, handoff, policy, and toolset
  discovery at supported semantic SDK boundaries
- definition/effective asset links, capability scopes, cross-layer aliases, and conservative
  duplicate correlation
- automatic experiment asset persistence with atomic worker-safe history merges and replayed
  `AssetUse` lineage
- `InstrumentAssetSpec` for custom SDK asset extraction without tracking decorators
- benchmark-level typed/YAML capture policy for privacy-controlled evidence and asset persistence
- sync, async, iterator, context-manager, and streaming lifecycle preservation
- typed fluent and YAML instrumentation configuration with versioned schema completion
- `autobench instrumentation doctor` compatibility diagnostics
- `autobench instrumentation trace` replay-only trace summaries
- real offline instrumentation, layering, and replay/extraction examples
- Python 3.11, 3.12, 3.13, and 3.14 quality matrix
- built-wheel/no-extras and target-library compatibility gates

## Compatibility

Existing `0.1.0` benchmark specs and RunRecords remain loadable. ABP evidence is additive: manual
spans and method instrumentation now materialize through the same protocol used by native
instrumentors. Replaying ABP records does not import optional provider SDKs.

## Intentionally Deferred

- OTLP and vendor exporters
- distributed context propagation
- import-hook auto-instrumentation
- execution cassette replay
- visualization

ABP protocol version `1` is the initial public protocol. Autobench `0.2.x` will preserve its
serialized meaning; additive fields remain forward-compatible through extension maps. A breaking
wire-format change requires a new ABP protocol version.

---

## 0.1.0

Canonical page: https://vcoderun.github.io/autobench/release-notes/0.1.0/

# 0.1.0

Autobench `0.1.0` is the first release-shaped core.

## Included

- YAML-first benchmark specs
- deterministic task runtime
- semantic observations and projection
- scoring, derivation, post-derivation, and policies
- immutable YAML recording and replay
- Markdown, YAML, and CSV reporting
- offline minimal, basic, mid, and advanced examples
- real optional CodeMode dogfood integration
- portable CLI source provenance
- Python 3.11, 3.12, and 3.13 quality matrix

## Intentionally Not Included

- autoptimize orchestration
- GEPA integration
- OpenTelemetry bridge
- hosted dashboard features
- distributed execution
- full Pydantic Evals dataset/evaluator execution

`PydanticEvalsBridge` in this release is an optional payload and availability bridge. It does not
claim to execute Pydantic Evals datasets. The full internal evaluation runtime remains a later
integration milestone.
