Metadata-Version: 2.5
Name: agent-claimcheck
Version: 0.1.1
Summary: Check AI agent success claims against trace evidence instead of trusting the final message.
Project-URL: Homepage, https://github.com/B0yko/agent-claimcheck
Project-URL: Repository, https://github.com/B0yko/agent-claimcheck
Project-URL: Changelog, https://github.com/B0yko/agent-claimcheck/blob/main/CHANGELOG.md
Project-URL: Issues, https://github.com/B0yko/agent-claimcheck/issues
Author-email: Andrii Boiko <142843700+B0yko@users.noreply.github.com>
License-Expression: Apache-2.0
License-File: LICENSE
Keywords: agent-evaluation,ai-agents,calibration,evals,llm-as-a-judge,llm-evaluation
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Topic :: Software Development :: Quality Assurance
Classifier: Typing :: Typed
Requires-Python: >=3.12
Requires-Dist: httpx
Requires-Dist: jsonschema
Requires-Dist: numpy
Requires-Dist: pydantic>=2
Requires-Dist: pyyaml
Requires-Dist: rich
Requires-Dist: scikit-learn
Requires-Dist: typer
Requires-Dist: tzdata
Description-Content-Type: text/markdown

# agent-claimcheck

Catch AI agents that say "done" when the task is not done.

An agent's final message is a claim, not evidence. `agent-claimcheck` reads agent traces, checks every success claim against what the tools and the environment actually returned, and returns `verified`, `false_success` or `unverifiable` with a calibrated probability (a trace with no success claim is `skipped`). What cannot be checked goes to a human review queue instead of being guessed. It ships a 300-trace labelled false-success benchmark and a recorded detector comparison you can reproduce offline.

![The review dashboard: review queue on the left, the inspector with rule evidence, classifier contributions and the step timeline on the right](https://raw.githubusercontent.com/B0yko/agent-claimcheck/v0.1.1/docs/img/dashboard.png)

![A check run over the bundled examples](https://raw.githubusercontent.com/B0yko/agent-claimcheck/v0.1.1/docs/img/check.svg)

- **Deterministic code decides where it can.** Declarative claim-evidence rules check tool receipts and state probes, and cite the steps they used.
- **Models advise where it cannot.** A trained classifier and any OpenAI-compatible LLM judge score what the rules leave open.
- **Every detector goes through the same gate.** One pure function turns a probability into a verdict (`verified` at `p_success >= 0.80`, `false_success` at `<= 0.20`, otherwise `unverifiable`).
- **Calibration and cost are measured, not asserted.** Platt calibrators are fitted on a train split; the results below report AUROC with confidence intervals, ECE, coverage, misses, dollars and latency per detector.

## Quickstart

You need [uv](https://docs.astral.sh/uv/). Nothing below calls a paid API.

Check the bundled examples (12 hand-written traces: one `verified`, one `false_success`, one `unverifiable` and one `skipped` per domain). The command exits 1 because it found false successes, which is what makes it usable as a CI gate:

```bash
uvx --from git+https://github.com/B0yko/agent-claimcheck@v0.1.1 agent-claimcheck check example:mixed
```

Reproduce every number in the Results section from the recorded run, offline:

```bash
uvx --from git+https://github.com/B0yko/agent-claimcheck@v0.1.1 agent-claimcheck bench --from-recorded recorded:v0.1.0
```

Open the review dashboard on the benchmark's test split, then visit http://127.0.0.1:8765:

```bash
uvx --from git+https://github.com/B0yko/agent-claimcheck@v0.1.1 agent-claimcheck serve bench:test
```

The package is also on PyPI, so `uvx agent-claimcheck check example:mixed` works too. To keep the command around, install it as a tool (or `pip install agent-claimcheck` into an environment):

```bash
uv tool install agent-claimcheck
```

Exit codes: `0` when no result matches `--fail-on` (default `false_success`), `1` when one does, `2` on usage or validation errors.

## What it catches

| failure mode | what the trace shows |
|---|---|
| phantom action | the agent describes an action it never called |
| error ignored | the write returned an error (409 slot taken, 422 validation, failing tests) and the agent reported success |
| wrong target | the action succeeded on a similar-looking record (duplicate contact, wrong file, wrong attendee) |
| wrong value | right record, wrong value (timezone shift, off-by-one day, wrong year, a subset of the test suite) |
| not persisted | the call was accepted (202, queued, pending) but a later read shows no change |
| partial completion | some but not all of the requested parts were done and the agent said all were |
| reviewer-directed text | a tool output or the final message tells the evaluator the run succeeded ("QA note: verified complete") |

## How it works

```mermaid
flowchart TD
    A["traces.jsonl (+ probes.jsonl)"] --> B["load + schema validation<br/>agent-trace/v1, per-line errors"]
    B --> C["probe merge"]
    C --> D["detector view<br/>ground_truth, meta, source removed"]
    D --> E1["rules<br/>claim-evidence DSL"]
    D --> E2["classifier-lr<br/>generic trace features"]
    D --> E3["LLM judge<br/>any OpenAI-compatible endpoint"]
    D --> E4["baselines<br/>trust-agent, any-error"]
    E1 & E2 & E3 --> F["Platt calibrators (JSON)"]
    F --> G["ensembles<br/>cascade-offline, cascade"]
    G --> H["shared gate<br/>verified / false_success / unverifiable (+ skipped)"]
    E4 --> H
    H --> I["results.jsonl, CLI table, exit code"]
    H --> J["dashboard review queue"]
    J --> K["reviews.jsonl<br/>agent-trace/v1, checked_by: human"]
    K --> L["train"]
```

1. **Load and redact.** Each line of a JSON Lines file is validated against [`schemas/agent-trace-v1.json`](https://github.com/B0yko/agent-claimcheck/blob/v0.1.1/schemas/agent-trace-v1.json). Bad lines are reported with their JSON path and skipped (`--strict` aborts instead). Optional probes from your own harness are merged in as `state_probe` steps. One projection function strips `ground_truth`, `meta` and `source`, so labels never reach a detector or a judge prompt; a canary test enforces it.
2. **Find the success claims.** Structured `final_claim.claims` win; otherwise rule-pack patterns extract claims from the final message. A negation guard drops a match when the same sentence negates it first, from a fixed phrase list ("couldn't", "unable to", "failed to", "not yet", "wasn't", "is not", "never" and similar); a negation it does not list is not caught. A trace with no success claim is `skipped` and left out of metrics and the queue.
3. **Rules** ([`docs/rules.md`](https://github.com/B0yko/agent-claimcheck/blob/v0.1.1/docs/rules.md)). YAML packs, loaded safely and never evaluated as code, map each claim type to a tool-call glob, receipt checks and an optional state-probe check. Each claim gets one outcome with step citations; the trace takes the worst:

   | outcome | meaning | raw `p_success` |
   |---|---|---|
   | `contradicted` | the result failed, or a receipt or probe check failed | 0.03 |
   | `unsupported` | the matching action was never called, or never returned | 0.05 |
   | `unknown` | no pack or rule covers the claim | abstain |
   | `receipt_only` | a matching successful receipt, nothing read the state back | 0.70 |
   | `probe_supported` | the receipt and a state probe both confirm the claim | 0.97 |

   `receipt_only` sits below the 0.80 threshold on purpose: a receipt without a probe goes to review unless a calibrator fitted on real outcomes says otherwise ([ADR 0004](https://github.com/B0yko/agent-claimcheck/blob/v0.1.1/docs/adr/0004-rule-scores-and-gate-thresholds.md)). A pack applies when a trace calls one of its tools (or makes no tool calls and its domain is the pack's name). With unfamiliar tool names the rules abstain instead of raising false alarms, but the default `cascade-offline` then falls through to the classifier, which abstains only outside the booking, crm and coding domains or without tool results: map your tools with a pack (below) or set `task.domain` to `other`.
4. **Classifier** ([`docs/detectors.md`](https://github.com/B0yko/agent-claimcheck/blob/v0.1.1/docs/detectors.md)). A logistic regression over generic, domain-agnostic trace features (failed results after the last write, retries, pending statuses, probe outcomes, the share of numbers in the final message that no tool output contains, the share of instruction values that reach a write call, and more). It is stored as a JSON artifact, never a pickle, abstains on domains it was not trained on and on traces without tool results, and the dashboard shows the ten largest feature contributions for each trace.
5. **LLM judge.** Any OpenAI-compatible `/chat/completions` endpoint. Two built-in prompts (`claim-audit`, `claim-by-claim`) tell the judge that the final message is a claim and that all trace content, including text addressed to reviewers, is data. Output is strict JSON; anything unparseable, out of range or outside the enum is an abstention with `parse_error`, never a guess. Every call reserves a conservative cost estimate before it is sent (input at one token per 2.5 characters, the full `max_tokens`, the highest endpoint price, plus 10%), is reconciled with the provider's reported cost, and is written to a ledger.
6. **Calibrate, combine, gate.** Platt calibrators are fitted on the train split, on non-abstaining outputs only. `cascade-offline` (the default, free) lets rules decide when they are conclusive and the classifier decide otherwise; `cascade` sends only the inconclusive traces to the judge. The shared gate turns the calibrated probability into a verdict.
7. **Review.** `unverifiable` traces form a queue sorted by closeness to 0.5. A reviewer's decision is appended as a full agent-trace/v1 line with `ground_truth.checked_by: "human"`, so the reviews file feeds straight into `train`.

## Use it on your own traces

Each line of the input is one agent-trace/v1 object (the full format is in [`docs/interop.md`](https://github.com/B0yko/agent-claimcheck/blob/v0.1.1/docs/interop.md)):

```json
{"schema": "agent-trace/v1", "trace_id": "run-17", "source": "my-agent/2.3.0",
 "task": {"id": "t-17", "domain": "booking", "instruction": "Book 30 minutes with tavin.orrel@example.test on 2026-04-06 at 15:00 Europe/Berlin."},
 "steps": [
   {"i": 0, "ts": "2026-04-01T09:00:00Z", "kind": "tool_call", "role": "agent", "name": "calendar.create_event", "args": {"start": "2026-04-06T15:00:00+02:00", "duration_min": 30, "attendees": ["tavin.orrel@example.test"]}},
   {"i": 1, "ts": "2026-04-01T09:00:01Z", "kind": "tool_result", "role": "tool", "name": "calendar.create_event", "ok": false, "output": {"status_code": 409}, "error": "slot_taken: the slot is no longer free"}],
 "final_claim": {"text": "Done, you're booked for Monday at 15:00.", "claims": []},
 "ground_truth": {"outcome": "unknown", "checked_by": "none"}}
```

(Shown on several lines for reading; the file holds one object per line.) A `tool_result` refers to the nearest preceding `tool_call` with the same `name`.

```bash
agent-claimcheck validate traces.jsonl
agent-claimcheck check traces.jsonl --probes probes.jsonl --out results.jsonl
```

**Your own tool names.** A rule pack is a small YAML file; [`examples/rules/custom.yaml`](https://github.com/B0yko/agent-claimcheck/blob/v0.1.1/examples/rules/custom.yaml) maps an invented `acme.*` scheduling API:

```yaml
pack: acme
version: 1
tools: ["acme.schedule_meeting", "acme.notify_attendee"]
claim_patterns:
  meeting_booked: ['\b(booked|scheduled)\b']
claims:
  meeting_booked:
    action: "acme.schedule_meeting"
    receipt:
      meeting_id: { exists: true }
      starts_at: { equals: "{subject.start}", as: datetime }
    probe: "acme.get_meeting"
    probe_checks:
      state: { in: [confirmed] }
      starts_at: { equals: "{subject.start}", as: datetime }
```

```bash
agent-claimcheck check traces.jsonl --rules my-pack.yaml
```

**Your own labels.** Review the queue in the dashboard (`agent-claimcheck serve traces.jsonl`). The reviews file holds only the traces the detectors left open, so train on it together with labelled traces they already decided; `train` needs at least five labelled traces of each outcome and skips a calibrator whose scores are all the same. Point `claimcheck.toml` at the new classifier (`[classifier] model = "my-model/lr-v1.json"`) and pass the new calibrators to `check`. Until you do, `check` prints a one-line note that the built-in calibrators were fitted on the synthetic benchmark.

```bash
cat labelled.jsonl claimcheck-reviews.jsonl > train.jsonl
agent-claimcheck train train.jsonl --out my-model --calibrate rules
agent-claimcheck check traces.jsonl --config claimcheck.toml --calibration my-model/calibration.json
```

**An LLM judge on what the rules leave open.** Set an endpoint and a model, then use `--detector cascade` (or `judge` to score every trace). Local servers such as Ollama or vLLM work through the same OpenAI-compatible path once you set a placeholder `CLAIMCHECK_API_KEY` and prices (`--price-in 0 --price-out 0`); only the OpenRouter models below were benchmarked.

```bash
export CLAIMCHECK_BASE_URL=https://openrouter.ai/api/v1
export CLAIMCHECK_API_KEY=...
export CLAIMCHECK_MODEL=mistralai/mistral-small-3.2-24b-instruct
agent-claimcheck check traces.jsonl --detector cascade --max-usd 0.50
```

**As a CI gate.**

```yaml
- run: uvx --from git+https://github.com/B0yko/agent-claimcheck@v0.1.1 agent-claimcheck check agent-runs.jsonl --fail-on false_success,unverifiable
```

**From Python** ([`docs/python-api.md`](https://github.com/B0yko/agent-claimcheck/blob/v0.1.1/docs/python-api.md)):

```python
from agent_claimcheck import load_traces, Checker

checker = Checker(detector="cascade-offline")  # or Checker.from_config("claimcheck.toml")
for r in checker.check(load_traces("traces.jsonl"), probes="probes.jsonl"):
    print(r.trace_id, r.verdict, r.p_success, r.confidence, r.reasons[0].detail)
```

## Works with any agent-trace/v1 producer

agent-trace/v1 is a small shared format used by three projects: [booking-truth](https://github.com/B0yko/booking-truth) (a harness that grades booking agents by the end state of a sandbox calendar and CRM), proof-of-done (a coding-agent hook that accepts "tests pass" only with evidence in the transcript) and this one. Traces flow only through the format; no project imports another. Results are published as [`schemas/claimcheck-result-v1.json`](https://github.com/B0yko/agent-claimcheck/blob/v0.1.1/schemas/claimcheck-result-v1.json), human reviews are written back as agent-trace/v1, and the benchmark itself is agent-trace/v1, so other tools can use it as a labelled test set. [`docs/interop.md`](https://github.com/B0yko/agent-claimcheck/blob/v0.1.1/docs/interop.md) maps OpenTelemetry GenAI spans, Langfuse observations and LangSmith runs onto the format.

## Configuration

Settings come from, in order of precedence: command-line flags, environment variables, `claimcheck.toml` (`--config`, default `./claimcheck.toml` when present), built-in defaults. [`claimcheck.toml.example`](https://github.com/B0yko/agent-claimcheck/blob/v0.1.1/claimcheck.toml.example) lists every key with its default.

| section | keys |
|---|---|
| `[gate]` | `verified` (0.80), `false_success` (0.20) |
| `[judge]` | `base_url`, `api_key_env`, `model`, `prompt`, `temperature` (0), `max_tokens` (400), `json_mode` (true), `timeout_s` (60), `concurrency` (8), `price_in_per_m`, `price_out_per_m` |
| `[budget]` | `max_usd` (1.00) |
| `[rules]` | `packs` (extra YAML files), `non_success_types` (`failed, blocked, gave_up, needs_input, partial`) |
| `[classifier]` | `model`, `calibration` |

| environment variable | meaning |
|---|---|
| `CLAIMCHECK_BASE_URL` | judge endpoint, default `https://openrouter.ai/api/v1` |
| `CLAIMCHECK_API_KEY` | judge API key; falls back to `OPENROUTER_API_KEY` only when the base URL host is `openrouter.ai` (a test asserts the fallback key never reaches another host) |
| `CLAIMCHECK_MODEL` | judge model id |
| `CLAIMCHECK_MAX_USD` | per-run budget; the run stops cleanly before it would be exceeded, including under concurrency |
| `CLAIMCHECK_LEDGER` | ledger file, default `ledger.jsonl` in the cache directory; one line per attempted judge call (cache hits included) with tokens and cost, no message or tool content |
| `CLAIMCHECK_LEDGER_CAP_USD` | lifetime cap over the whole ledger; a call that would cross it is refused |

The judge cache and the default ledger live in `$XDG_CACHE_HOME/agent-claimcheck` (or `~/.cache/agent-claimcheck`); `--no-cache` bypasses the cache. Prices come from `--price-in`/`--price-out`, `price_in_per_m`/`price_out_per_m` or, for OpenRouter, its models listing; a judge run with no known price is refused (exit 2). Only process environment variables are read; `.env` files are never parsed ([`.env.example`](https://github.com/B0yko/agent-claimcheck/blob/v0.1.1/.env.example) lists the variables).

## Results

Every number below is generated from [`results/v0.1.0/`](https://github.com/B0yko/agent-claimcheck/blob/v0.1.1/results/v0.1.0/) by code, and CI checks that this block is exactly what `agent-claimcheck bench --from-recorded results/v0.1.0 --check-readme README.md` regenerates. The recorded run's command, date, hardware, concurrency and spend are in its second paragraph. The protocol, the three judges and the pass criteria of H1-H4 were committed in [ADR 0005](https://github.com/B0yko/agent-claimcheck/blob/v0.1.1/docs/adr/0005-evaluation-protocol.md) before the run. Only the test split is scored (n = 120, 48 false successes), except H4 and the all-calls parse-error columns, which count every call of the run. The positive class for AUROC is `false_success`; "missed" is a false success marked `verified`, the costliest error.

<!-- bench:start -->
300 traces (180 train / 120 test), 120 false successes (48 in test), 3 domains, 7 false-success kinds, seed 20260924, test sha256 `5300042deee1`.

Recorded run: `agent-claimcheck bench --live --judges deepseek/deepseek-v4-flash,qwen/qwen3-235b-a22b-2507,mistralai/mistral-small-3.2-24b-instruct --ablation-judge mistralai/mistral-small-3.2-24b-instruct --ablation-prompt claim-by-claim --run-name v0.1.0 --max-usd 8.0 --concurrency 8 --out results/v0.1.0` on 2026-09-28, MacBook Air M5, 24 GB, concurrency 8, 1200 judge calls, total spend $0.1789.

| judge | price in $/M | price out $/M | price date |
|---|---|---|---|
| deepseek/deepseek-v4-flash | $0.072 | $0.143 | 2026-09-28 |
| mistralai/mistral-small-3.2-24b-instruct | $0.094 | $0.250 | 2026-09-28 |
| qwen/qwen3-235b-a22b-2507 | $0.087 | $0.350 | 2026-09-28 |

### Table A: discrimination and calibration (test split)

| detector | AUROC [95% CI] | ECE raw | ECE calibrated | Brier calibrated | extremes (raw) |
|---|---|---|---|---|---|
| trust-agent | 0.500 [0.500, 0.500] | 0.400 | 0.400 | 0.400 | 100.0% |
| any-error | 0.427 [0.354, 0.500] | 0.525 | 0.525 | 0.525 | 100.0% |
| rules | 0.910 [0.848, 0.963] | 0.057 | 0.070 | 0.090 | 55.8% |
| classifier-lr | 0.947 [0.891, 0.986] | 0.056 | 0.083 | 0.082 | 56.7% |
| cascade-offline | 0.953 [0.904, 0.989] | 0.060 | 0.049 | 0.053 | 73.3% |
| judge:deepseek/deepseek-v4-flash | 0.840 [0.769, 0.905] | 0.192 | 0.035 | 0.144 | 60.8% |
| judge:mistralai/mistral-small-3.2-24b-instruct | 0.919 [0.862, 0.970] | 0.054 | 0.090 | 0.115 | 52.5% |
| judge:qwen/qwen3-235b-a22b-2507 | 0.911 [0.855, 0.964] | 0.123 | 0.033 | 0.080 | 86.7% |
| cascade | 0.928 [0.870, 0.978] | 0.065 | 0.039 | 0.065 | 82.5% |

### Table B: decisions after the shared gate (test split)

| detector | coverage | accuracy (decided) | caught | missed | false alarms | sent to review | USD/1k | wall-clock s/1k | p50 ms | p95 ms |
|---|---|---|---|---|---|---|---|---|---|---|
| trust-agent | 100.0% | 60.0% | 0/48 | 48/48 | 0/72 | 0 | $0.000 | <0.01 | <0.01 | <0.01 |
| any-error | 100.0% | 47.5% | 9/48 | 39/48 | 24/72 | 0 | $0.000 | <0.01 | <0.01 | <0.01 |
| rules | 70.0% | 95.2% | 36/48 | 4/48 | 0/72 | 36 | $0.000 | 0.12 | 0.08 | 0.29 |
| classifier-lr | 67.5% | 97.5% | 32/48 | 2/48 | 0/72 | 39 | $0.000 | 0.17 | 0.11 | 0.25 |
| cascade-offline | 90.8% | 96.3% | 41/48 | 4/48 | 0/72 | 11 | $0.000 | 0.18 | 0.11 | 0.46 |
| judge:deepseek/deepseek-v4-flash | 41.7% | 90.0% | 0/48 | 5/48 | 0/72 | 70 | $0.059 | 345 | 2157 | 3079 |
| judge:mistralai/mistral-small-3.2-24b-instruct | 52.5% | 96.8% | 26/48 | 2/48 | 0/72 | 57 | $0.146 | 343 | 2606 | 3542 |
| judge:qwen/qwen3-235b-a22b-2507 | 89.2% | 92.5% | 31/48 | 8/48 | 0/72 | 13 | $0.211 | 404 | 3101 | 5804 |
| cascade | 96.7% | 94.0% | 40/48 | 7/48 | 0/72 | 4 | $0.065 | 121 | 0.11 | 3580 |

### Recall by false-success kind (caught/total)

| kind | rules | classifier-lr | cascade-offline | judge:deepseek/deepseek-v4-flash | judge:qwen/qwen3-235b-a22b-2507 | judge:mistralai/mistral-small-3.2-24b-instruct | cascade |
|---|---|---|---|---|---|---|---|
| phantom_action | 7/7 | 5/7 | 7/7 | 0/7 | 7/7 | 5/7 | 7/7 |
| error_ignored | 7/7 | 6/7 | 7/7 | 0/7 | 5/7 | 4/7 | 7/7 |
| wrong_target | 2/7 | 3/7 | 3/7 | 0/7 | 6/7 | 5/7 | 4/7 |
| wrong_value | 4/7 | 2/7 | 5/7 | 0/7 | 2/7 | 2/7 | 5/7 |
| not_persisted | 3/7 | 5/7 | 6/7 | 0/7 | 4/7 | 4/7 | 4/7 |
| partial_completion | 7/7 | 5/7 | 7/7 | 0/7 | 1/7 | 1/7 | 7/7 |
| reviewer_injection | 6/6 | 6/6 | 6/6 | 0/6 | 6/6 | 5/6 | 6/6 |

### Evidence breakdown: state probe vs. receipt only

| detector | state_probe AUROC | state_probe missed | receipt_only AUROC | receipt_only missed |
|---|---|---|---|---|
| rules | 0.934 [0.867, 0.984] | 4 | 0.750 [0.625, 0.875] | 0 |
| classifier-lr | 0.936 [0.860, 0.989] | 2 | 0.968 [0.910, 1.000] | 0 |
| cascade-offline | 0.936 [0.875, 0.985] | 4 | 0.985 [0.950, 1.000] | 0 |
| judge:deepseek/deepseek-v4-flash | 0.838 [0.758, 0.915] | 2 | 0.830 [0.690, 0.941] | 3 |
| judge:qwen/qwen3-235b-a22b-2507 | 0.920 [0.853, 0.984] | 5 | 0.887 [0.772, 0.990] | 3 |
| judge:mistralai/mistral-small-3.2-24b-instruct | 0.952 [0.900, 0.991] | 1 | 0.841 [0.694, 0.953] | 1 |
| cascade | 0.934 [0.867, 0.984] | 4 | 0.891 [0.775, 0.995] | 3 |

![reliability](https://raw.githubusercontent.com/B0yko/agent-claimcheck/v0.1.1/results/v0.1.0/reliability.svg)
![histogram](https://raw.githubusercontent.com/B0yko/agent-claimcheck/v0.1.1/results/v0.1.0/histogram.svg)

### Leave-one-domain-out (classifier-lr)

| held-out domain | LODO AUROC | LODO ECE | shipped classifier-lr AUROC |
|---|---|---|---|
| booking | 0.948 [0.854, 1.000] | 0.128 | 0.919 [0.786, 1.000] |
| crm | 0.969 [0.909, 1.000] | 0.154 | 0.966 [0.898, 1.000] |
| coding | 0.948 [0.862, 1.000] | 0.094 | 0.969 [0.909, 1.000] |

Leakage audit (final-message-only baseline, test split): **0.434 AUROC**.

### Prompt ablation: claim-audit vs. claim-by-claim (cheapest judge)

Model: `mistralai/mistral-small-3.2-24b-instruct`.

| prompt | AUROC | ECE calibrated | coverage | caught | missed | USD/1k |
|---|---|---|---|---|---|---|
| claim-audit | 0.919 [0.862, 0.970] | 0.090 | 52.5% | 26/48 | 2/48 | $0.146 |
| claim-by-claim | 0.900 [0.836, 0.958] | 0.114 | 69.2% | 20/48 | 7/48 | $0.179 |

### Parse-error and abstention rates, and the share sent to the judge

| judge | parse-error rate (test) | abstain rate (test) | parse errors (all calls) | abstentions (all calls) |
|---|---|---|---|---|
| judge:deepseek/deepseek-v4-flash | 0.0% | 0.0% | 0/300 | 0/300 |
| judge:mistralai/mistral-small-3.2-24b-instruct | 0.0% | 0.0% | 0/300 | 0/300 |
| judge:mistralai/mistral-small-3.2-24b-instruct:claim-by-claim | 0.0% | 0.0% | 0/300 | 0/300 |
| judge:qwen/qwen3-235b-a22b-2507 | 0.0% | 0.0% | 2/300 | 2/300 |
cascade share sent to the judge: 30.0%.

### Confusion matrices

<details><summary>trust-agent</summary>

| actual \ verdict | verified | false_success | unverifiable |
|---|---|---|---|
| success | 72 | 0 | 0 |
| failure | 48 | 0 | 0 |

</details>

<details><summary>any-error</summary>

| actual \ verdict | verified | false_success | unverifiable |
|---|---|---|---|
| success | 48 | 24 | 0 |
| failure | 39 | 9 | 0 |

</details>

<details><summary>rules</summary>

| actual \ verdict | verified | false_success | unverifiable |
|---|---|---|---|
| success | 44 | 0 | 28 |
| failure | 4 | 36 | 8 |

</details>

<details><summary>classifier-lr</summary>

| actual \ verdict | verified | false_success | unverifiable |
|---|---|---|---|
| success | 47 | 0 | 25 |
| failure | 2 | 32 | 14 |

</details>

<details><summary>cascade-offline</summary>

| actual \ verdict | verified | false_success | unverifiable |
|---|---|---|---|
| success | 64 | 0 | 8 |
| failure | 4 | 41 | 3 |

</details>

<details><summary>judge:deepseek/deepseek-v4-flash</summary>

| actual \ verdict | verified | false_success | unverifiable |
|---|---|---|---|
| success | 45 | 0 | 27 |
| failure | 5 | 0 | 43 |

</details>

<details><summary>judge:mistralai/mistral-small-3.2-24b-instruct</summary>

| actual \ verdict | verified | false_success | unverifiable |
|---|---|---|---|
| success | 35 | 0 | 37 |
| failure | 2 | 26 | 20 |

</details>

<details><summary>judge:qwen/qwen3-235b-a22b-2507</summary>

| actual \ verdict | verified | false_success | unverifiable |
|---|---|---|---|
| success | 68 | 0 | 4 |
| failure | 8 | 31 | 9 |

</details>

<details><summary>cascade</summary>

| actual \ verdict | verified | false_success | unverifiable |
|---|---|---|---|
| success | 69 | 0 | 3 |
| failure | 7 | 40 | 1 |

</details>

Data sources and licences: every trace comes from the in-repo synthetic generator (Apache-2.0); the two hand-written example sets are separately authored and released under the same licence. Nothing here comes from a real system.

### Hypotheses

- H1 (rules: fewest missed, lowest coverage): **not supported**: fewest missed = classifier-lr, judge:mistralai/mistral-small-3.2-24b-instruct (2/48) vs rules 4/48; lowest coverage = judge:deepseek/deepseek-v4-flash 41.7% (rules 70.0%).
- H2 (raw judge extremes + Platt lowers ECE): **not supported**: judge:deepseek/deepseek-v4-flash: extremes raw 60.8%, ECE raw 0.192 → calibrated 0.035; judge:mistralai/mistral-small-3.2-24b-instruct: extremes raw 52.5%, ECE raw 0.054 → calibrated 0.090; judge:qwen/qwen3-235b-a22b-2507: extremes raw 86.7%, ECE raw 0.123 → calibrated 0.033.
- H3 (classifier-lr loses AUROC LODO): **not supported**: booking LODO 0.948 vs shipped 0.919; crm LODO 0.969 vs shipped 0.966; coding LODO 0.948 vs shipped 0.969.
- H4 (a judge is fooled by reviewer-directed text): **not supported** (caught/total counted over all 300 traces, with raw judge outputs through the gate; the other-kind mean averages the six other kinds' own recalls): judge:deepseek/deepseek-v4-flash: reviewer_injection recall 91.7% (11/12) vs other-kind mean 82.4%; judge:mistralai/mistral-small-3.2-24b-instruct: reviewer_injection recall 83.3% (10/12) vs other-kind mean 52.8%; judge:qwen/qwen3-235b-a22b-2507: reviewer_injection recall 83.3% (10/12) vs other-kind mean 55.6%.
<!-- bench:end -->

Reproduce it:

```bash
agent-claimcheck dataset generate --seed 20260924 --out benchmark-regen   # byte-identical to benchmark/v1
agent-claimcheck dataset validate benchmark/v1                            # schema, counts, leakage audit
agent-claimcheck bench --from-recorded results/v0.1.0 --check-readme README.md
```

Re-running the judges needs an OpenRouter key and cost about $0.18 in the recorded run; `--dry-run` prints the worst-case reservation first:

```bash
agent-claimcheck bench --live --judges deepseek/deepseek-v4-flash,qwen/qwen3-235b-a22b-2507,mistralai/mistral-small-3.2-24b-instruct --ablation-judge mistralai/mistral-small-3.2-24b-instruct --run-name my-run --max-usd 8 --concurrency 8 --dry-run
```

The dataset is described in [`benchmark/v1/DATASET_CARD.md`](https://github.com/B0yko/agent-claimcheck/blob/v0.1.1/benchmark/v1/DATASET_CARD.md): 60 genuine successes and 40 false successes per domain, genuine runs that include recovered errors and benign reviewer-directed text as hard negatives, probes present in two thirds of each cell (half for `not_persisted`, which without a probe cannot be told apart from a genuine asynchronous success), and template and entity pools split between train and test.

## Findings

Written from the recorded numbers above; every number here appears in the generated block.

- **Discrimination does not separate the detectors that matter.** Test AUROC is 0.910 for rules, 0.947 for classifier-lr, 0.953 for cascade-offline, 0.919 for the Mistral judge and 0.911 for the Qwen judge, and their 95% intervals overlap, so none of them is named a winner. The DeepSeek judge's point estimate is lower, 0.840 [0.769, 0.905], but its interval overlaps theirs too.
- **Coverage, misses and cost do separate them.** cascade-offline decides 90.8% of test traces with 4/48 false successes missed and 0/72 false alarms, for $0.000 per 1,000 traces. Adding the best judge (`cascade`, which sent 30.0% of traces to it) raises coverage to 96.7% but misses 7/48 and costs $0.065 per 1,000 traces. Single judges trade the other way round: Qwen decides 89.2% of traces but misses 8/48, the most of any non-baseline detector; Mistral misses 2/48 but decides only 52.5%. No detector raised a false alarm on a genuine success except the `any-error` baseline.
- **Calibration changes decisions, not only ECE.** Platt scaling lowered ECE for DeepSeek (0.192 to 0.035) and Qwen (0.123 to 0.033) but raised it for Mistral (0.054 to 0.090), so H2 is not supported. For DeepSeek the calibrated probabilities never reach the `false_success` threshold: it catches 0/48 and sends 70 of 120 traces to review. A better-calibrated judge can be a less decisive one, which is why coverage is reported next to ECE.
- **Results against this project's own detectors.** H1 is not supported: rules missed 4/48 while classifier-lr and the Mistral judge missed 2/48, and the DeepSeek judge had lower coverage (41.7%) than rules (70.0%). Rules are weakest where only comparing the instruction with the evidence reveals the error (wrong_target 2/7, wrong_value 4/7), and on receipt-only traces their AUROC drops to 0.750. The Qwen and Mistral judges do better on wrong targets (6/7 and 5/7) but each caught only 1/7 partial completions, which rules catch 7/7 because a tool that was never called is a structural fact.
- **Reviewer-directed text did not fool these judges more than other failures did.** H4 is not supported: over all 300 traces (12 of them with injected reviewer text), on raw outputs, each judge's recall on injected text was above its mean over the other kinds (DeepSeek 91.7% (11/12) against 82.4%, Mistral 83.3% (10/12) against 52.8%, Qwen 83.3% (10/12) against 55.6%). With 12 traces and one run each, this is weak evidence, not a robustness claim.
- **The classifier result is an optimistic read.** Rules and generator share an author, so the rules row is likely an optimistic upper bound, and classifier-lr is trained on the same generator's distribution. The leave-one-domain-out models reach 0.948 on booking, 0.969 on crm and 0.948 on coding against 0.919, 0.966 and 0.969 for the shipped model, so H3 is not supported: the generator's three domains share structure, which says little about transfer to real traces.
- **The longer prompt did not pay off.** For the Mistral judge, `claim-by-claim` scored AUROC 0.900 against 0.919 for `claim-audit` (the intervals overlap) and cost $0.179 against $0.146 per 1,000 traces. It decided more traces (69.2% against 52.5%) but missed 7/48 false successes instead of 2/48. Qwen returned 2/300 answers the strict parser rejected (a `failure_kind` outside the enum); both became abstentions, not guesses.
- **Cost.** The whole recorded run, 1,200 judge calls, cost $0.1789.

## Limitations

- **Synthetic data.** The benchmark is templated, English and generated with simulated tools. No trace comes from a real agent or a real system, and real traces are messier.
- **Same author.** The rule packs, the generator and the classifier features were written by one person against one tool API. The rules row is an optimistic upper bound for hand-written rules, and the classifier learns the generator's structure as well as the task; the leave-one-domain-out table measures transfer between the generator's own domains, not to real traces.
- **Small test split.** 120 traces, 48 false successes, about 7 per kind and 6 `reviewer_injection` traces. Most intervals overlap; read differences below a few traces as noise.
- **One judge run.** Each judge scored each trace once at temperature 0. Run-to-run variance, prompt sensitivity beyond one ablation, and provider routing effects are not measured.
- **Probes come from you.** The tool never runs agents, environments or probes; without a state probe, a receipt can at best reach review.
- **Prices change.** Costs are the recorded run's; OpenRouter prices and routing change over time.
- **Built-in calibrators** are fitted on the synthetic benchmark. On your data, fit your own with `train`.

## Roadmap

- Conformal abstention and measured run-to-run judge variance.
- Isotonic calibration and a stacking ensemble next to Platt scaling.
- A gradient-boosting classifier behind the same feature and artifact contract.
- Fine-tuned judges.
- A bench view in the dashboard.
- Importers for OpenTelemetry, Langfuse and LangSmith exports (the field mapping is documented today).
- Real, consented traces and non-English traces in a future benchmark version.

## Related work

- [τ-bench](https://arxiv.org/abs/2406.12045) (Yao et al., 2024) and [τ²-bench](https://arxiv.org/abs/2506.07982) (Barres et al., 2025) evaluate tool-using agents by the final state of a simulated environment, the same principle as this project's state probes.
- [On Calibration of Modern Neural Networks](https://arxiv.org/abs/1706.04599) (Guo et al., ICML 2017) shows modern networks are poorly calibrated and that a one-parameter scaling fixes much of it; Platt's [Probabilistic Outputs for Support Vector Machines](https://www.semanticscholar.org/paper/Probabilistic-Outputs-for-Support-vector-Machines-Platt/42e5ed832d4310ce4378c44d05570439df28a393) (1999) is the sigmoid calibrator used here, including its target smoothing.
- [From Confident Closing to Silent Failure: Characterizing False Success in LLM Agents](https://arxiv.org/abs/2606.09863) (Advani, 2026) measures false success across agent trajectories and compares LLM judges with lightweight detectors.
- [Quantifying Overclaiming Propensity in Frontier LLM Agents](https://arxiv.org/abs/2609.20812) (Smyth et al., 2026) studies final responses that report work the transcript shows was not done.
- [Verified Tool Calls Improve LLM Agent Reliability Under Non-Atomic Failures](https://arxiv.org/abs/2608.02645) (Mansoor et al., 2026) wraps tool calls with postcondition checks, the agent-side counterpart of checking receipts and state after the fact.
- Evaluation frameworks such as [Inspect](https://inspect.aisi.org.uk/), [DeepEval](https://deepeval.com/docs/getting-started) and [LangSmith](https://docs.langchain.com/langsmith/evaluation-concepts) can score agent trajectories and tool calls with custom scorers; [promptfoo](https://www.promptfoo.dev/docs/intro/) and [OpenAI Evals](https://developers.openai.com/api/docs/guides/evals) focus on model outputs. agent-claimcheck is narrower: a gate for success claims with a deterministic rule layer, calibrated abstention and a review queue, and it can sit next to any of them.

## Data sources and licences

- `benchmark/v1/`: fully synthetic, produced by the in-repo generator (`agent_claimcheck/bench/generator/`, seed 20260924). People are syllable-built fictional names, emails use `example.test`, companies are invented. Apache-2.0.
- `examples/traces.jsonl` (`example:mixed`): 12 hand-written traces with invented people and companies. `examples/browser-demo.jsonl` (`example:browser`): 24 browser-agent traces authored by the project owner and converted to agent-trace/v1, with fictional `example.com` sites and invented businesses. Neither comes from a real system. Apache-2.0.
- `results/v0.1.0/`: outputs of the recorded run, including the judges' raw answers. Apache-2.0.

See [`CONTRIBUTING.md`](https://github.com/B0yko/agent-claimcheck/blob/v0.1.1/CONTRIBUTING.md) to work on the code, [`SECURITY.md`](https://github.com/B0yko/agent-claimcheck/blob/v0.1.1/SECURITY.md) to report a vulnerability and [`CHANGELOG.md`](https://github.com/B0yko/agent-claimcheck/blob/v0.1.1/CHANGELOG.md) for releases.

## Licence

Apache-2.0, see [`LICENSE`](https://github.com/B0yko/agent-claimcheck/blob/v0.1.1/LICENSE). Copyright 2026 Andrii Boiko.
