Metadata-Version: 2.4
Name: agent-loss-map
Version: 0.18.0
Summary: Measures what an agent definition loses when it crosses a format boundary — MCP, OpenAI function-calling, A2A, UACP. Computed, not asserted.
License-Expression: MIT
Project-URL: Homepage, https://github.com/wippa-studios/agent-loss-map
Project-URL: Repository, https://github.com/wippa-studios/agent-loss-map
Project-URL: Issues, https://github.com/wippa-studios/agent-loss-map/issues
Project-URL: Protocol, https://github.com/wippa-studios/wippa-uacp
Keywords: interoperability,agent,mcp,model-context-protocol,a2a,openai,json-schema,schema-compatibility,bridge,interop
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Software Development :: Testing
Classifier: Topic :: Software Development :: Quality Assurance
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: jsonschema>=4.0
Provides-Extra: dev
Requires-Dist: pytest>=7.0; extra == "dev"
Requires-Dist: build>=1.0; extra == "dev"
Requires-Dist: twine>=5.0; extra == "dev"
Dynamic: license-file

# agent-loss-map

[![PyPI](https://img.shields.io/pypi/v/agent-loss-map)](https://pypi.org/project/agent-loss-map/)
[![Python](https://img.shields.io/pypi/pyversions/agent-loss-map)](https://pypi.org/project/agent-loss-map/)
[![CI](https://github.com/wippa-studios/agent-loss-map/actions/workflows/ci.yml/badge.svg)](https://github.com/wippa-studios/agent-loss-map/actions)

## What does your agent definition lose when it crosses a format boundary?

![agent-loss-map report](docs/scorecard.svg)

Every agent framework invented its own idea of what a tool is. MCP has
`tools/list`. OpenAI function-calling has a `parameters` block with a narrow
schema subset. UACP has a capability schema that uses JSON Schema properly. The
differences between them are invisible until an agent fails in production, and
nobody has tooling for measuring them.

**This measures them.** It takes a format-native agent description, runs the
resulting artifact through a real JSON Schema validator, and reports every piece
of information that does not survive — with evidence, and with a confidence tier
saying how much the finding is worth.

```bash
uvx agent-loss-map
```

No install, no clone, no API key. A run takes **under 200 ms** (median 139 ms over 20 runs).

## Who this is for

**Bridge and gateway authors.** If you are shipping MCP↔A2A, MCP↔OpenAI, or any
pair of agent formats, your bridge is losing information at the boundary and
your current answer is a hand-written `lossy` dict. This computes it instead —
and it is fast enough to run in CI before merge.

**Protocol and spec authors.** If your spec has normative text and JSON
Schemas, this machine-checks whether they agree with each other — a class of
defect that a conformance checklist cannot see, because a checklist proves your
code matches *your own reading* of the spec, not that your reading and your code
agree.

**Framework maintainers.** If your framework publishes agent definitions, this
tells you what a consumer in another format will not be able to represent.

## What it measures

Four checks, all computed rather than asserted:

1. **Schema versus prose.** Build a message the spec *requires* an
   implementation to accept, validate it against the shipped schema, report the
   rejection.
2. **Schema versus the implementations.** Check each reference implementation
   can represent what the schema declares, and that they model the same envelope.
3. **Cross-format divergence, both directions.** Map a framework-native agent
   description into a target descriptor and validate the artifact actually
   produced; separately, test whether a capability survives being expressed in a
   format's native tool schema.
4. **Target-format rules JSON Schema cannot express.** An
   OpenAI-compatible `parameters` block must carry `properties`, so a capability
   publishing `{"type": "object"}` gets the whole tool list rejected with a 400.
   The capability schema accepts that shape perfectly well, so the constraint
   only exists in the target format — which is the same class as a schema
   keyword with no representation there.

Losses are **computed, not asserted**. `SCHEMA_VIOLATION` findings come from
running the real validator over the artifact an adapter produced — a hand-copied
regex predicting the same thing was removed in `0.1.0` because it double-counted
every finding.

## Every finding carries evidence and a confidence tier

This is the part most tooling in this space skips, and it is the reason to trust
the output:

| Tier | Meaning |
|---|---|
| `observed-serialization` | we read a real serialized artifact |
| `documented-api` | from official documentation of the public API |
| `inferred` | our modelling choice, not a documented shape |

The report prints the tier for every entry and says which findings rest on
weaker evidence. An entry that cannot be sourced honestly is worse than a missing
one, and the loader refuses to guess: a missing or misspelt `confidence` is an
error, not a silent downgrade to the weakest tier.

## Proof it works: eleven defects in a protocol we wrote

We built this to grade [wippa-uacp](https://github.com/wippa-studios/wippa-uacp),
the protocol we wrote. The first run reported five blockers. All five were real,
and all five are fixed:

| Was | Defect | Status |
|---|---|---|
| 🔴 2 blockers | `metadata.auth` existed in the TypeScript model and not the Python one, so an authenticated message decoded as **anonymous, with no error** | fixed |
| 🟠 3 majors | `capability` was in the schema's global `required` list, so `heartbeat` and `register` had to invent a placeholder — and both implementations sent the literal `"_internal"` | fixed |
| 🟠 1 major | `flowControl` was specified in the prose, modelled by both implementations, and declared by **no schema** | fixed |
| 🟠 1 major | `async` defaulted to `true`, so a plain request/response capability was documented as streaming — the consumer waits for a `stream.end` that never comes, and sees a **hang**, not an error | fixed |
| 🔴 1 blocker | the audit trail persisted bearer tokens in plaintext | fixed |

The most useful one nothing else could have caught: **a security control that
was never wired.** SPEC.md said a bus "can enforce capability allow-lists per
caller". Both implementations shipped a correct, unit-tested
`check_authorization`. `Bus.send` never called it. A schema-versus-prose
comparison cannot see this — both were fine, because the prose under-committed.
"Can" reads as a capability, not an obligation. Only reading the code can.

The check that now guards it is a source probe, not a behavioural test, and the
README says so rather than implying otherwise. An earlier version recorded
`enforced: true` by hand, which meant it was asserting the author's own reading
back at him; had the call been deleted later it would have reported clean
forever. A source probe is strictly better than that and still not the same
thing as running the code.

It also found a compatibility defect by measuring live servers rather than
reading the schema: UACP rejected hyphens in capability and agent names, so
`resolve-library-id` — a name OpenAI's own grammar accepts — could not be named
without renaming, and a renamed tool is a different tool to any client that
refers to it by name. Two of six live MCP servers could not describe themselves
at all. No amount of self-consistent testing would have found this; only
measuring servers we did not build did.

## Add your format

A corpus entry is pure data and the checks read nothing else, so measuring a
format needs no Python and no pull request:

```bash
agent-loss-map --corpus my-format.json
```

```json
{
  "framework": "my-framework",
  "agent_id": "my_agent",
  "description": "What this is.",
  "confidence": "documented",
  "provenance": "Where you read the shape, and when",
  "capabilities": [
    { "name": "web.search", "description": "Search the web.",
      "parameters": { "type": "object",
                      "properties": { "query": { "type": "string" } } } }
  ]
}
```

It gets its own section in the report, computed by the same checks as everything
else. `--corpus` also takes a directory and is repeatable.

A worked example, captured from a live server rather than synthesised:

```bash
uvx agent-loss-map --corpus examples/deepwiki-mcp.entry.json
```

`examples/deepwiki-mcp.capture.json` and `examples/context7.capture.json` are raw
`tools/list` results from two live servers — DeepWiki 2.14.3 and Context7 4.1.1 —
each taken over a real MCP handshake, with the matching `.entry.json` a
projection of it.

`scripts/capture_mcp.py` does both halves: it performs the handshake and then
projects the result, and it runs the fidelity check before writing an entry, so a
projection that disagrees with its own artifact is refused rather than published.
That check exists because an `outputSchema` was dropped in silence during
exactly this conversion, by hand, and the harness then reported the omission as a
major finding on all three tools — as if the server had left it out. It had not.

```bash
git clone https://github.com/wippa-studios/agent-loss-map && cd agent-loss-map
python scripts/capture_mcp.py https://mcp.deepwiki.com/mcp --out deepwiki-mcp
```

The scripts are repository tooling rather than installed console entry points, so
run them from a clone.

To make a format permanent, append a `ForeignAgent` to
`agent_loss_map/corpus.py` — about twenty minutes, and
[CONTRIBUTING.md](CONTRIBUTING.md) has the recipe and the provenance rules.

## Install

```bash
uvx agent-loss-map                       # or: pipx run agent-loss-map
```

```bash
pip install -e .                       # from a checkout
agent-loss-map
```

| Flag | |
|---|---|
| `--scorecard` | the per-section table on its own |
| `--markdown` | the full report as Markdown, for an issue or a PR |
| `--json` | machine-readable |
| `--badge` | status-badge JSON for the worst finding against the spec |
| `--corpus PATH` | measure a format that is not bundled |
| `--fail-on-blocker` | exit 1 if any blocker is found |

Every renderer derives its counts from the same findings, and a test asserts they
cannot disagree — a renderer that tallied severities independently is the
`0.1.0` double-counting bug in a new disguise. `--badge` reports only the spec
and implementation audit, because a blocker in a *foreign* format is a finding
about that format and reddening the badge would misattribute the fault.

## It cannot go stale quietly

The report is only worth anything if it is current, so:

- the vendored schemas are pinned to an upstream commit in
  `agent_loss_map/schemas/PROVENANCE.json`, and CI fails on drift
- a scheduled workflow re-runs the probe weekly and opens an issue if the result
  stops matching the committed artifact
- `docs/scorecard.svg` in this README is generated from a live run, and CI fails
  if it goes stale

A conformance report that silently tests an old protocol is worse than no
report.

## Known limits

- **One external `$ref` needs a local registry to resolve.** The schemas claim
  canonical URLs on a host that does not resolve, so a plain offline validator
  fetching `agent-descriptor` will attempt a network call. This harness builds a
  `referencing.Registry` for exactly that reason, and the finding is reported
  rather than hidden.
- **No behavioural half.** Everything here is static: *these two definitions
  cannot both be true*, and, for the security claims, *this code path does or
  does not call this function*. Nothing executes an implementation. The claim
  that found the unwired authorization was a person reading the code; the check
  that guards it afterwards is a source probe, not a behavioural test, and the
  difference matters if you are deciding how much to trust a clean report.
- **An absent protocol checkout makes the code checks unverified, not passing.**
  Set `UACP_SOURCE_ROOT` to a checkout to have them run. Silently skipping them
  and reporting clean would be the failure this harness exists to catch.
- **The corpus is small**, because the honest entry is one you can source. Three
  ship; a LangGraph entry is [open](https://github.com/wippa-studios/agent-loss-map/issues/3).
- **Not yet format-to-format.** It measures loss crossing into a target
  descriptor and loss expressing one in a native tool schema, not directly
  between two foreign formats. That is the obvious next thing and it is not
  built.

## Verifying it

```bash
pip install -e ".[dev]"
pytest -q
```

96 tests. The suite is mostly meta-tests: each mutates a copy of the input so the
defect it looks for is absent, and asserts the check goes quiet. A check that
quietly stopped working fails the suite rather than passing silently — which has
twice caught a guard in this repo that had quietly stopped guarding anything,
both times because relaxing a rule invalidated the fixture it used to trigger on.

MIT.
