Metadata-Version: 2.4
Name: agent-loss-map
Version: 0.21.0
Summary: Measures what an agent tool definition loses when converted between MCP, OpenAI, A2A, Anthropic, Gemini, Bedrock and 10 more formats. Computed, not asserted.
License-Expression: MIT
Project-URL: Homepage, https://github.com/wippa-studios/agent-loss-map
Project-URL: Repository, https://github.com/wippa-studios/agent-loss-map
Project-URL: Issues, https://github.com/wippa-studios/agent-loss-map/issues
Project-URL: Protocol, https://github.com/wippa-studios/wippa-uacp
Keywords: interoperability,agent,mcp,model-context-protocol,a2a,openai,json-schema,schema-compatibility,bridge,interop,function-calling,tool-schema,agents,anthropic,gemini,bedrock,tool-use
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Software Development :: Libraries :: Python Modules
Classifier: Topic :: Software Development :: Quality Assurance
Classifier: Topic :: Software Development :: Testing
Classifier: Topic :: Software Development :: Quality Assurance
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: jsonschema>=4.0
Provides-Extra: dev
Requires-Dist: pytest>=7.0; extra == "dev"
Requires-Dist: build>=1.0; extra == "dev"
Requires-Dist: twine>=5.0; extra == "dev"
Provides-Extra: behavioural
Requires-Dist: openai==1.109.1; extra == "behavioural"
Requires-Dist: anthropic==1.11.0; extra == "behavioural"
Requires-Dist: mcp==1.20.0; extra == "behavioural"
Dynamic: license-file

# agent-loss-map

<!-- The ?style=flat suffix is a cache-buster. shields.io caches a negative
     result too, so a badge can sit reading "package or version not found"
     for weeks after the first publish until the cache happens to expire. The
     suffix forces a fresh fetch without changing what the badge reports. -->
[![PyPI](https://img.shields.io/pypi/v/agent-loss-map?style=flat)](https://pypi.org/project/agent-loss-map/)
[![Python](https://img.shields.io/pypi/pyversions/agent-loss-map?style=flat)](https://pypi.org/project/agent-loss-map/)
[![CI](https://github.com/wippa-studios/agent-loss-map/actions/workflows/ci.yml/badge.svg)](https://github.com/wippa-studios/agent-loss-map/actions)

## What does your agent definition lose when it crosses a format boundary?

![agent-loss-map report](docs/scorecard.svg)

Every agent framework invented its own idea of what a tool is, and the differences are invisible until an agent fails in production. **This measures them.** It projects a format-native tool definition into another format, runs the artifact it actually produced, and reports every piece of information that did not survive — with evidence, and a confidence tier saying how much the finding is worth.

```bash
uvx agent-loss-map --matrix
```

No install, no clone, no API key. The bundled run makes no network call at all — `--matrix`, `--chain` and `--behavioural` are computed entirely from the shipped corpus and the vendored schemas. 16 formats in 7 classes, 208 crossings in the matrix, 325 tests. The whole matrix computes in about 400 ms — though that is lopsided: the Stripe entry alone is 361 ms of it, because its 20-property schemas are walked against fifteen dialects. The other twelve sources together are 48 ms. (`--mcp` is the one flag that reaches the network, because its job is to measure a server you have not read yet.)

---

## Some formats cannot hold a tool contract at all

**A2A's protocol data model has no JSON Schema in it.** `message AgentSkill` in the normative source — [`specification/a2a.proto`](https://github.com/a2aproject/A2A/blob/v1.0.1/specification/a2a.proto) at tag `v1.0.1` — has eight fields: `id`, `name`, `description`, `tags`, `examples`, `input_modes`, `output_modes`, `security_requirements`. The words `schema` and `json_schema` do not appear anywhere in that file's 811 lines. A skill is described by prose and by media-type strings.

So a real MCP server crossing into A2A loses its entire structured contract, and the reverse direction cannot recover what was never there:

```console
$ agent-loss-map --from deepwiki --to a2a

  source    findings  blocker  major  minor
  ---------  --------  -------  -----  -----
  deepwiki         6        6      0      0
```

Six blockers, zero majors, zero minors — because there is no degraded path to report. Every schema is destroyed rather than downgraded, and a bridge that papers over this by inventing an `inputSchema` has written a private convention and called it a protocol.

This is a property of the format, not a criticism of it: A2A is opaque-by-design about internals, which is a legitimate goal for agent-to-agent exchange. It does mean any `inputSchema` a bridge attaches is a private convention, not part of the protocol. Section 1.4 of [`specification.md`](https://github.com/a2aproject/A2A/blob/v1.0.1/docs/specification.md) makes the proto the single authoritative definition and describes the generated JSON artifact as a non-normative build artifact.

---

## The formats are not one kind of thing

| | formats | what a schema crossing does |
|---|---|---|
| **can carry a schema** | 7 — `mcp`, `openai-fc`, `anthropic`, `gemini`, `bedrock-converse`, `bedrock-agent`, `langchain` | the schema crosses, minus whatever keywords the target's dialect refuses |
| **contract in pieces, no join rule** | 2 — `openapi`, `stripe` | the parameters exist; nothing specifies how to flatten them into one object |
| **a different type system** | 2 — `shopify`, `langgraph` | GraphQL arguments and state channels are not a JSON Schema dialect |
| **media types only** | 1 — `a2a` | **nothing survives** |
| **taxonomy references** | 1 — `oasf` | skills are `{name, id}` vocabulary refs, not contracts |
| **prose** | 1 — `agent-skills` | five frontmatter fields and Markdown |
| **no machine-readable artifact** | 2 — `anp`, `emvco` | nothing to project into |

Seven classes, and the split is the finding rather than a category list: two formats that look interchangeable in a README destroy and downlevel the same schema by completely different mechanisms. Where a format's inputs are several located declarations with no join rule — OpenAPI's `parameters[]`, and Stripe's — no specification says how to flatten them into one object, and three real implementations produce three incompatible answers, so the harness reports the ambiguity rather than inventing a mapping and calling the crossing lossless. [DESIGN.md](docs/DESIGN.md#the-format-registry) has the registry, the guard that refuses a self-contradictory record, and [OpenAPI/Stripe](docs/DESIGN.md#openapi-and-stripe) in full.

---

## Multi-hop loss is not additive, and no single crossing shows it

A single crossing answers "what does this source lose entering that format". A bridge asks a third question: *what survives the whole path*. `--chain` walks it, feeding each hop the artifact the previous hop actually produced:

```bash
agent-loss-map --from deepwiki --chain mcp,a2a,anthropic,openai-fc
```

| hop | format | findings | cumulative intact | lost |
|---:|---|---:|---:|---:|
| 1 | `mcp` | 0 | 100% | 0 |
| 2 | `a2a` | 6 blockers | 46% | 7 |
| 3 | `anthropic` | **0** | 46% | 7 |
| 4 | `openai-fc` | 3 blockers | 46% | 7 |

Cumulative intact does not rise at hop 3 or hop 4: the fields were already gone, and no later hop invented any back. Two things in that table are worth more than the whole of `--matrix`.

**Hop 3 reporting zero findings is not a clean hop.** Anthropic has a schema field, and it reported nothing at all — because A2A's `AgentSkill` left it nothing to report about. Summing per-hop counts renders that as a clean second hop. It was a hop that ran on an empty document.

**No pair of single-hop crossings reveals hop 4's blockers.** Three blockers appear only at the end of the chain; the same two hops measured independently produce nothing. And they land on a `{"type": "object"}` **the chain itself fabricated at hop 3**, because A2A had already destroyed the real contracts — so Anthropic was handed nothing and emitted an empty schema, and the target's own rule that a `parameters` block must carry `properties` then refused it. The finding is real and the attribution belongs to the chain, not to Anthropic or OpenAI. That is the compound failure this whole feature exists to name, and a per-crossing matrix scores it clean.

The intermediate is **this harness's own re-read, not a client's parse**: hop N is fed the artifact hop N−1 actually produced, converted back into a capability list by `agent_loss_map.chain.reread`. `intact` means the artifact carries that value, not that any client took it. The chain report prints that caveat with the result and `--json` carries it in the payload. [DESIGN.md](docs/DESIGN.md#multi-hop) covers field fate, terminal states, and the re-read's known distortions.

---

## The executed half: third-party code, actually run

Everything above reasons. `--behavioural` does not — it imports client libraries, calls their own code with the artifacts this harness produced, and writes down what came back. Its vocabulary has five words and none of them is a pass by another name: `accepted`, `REFUSED`, `altered`, `DIVERGED`, `NOT RUN`.

**Only three formats have a client library to run, and you have to name one of them.** `openai-fc`, `anthropic` and `mcp` are bound to the vendors' own SDKs. The other 13 report `NOT RUN` and say why — and the reason is not always "no library": some have no client binding, and some have no input contract to execute at all.

The default run measures against UACP, which has no client library, so:

```console
$ agent-loss-map --behavioural
  accepted=0  refused=0  altered=0  diverged=0  not run=5
```

Nothing executed, at exit 0, because that is an honest result and not a crash. To see real client code run:

```bash
agent-loss-map --from mcp --to mcp --behavioural        # all four probes
agent-loss-map --matrix --behavioural                    # every crossing
```

Three decisions keep it from inflating the result:

- It is a separate type with no severity and no code, so it cannot enter the finding totals, the scorecard, the badge or the gate. An executed result is not a `Divergence`; concatenating the two raises, and a test asserts it. The default run, `--matrix`, `--chain` and `--gate` are byte-identical without `--behavioural`.
- `NOT RUN` is never a pass. A machine that cannot run a check and a machine whose check passed would otherwise produce identical bytes. `clean`, `pass`, `passed`, `ok` and `success` are not in the vocabulary and a test asserts they cannot get in.
- A local type check is not a service acceptance. Every executed result says which of the two it is, and the test suite asserts the wording.

`--behavioural` is refused with `--chain` (exit 2) rather than ignored. What each of the four probes runs, which callable, and which pinned version: [DESIGN.md](docs/DESIGN.md#the-executed-half).

---

## Who this is for

**Bridge and gateway authors.** If you are converting between any of these formats, your bridge is losing information at the boundary and the usual answer is a hand-written `lossy` dict. This computes it — and for a multi-hop path, it computes the compounding the dict cannot show.

**Anyone choosing a format.** The matrix is the comparison: which targets a source crosses into intact, and which destroy it. A format that can hold a schema and one that cannot look identical in a README and behave nothing alike.

**Protocol and spec authors.** If your spec has normative text and JSON Schemas, this machine-checks whether they agree with each other — a class of defect a feature checklist cannot see, because a checklist proves your code matches *your own reading* of the spec, not that your reading and your code agree. It found eleven such defects in our own protocol: eight in the schema-versus-prose ledger, one security control that was correct, unit-tested, and never called, and two naming-rule incompatibilities found by measuring servers we did not build. Every one is fixed upstream and retained as a regression guard. [DESIGN.md](docs/DESIGN.md#the-eleven-uacp-defects) has the ledger, its arithmetic, and the defect-by-defect account.

**Framework maintainers.** If your framework publishes agent definitions, this tells you what a consumer in another format will not be able to represent.

---

## Install

```bash
uvx agent-loss-map                       # or: pipx run agent-loss-map
```

```bash
pip install -e .                       # from a checkout
agent-loss-map
```

`--behavioural` needs three client libraries, which are an optional extra rather
than a default dependency:

```bash
pip install 'agent-loss-map[behavioural]'
```

The client libraries are pinned exactly (`openai==1.109.1`, `anthropic==1.11.0`, `mcp==1.20.0`) rather than with a floor, because every executed result names the installed version in its evidence — "the client's types accept this" is worth nothing without knowing whose types they are.

`--mcp` measures a live server: the harness performs a real handshake, projects `tools/list` faithfully, refuses to report anything if its own projection dropped a field, and reports the era it spoke. Three MCP captures ship with the package, so `agent-loss-map --from mcp --to openai-fc` works with nothing else installed.

| Flag | |
|---|---|
| `--matrix` | every source against every target in one table |
| `--chain F,F,F` | measure a multi-hop path, feeding each hop the previous hop's artifact |
| `--behavioural` | also run the executed half against real client-library code |
| `--scorecard` | the per-section table on its own |
| `--markdown` | the full report as Markdown, for an issue or a PR |
| `--json` | machine-readable |
| `--badge` | status-badge JSON for the worst finding against the spec |
| `--mcp URL` | capture `tools/list` from a live server and measure it now |
| `--from SOURCE` | only measure sources matching a framework id or format family |
| `--to TARGET` | project into a target format and validate the artifact produced |
| `--corpus PATH` | measure a format that is not bundled |
| `--fail-on-blocker` | exit 1 if any blocker is found |
| `--gate` | exit 1 only on a blocker against UACP itself (the release gate) |

Modes answer different questions and refuse to be combined: `--chain` rejects `--to`, `--matrix` and `--behavioural` with exit 2 and says why. `--gate` is not evaluated on a chain run, because a chain cannot answer whether UACP's own schemas agree with its own normative text, and answering it "clean" would be a false green.

A report that silently tests an old protocol is worse than no report, so two workflows guard that: CI fails a build on real schema drift and on the test suite, while a scheduled probe re-runs the measurement and opens a PR when the committed report, badge or scorecard differ. Staleness becomes visible in review rather than breaking a build — [DESIGN.md](docs/DESIGN.md#it-cannot-go-stale-quietly) has both workflows and the pinned vendored-schema provenance.

---

## Add your format

A corpus entry is pure data and the checks read nothing else, so measuring a format needs no Python and no pull request:

```bash
agent-loss-map --corpus my-format.json
```

```json
{
  "framework": "my-framework",
  "agent_id": "my_agent",
  "description": "What this is.",
  "confidence": "documented",
  "provenance": "Where you read the shape, and when",
  "capabilities": [
    { "name": "web.search", "description": "Search the web.",
      "parameters": { "type": "object",
                      "properties": { "query": { "type": "string" } } } }
  ]
}
```

It gets its own section in the report, computed by the same checks as everything else. `--corpus` also takes a directory and is repeatable. Every entry declares how well-sourced it is, and the loader refuses to guess — a missing or misspelt `confidence` is an error, not a silent downgrade:

| Tier | Meaning |
|---|---|
| `observed-serialization` | we read a real serialized artifact |
| `documented-api` | from official documentation of the public API |
| `inferred` | our modelling choice, not a documented shape |

The corpus is 13 entries: **5 `observed-serialization`** (three MCP captures, a UACP descriptor quoted from its spec, and the Stripe subset derived by script) and **8 `documented-api`**. The default run reports 44 findings — 4 spec-audit (2 major, 2 minor) and 40 cross-format (1 blocker, 13 major, 26 minor). Today's run against `wippa-uacp` reports zero blockers in the spec audit; the run's one blocker is loss measured in a *foreign* format, which is what the harness is for. Captured in [`docs/sample-report.txt`](docs/sample-report.txt).

To make a format permanent, append a `ForeignAgent` to `agent_loss_map/corpus.py` — about twenty minutes. [CONTRIBUTING.md](CONTRIBUTING.md) has the recipe and the provenance rules; [DESIGN.md](docs/DESIGN.md#how-to-add-a-format) has the registry guards a new format must satisfy.

---

## Known limits

- **The security claims are still guarded by a source probe, not by execution, and that is the limit that matters most.** The unwired authorization was found by a person reading the code and is now guarded by a regex over the routing path. `--behavioural` does not change that: no probe runs UACP's own `Bus`, because the package must work without a checkout of the protocol repo. A source probe can be satisfied by code that is never called, and this one is still open to that.
- **A chain is fed forward through this harness's own re-read, not a client's parse.** `--chain mcp,openai-fc,a2a` measures the path, but hop 2 consumes the artifact hop 1 *actually produced*, converted back into a capability list by `agent_loss_map/chain.py`. No OpenAI server, A2A client or Anthropic API was involved, and a real client may refuse a shape the re-read accepts. `intact` means the artifact carries that value, not that a client took it. The re-read's known distortions are enumerated per hop and printed with the chain, and the intermediate is labelled `inferred` because a real system did not serialize it — this package did. `--matrix` still computes each crossing independently; a chain is the third question, not a substitute for the other two.
- **Per-hop finding counts in a chain are not additive, and the report will not pretend they are.** The headline is the cumulative share of the source's own declared fields the artifact still carries, and a hop that reported nothing because it had nothing left to lose is called out by name.
- **The executed half is real, and much smaller than the static one.** `--behavioural` covers 3 formats with a maintained Python client; the other 13 report `NOT RUN` and say why. `unavailable` is never rendered as a pass, and `clean`/`pass` are not statuses the vocabulary can produce. A local type check is not a service acceptance: when the `openai` model's pydantic validation accepts a function block, no request was built, no key was needed, and nothing is known about what the API would do.
- **A format with no published keyword subset cannot produce a keyword finding.** `mcp`, `langchain`, `openapi`, `shopify` and `stripe` declare no accepted subset, so a clean cell for them means "the rules they publish were satisfied", not "fully supported". A constraint on what a format accepts is not a thing it has documented.
- **Two records rest on mirror-sourced or disputed facts, and say so.** The Anthropic keyword subset was read from a mirror of the structured-outputs limitations page, so it accepts rather than reports the keywords in doubt; the Anthropic and Gemini name-length limits are disputed between two official surfaces, so no length finding is reported for either.
- **The bundled MCP captures are not all the same protocol era, and each one says which it is.** `context7` is `2026-07-28` — modern, captured over `server/discover`, with per-request `_meta` and no session. `deepwiki-mcp` and `adoraads-beauty` are `2025-06-18` — legacy, captured over an `initialize` handshake. Every capture was probed for `server/discover` and labelled by what the server answered, not inferred from the corpus. The `Tool` data type — what the harness actually measures — has not changed in the ways that matter, but the two eras negotiate the transport differently, and a legacy client against a modern server fails on the current revision's own compatibility matrix.
- **A standards-compliant validator still cannot resolve the UACP schemas offline.** The `$id`s claim canonical URLs on a host that does not resolve, so loading `agent-descriptor` and following its relative `capability.schema.json` ref attempts a network call. This harness resolves those refs locally by filename as well as by `$id`, and a test proves the resolution is offline rather than incidentally working. The finding stays major because the defect is in the published artifact, not in the workaround.
- **An absent protocol checkout makes the code checks unverified, not passing.** Set `UACP_SOURCE_ROOT` to a checkout to have them run. Silently skipping them and reporting clean would be the failure this harness exists to catch.
- **Release notes are per version, in [`docs/`](docs/).** [0.21.0](docs/RELEASE-0.21.0.md) is the executed half's self-comparison, six keyword absences, and unreachable version negotiation. [0.20.0](docs/RELEASE-0.20.0.md) closes an SSRF in `--behavioural` — if you are on 0.19.0 or earlier, upgrade; `agent-loss-map --version`. [0.19.0](docs/RELEASE-0.19.0.md) adds `--chain` and `--behavioural`.
- **`FINDINGS.md` covers the spec audit, not the whole run.** It explains what the schema-versus-prose and schema-versus-implementation checks found in `wippa-uacp`, which is the durable part. For current counts across all 13 corpus sources, read [`docs/sample-report.txt`](docs/sample-report.txt), which the scheduled probe regenerates from a live run.

---

## Verifying it

```bash
pip install -e ".[dev]"
pytest -q
```

325 tests, 1 skipped. The suite is mostly meta-tests: each mutates a copy of the input so the defect it looks for is absent, and asserts the check goes quiet. A check that quietly stopped working fails the suite rather than passing silently — which has twice caught a guard in this repo that had quietly stopped guarding anything, both times because relaxing a rule invalidated the fixture it used to trigger on.

MIT.
