Metadata-Version: 2.4
Name: uacp-interop
Version: 0.6.0
Summary: Interop conformance probe for UACP: computes where UACP's schemas, prose and reference implementations diverge, and where a UACP agent definition cannot survive a round trip through a framework-native tool format.
License-Expression: MIT
Project-URL: Homepage, https://github.com/wippa-studios/uacp-interop
Project-URL: Repository, https://github.com/wippa-studios/uacp-interop
Project-URL: Issues, https://github.com/wippa-studios/uacp-interop/issues
Project-URL: Protocol, https://github.com/wippa-studios/wippa-uacp
Keywords: uacp,interoperability,conformance,agent,json-schema,protocol
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Software Development :: Testing
Classifier: Topic :: Software Development :: Quality Assurance
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: jsonschema>=4.0
Provides-Extra: dev
Requires-Dist: pytest>=7.0; extra == "dev"
Requires-Dist: build>=1.0; extra == "dev"
Requires-Dist: twine>=5.0; extra == "dev"
Dynamic: license-file

# uacp-interop

[![uacp spec](https://img.shields.io/badge/dynamic/json?url=https%3A%2F%2Fraw.githubusercontent.com%2Fwippa-studios%2Fuacp-interop%2Fmain%2Fdocs%2Fbadge.json&label=uacp%20spec&prefix=)](https://github.com/wippa-studios/uacp-interop)
[![PyPI](https://img.shields.io/pypi/v/uacp-interop)](https://pypi.org/project/uacp-interop/)
[![Python](https://img.shields.io/pypi/pyversions/uacp-interop)](https://pypi.org/project/uacp-interop/)

## What does your agent definition lose when it crosses a format boundary?

![uacp-interop report](docs/scorecard.svg)

Every agent framework invented its own idea of what a tool is. MCP has
`tools/list`. OpenAI function-calling has a `parameters` block with a narrow
schema subset. UACP has a capability schema that uses JSON Schema properly. The
differences between them are invisible until an agent fails in production, and
nobody has tooling for measuring them.

This measures them. It takes a format-native agent description, runs the
resulting artifact through a real JSON Schema validator, and reports every piece
of information that does not survive — with evidence, and with a confidence tier
saying how much the finding is worth.

```bash
uvx uacp-interop
```

No install, no clone, no API key. A run takes **under 200 ms** (median 139 ms over 20 runs).

## It found five real defects in its own protocol

We built this to grade [wippa-uacp](https://github.com/wippa-studios/wippa-uacp),
the protocol we wrote. The first run reported five blockers. All five were real,
and all five are fixed:

| Was | Defect | Status |
|---|---|---|
| 🔴 2 blockers | `metadata.auth` existed in the TypeScript model and not the Python one, so an authenticated message decoded as **anonymous, with no error** | fixed |
| 🟠 3 majors | `capability` was in the schema's global `required` list, so `heartbeat` and `register` had to invent a placeholder — and both implementations sent the literal `"_internal"` | fixed |
| 🟠 1 major | `flowControl` was specified in the prose, modelled by both implementations, and declared by **no schema** | fixed |
| 🟠 1 major | `async` defaulted to `true`, so a plain request/response capability was documented as streaming — the consumer waits for a `stream.end` that never comes, and sees a **hang**, not an error | fixed |
| 🔴 1 blocker | the audit trail persisted bearer tokens in plaintext | fixed |

That is the part worth caring about. A conformance checklist proves your
implementation matches *your own reading* of a spec. It says nothing about
whether your reading and your code agree with each other, or with anyone
else's. That is what this measures instead, and every finding it produced here
was one a checklist would have passed.

## What it actually checks

1. **Schema versus prose.** Build a message the spec *requires* an
   implementation to accept, validate it against the shipped schema, report the
   rejection.
2. **Schema versus the implementations.** Check each reference implementation can
   represent what the schema declares, and that they model the same envelope.
3. **Cross-format divergence, both directions.** Map a framework-native agent
   description into a target descriptor and validate the artifact actually
   produced; separately, test whether a capability survives being expressed in a
   format's native tool schema.
4. **Target-format rules JSON Schema cannot express.** An
   OpenAI-compatible `parameters` block must carry `properties`, so a capability
   publishing `{"type": "object"}` gets the whole tool list rejected with a 400.
   The capability schema accepts that shape perfectly well, so the constraint
   only exists in the target format — which is the same class as a schema
   keyword with no representation there.

Losses are **computed, not asserted**. `SCHEMA_VIOLATION` findings come from
running the real validator over the artifact an adapter produced — a hand-copied
regex predicting the same thing was removed in `0.1.0` because it double-counted
every finding. Divergences in the other direction are found by walking the source
JSON Schema and collecting keywords outside the target format's supported
subset.

## Every finding carries evidence and a confidence tier

This is the part most tooling in this space skips, and it is the reason to trust
the output:

| Tier | Meaning |
|---|---|
| `observed-serialization` | we read a real serialized artifact |
| `documented-api` | from official documentation of the public API |
| `inferred` | our modelling choice, not a documented shape |

The report prints the tier for every entry and says which findings rest on
weaker evidence. The bundled AutoGen entry models a documented constructor
surface rather than an observed serialization, so it is labelled as the weakest
in the corpus. An entry that cannot be sourced honestly is worse than a missing
one, and the loader refuses to guess: a missing or misspelt `confidence` is an
error, not a silent downgrade to the weakest tier.

## Add your format

A corpus entry is pure data and the checks read nothing else, so measuring a
format needs no Python and no pull request:

```bash
uacp-interop --corpus my-format.json
```

```json
{
  "framework": "my-framework",
  "agent_id": "my_agent",
  "description": "What this is.",
  "confidence": "documented",
  "provenance": "Where you read the shape, and when",
  "capabilities": [
    { "name": "web.search", "description": "Search the web.",
      "parameters": { "type": "object",
                      "properties": { "query": { "type": "string" } } } }
  ]
}
```

It gets its own section in the report, computed by the same checks as everything
else. `--corpus` also takes a directory and is repeatable. To make it permanent,
append a `ForeignAgent` to `uacp_interop/corpus.py` — about twenty minutes, and
[CONTRIBUTING.md](CONTRIBUTING.md) has the recipe and the provenance rules.

## Install

```bash
uvx uacp-interop                       # or: pipx run uacp-interop
```

```bash
pip install -e .                       # from a checkout
uacp-interop
```

| Flag | |
|---|---|
| `--scorecard` | the per-section table on its own |
| `--markdown` | the full report as Markdown, for an issue or a PR |
| `--json` | machine-readable |
| `--badge` | status-badge JSON for the worst finding against the spec |
| `--corpus PATH` | measure a format that is not bundled |
| `--fail-on-blocker` | exit 1 if any blocker is found |

Every renderer derives its counts from the same findings, and a test asserts they
cannot disagree — a renderer that tallied severities independently is the
`0.1.0` double-counting bug in a new disguise. `--badge` reports only the spec
and implementation audit, because a blocker in a *foreign* format is a finding
about that format and reddening the badge would misattribute the fault.

## It cannot go stale quietly

The report is only worth anything if it is current, so:

- the vendored schemas are pinned to an upstream commit in
  `uacp_interop/schemas/PROVENANCE.json`, and CI fails on drift
- a scheduled workflow re-runs the probe weekly and opens an issue if the result
  stops matching the committed artifact
- `docs/scorecard.svg` in this README is generated from a live run, and CI fails
  if it goes stale

A conformance report that silently tests an old protocol is worse than no
report.

## Known limits

- **One external `$ref` needs a local registry to resolve.** The schemas claim
  canonical URLs on a host that does not resolve, so a plain offline validator
  fetching `agent-descriptor` will attempt a network call. This harness builds a
  `referencing.Registry` for exactly that reason, and the finding is reported
  rather than hidden.
- **No behavioural half.** Everything here is structural: *these two definitions
  cannot both be true*, not *these two implementations produced different
  messages*. That is the harder half and it is not built.
- **The corpus is small**, because the honest entry is one you can source. Three
  ship; a LangGraph entry is [open](https://github.com/wippa-studios/uacp-interop/issues/3).
- **Not yet format-to-format.** It measures loss crossing into a target
  descriptor and loss expressing one in a native tool schema, not directly
  between two foreign formats. That is the obvious next thing and it is not
  built.

## Verifying it

```bash
pip install -e ".[dev]"
pytest -q
```

The suite is mostly meta-tests: each mutates a copy of the input so the defect it
looks for is absent, and asserts the check goes quiet. A check that quietly
stopped working fails the suite rather than passing silently — which has twice
caught a guard in this repo that had quietly stopped guarding anything.

MIT.
