Metadata-Version: 2.4
Name: uacp-interop
Version: 0.3.0
Summary: Interop conformance probe for UACP: computes where UACP's schemas, prose and reference implementations diverge, and where a UACP agent definition cannot survive a round trip through a framework-native tool format.
License-Expression: MIT
Project-URL: Homepage, https://github.com/wippa-studios/uacp-interop
Project-URL: Repository, https://github.com/wippa-studios/uacp-interop
Project-URL: Issues, https://github.com/wippa-studios/uacp-interop/issues
Project-URL: Protocol, https://github.com/wippa-studios/wippa-uacp
Keywords: uacp,interoperability,conformance,agent,json-schema,protocol
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Software Development :: Testing
Classifier: Topic :: Software Development :: Quality Assurance
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: jsonschema>=4.0
Provides-Extra: dev
Requires-Dist: pytest>=7.0; extra == "dev"
Requires-Dist: build>=1.0; extra == "dev"
Requires-Dist: twine>=5.0; extra == "dev"
Dynamic: license-file

# uacp-interop

Interop conformance probe for [wippa-uacp](https://github.com/wippa-studios/wippa-uacp).

UACP's own L0 to L3 conformance levels are a feature checklist. Passing proves an
implementation matches its author's reading of the spec, and says nothing about
whether two implementations built by different people interoperate, which is the
premise the project rests on.

This harness tests the three things the checklist cannot:

1. **Schema versus prose.** Build a message the spec *requires* an
   implementation to accept, validate it against the shipped schema, report the
   rejection.
2. **Schema versus the implementations.** Check that each reference
   implementation can represent what the schema declares, and that the two
   implementations model the same envelope. This class found a silent credential
   downgrade between the TypeScript and Python implementations.
3. **Cross-format divergence, both directions.** Map a framework-native agent
   description into a UACP descriptor, validate the artifact actually produced,
   and separately test whether a UACP capability survives being expressed in a
   framework's native tool format.

## Results

See [FINDINGS.md](FINDINGS.md). Against `wippa-uacp@b3fe7a1f`: **2 schema
conflicts and 12 cross-format findings (3 blocker, 7 major, 4 minor).**

UACP itself is at **zero blockers**: its schemas now agree with its normative
text, and the TypeScript and Python reference implementations model the same
envelope. All three remaining blockers are measured loss in a *foreign* format,
which is what the harness is for — a UACP capability using `const` and `oneOf`
that an OpenAI function-calling block cannot express, and a corpus entry whose
agent id violates UACP's identifier pattern.

Every blocker this harness once reported against its own protocol has been fixed
upstream and is retained here as a regression guard, so a regression re-opens it
rather than passing silently. See [Fixed upstream](#fixed-upstream).

## Install

```bash
uvx uacp-interop                       # or: pipx run uacp-interop
```

Nothing to clone and no virtualenv to build. From a checkout:

```bash
pip install -e .                       # runtime only
uacp-interop                           # the report
```

## Quick start

```bash
uacp-interop                            # scorecard, then the full report
uacp-interop --scorecard                # the one-glance per-section table
uacp-interop --markdown                 # full report as Markdown, for an issue or a PR
uacp-interop --json                     # machine-readable
uacp-interop --badge                    # status-badge JSON for the worst UACP finding
uacp-interop --fail-on-blocker          # exit 1 if any blocker is found

pip install -e ".[dev]" && pytest -q    # only to run the tests
```

Every renderer derives its counts from the same findings via `report.tally()`,
and a test asserts they cannot disagree — a renderer that tallied severities
independently would be the double-counting defect from 0.1.0 in a new disguise.
`--badge` deliberately reports only the spec and implementation audit: a
blocker in a *foreign* format is a finding about that format, and reddening the
badge for it would misattribute the fault to UACP.

The scorecard, which is the shareable view:

```
uacp-interop 0.2.0

  section                  findings  blocker  major  minor
  -----------------------  --------  -------  -----  -----
  uacp-spec                       2        0      2      0
  autogen-agentchat               6        0      4      2
  openai-function-calling         4        1      1      2
  uacp-native                     2        2      0      0
  -----------------------  --------  -------  -----  -----
  TOTAL                          14        3      7      4
```

A captured run of the current corpus is committed at
[docs/sample-report.txt](docs/sample-report.txt), so you can see the output
before installing anything.

## Method

Losses are computed, not asserted. `SCHEMA_VIOLATION` findings are observed by
running the real validator over the artifact an adapter produced, and it is the
*only* source of them: the adapters do not predict schema violations from
hand-copied patterns, because a copied pattern is a second source of truth that
can drift from the schema, and because predicting as well as observing counted
each such defect twice. Divergences in the reverse direction are found by
walking the source JSON Schema and collecting keywords outside the target
format's supported subset.

Each field yields at most one finding. When both reference implementations
disagree the same way about a field, that is one defect with two witnesses and
both are cited in the evidence line, not two findings. `duplicate_findings()`
enforces this over the whole report, and the suite asserts it holds for the real
corpus and that the guard is able to see a duplicate when one is injected.

## Fixed upstream

The harness was written against its own protocol, so its first reports were
mostly defects in UACP. All of them are fixed, and each is now a test:

| Was | Defect | Fixed in |
|---|---|---|
| 2 blockers | `metadata.auth` modelled in TypeScript, absent in Python — an authenticated message decoded silently as anonymous | `wippa-uacp@b3fe7a1f` |
| 3 majors | `capability` in the schema's global `required`, so `heartbeat`, `register` and `discover` had to invent a placeholder. Both implementations sent the literal `_internal` | `wippa-uacp@b3fe7a1f` |
| 1 major | `flowControl` specified in SPEC.md and modelled by both implementations, declared by no schema | `wippa-uacp@b3fe7a1f` |
| 1 blocker | audit-trail credential persistence in plaintext | `wippa-uacp@e0463b4` |

That sequence is the argument for the tool: it found real defects in the
protocol it ships alongside, including a silent security downgrade, and the
fixes are verified by the same harness that reported them.

Every corpus entry declares a `confidence`:

- `observed-serialization`: a real serialized artifact was read
- `documented-api`: taken from official documentation of the public API
- `inferred`: our modelling choice, not a documented shape

Findings derived from weaker entries are weaker evidence, and the CLI says so.
The AutoGen entry currently models a documented constructor surface rather than
an observed serialization, which makes it the weakest evidence in the corpus.

Each check is backed by a meta-test that mutates a copy of the schema and
confirms the check *fires*, so a check that quietly stopped working fails the
suite rather than passing silently.

## Vendored schemas

`uacp_interop/schemas/` holds a copy of UACP's JSON Schemas so the harness can
validate offline. A conformance report that silently tests an old protocol is
worse than no report, so the upstream commit is pinned in
`uacp_interop/schemas/PROVENANCE.json` and CI fails on drift:

```bash
python scripts/refresh_schemas.py            # refresh and re-pin
python scripts/refresh_schemas.py --check    # verify only, exit 1 on drift
```

## Known limits

- No behavioural half. Everything here is structural: "these two definitions
  cannot both be true", not "these two implementations produced different
  messages". That is the harder half and the one not yet built.
- No LangGraph, CrewAI or OpenAI Assistants entries. The LangGraph docs did not
  yield a sourceable serialization format, and an entry that cannot be sourced
  honestly is worse than a missing one.
- The implementation audit records observed field sets as data with provenance
  rather than parsing the dataclass and TypeScript interface at runtime.
  Re-reading them is a manual step, so they can drift.

## License

MIT.
