Metadata-Version: 2.4
Name: schemair
Version: 0.3.0
Summary: Python bindings for the SchemaIR portable contract protocol
Author: SchemaIR contributors
License-Expression: Apache-2.0
Project-URL: Repository, https://github.com/schemair/schemair
Project-URL: Documentation, https://github.com/schemair/schemair#readme
Project-URL: Issues, https://github.com/schemair/schemair/issues
Keywords: schema,types,contracts,protocol
Classifier: Development Status :: 3 - Alpha
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Topic :: Software Development :: Libraries
Requires-Python: >=3.11
Description-Content-Type: text/markdown
License-File: LICENSE
Dynamic: license-file

# SchemaIR for Python

[![PyPI version](https://img.shields.io/pypi/v/schemair)](https://pypi.org/project/schemair/)
[![CI](https://github.com/schemair/schemair/actions/workflows/ci.yml/badge.svg)](https://github.com/schemair/schemair/actions/workflows/ci.yml)
[![License: Apache 2.0](https://img.shields.io/badge/License-Apache%202.0-blue.svg)](https://github.com/schemair/schemair/blob/main/LICENSE)

**Types as data.**

SchemaIR is a language-independent protocol and intermediate representation
for portable data contracts, operation definitions, and type calculations.

The [`schemair` package](https://pypi.org/project/schemair/) is the Python
implementation. It represents contracts, operations, and type expressions as
serializable data.

For the protocol's theory, IR model, design philosophy, standard vocabulary,
cross-language boundaries, and conformance rules, see the [SchemaIR project
README](https://github.com/schemair/schemair#readme) and the [protocol
specification](https://github.com/schemair/schemair/tree/main/spec). This page
focuses on using SchemaIR from Python.

## About this implementation

SchemaIR declarations are ordinary JSON-compatible dictionaries. You can store
a contract with a project, exchange it with another tool, and inspect or
calculate types without executing the application behind the contract.

This binding implements the shared protocol and conformance rules rather than
Python annotations or Pydantic models. See the [implementation
matrix](https://github.com/schemair/schemair/blob/main/IMPLEMENTATIONS.md) for
coverage and status.

## Install

Use Python 3.11 or newer. Install the published package from PyPI:

```sh
python -m pip install schemair
```

To work on the implementation from this repository, install the local package
instead:

```sh
python -m pip install -e ./python
```

The package is published on PyPI as `schemair`.

## Quick start

```python
import json
from schemair import authoring as s
from schemair import (
    check_schema_relation,
    evaluate_expression,
    validate_payload,
    validate_schema,
)

# A SchemaNode describes the concrete output shape as data.
user = s.record([
    s.required_field("id", s.string()),
    s.required_field("name", s.string()),
])

# Reload the declaration without importing its original application.
loaded = json.loads(json.dumps(user))
assert validate_schema(loaded).status == "accepted"

# Build a SchemaTypeExpr, then evaluate it before executing an operation.
selection = s.field_of(loaded, "name")
name_type = evaluate_expression(selection)
assert name_type.status == "evaluated"
assert check_schema_relation(name_type.value, s.string()).status == "assignable"

# Validate actual values independently of type calculation.
payload = validate_payload({"id": "123", "name": "Ada"}, loaded)
assert payload.status == "accepted"
```

Python exposes snake_case helpers directly from `schemair.authoring` and also
provides `schemair.authoring.schema` and `schemair.authoring.expression`
namespaces that mirror the TypeScript package's authoring API. Wire fields
retain the protocol's camelCase spelling, such as `elementSchema`, `refPath`, and
`semanticPath`.

## API and results

| Task | Entry points | Results |
| --- | --- | --- |
| Validate declarations | `validate_schema`, `validate_operation`, `validate_operation_container`, `validate_structured_operation`, `validate_structured_operation_container`, `validate_schema_type_expr`, `validate_schema_type_term` | `accepted`, `rejected` |
| Load references | `resolve_schema_references`, `resolve_operation_closure`, and their `_async` variants (from `schemair.resolve`) | `resolved`, `incomplete`, `rejected` |
| Validate data | `validate_payload` | `accepted`, `rejected`, `unknown` |
| Compare contracts | `check_schema_relation`, `check_operation_relation` | `assignable`, `incompatible`, `unknown`, `rejected` |
| Report compatibility | `report_schema_compatibility`, `report_operation_compatibility` | `compatible`, `breaking`, `unknown`, `invalid` |
| Calculate types | `evaluate_expression`, `evaluate_schema_type_term` | `evaluated`, `incomplete`, `rejected` |
| Compare type terms | `check_type_term_relation` | Relation statuses |
| Infer variables | `solve_type_variables` | `solved`, `unknown`, `incompatible`, `rejected` |

Calculation functions return a `Result` with `status`, `issues`, and `value`.
Reference loaders return `SchemaReferencesResult` with `status`, `issues`, and
`graph`.
Expression results put the evaluated term in `value` and expose unresolved
variable names through `unresolved_type_vars`; solver results put the binding
dictionary there. An evaluated term may still be an expression, so check its
shape before using it as a concrete schema. Diagnostics contain a `code`,
`path`, `message`, and optional `details`.

`validate_expression(value, term=True)` remains available for validating either
a schema or an expression. Definition validation does not execute operations
or resolve external references.

Compatibility reports run both substitution directions for a previous and next
declaration. The `value` contains `previous_to_next` and `next_to_previous`
checks, each with its mapped status and original relation result. The report
does not select migrations or a publication policy.

## Solve a type variable

```python
from schemair import authoring as s, evaluate_expression, solve_type_variables

solution = solve_type_variables(
    ["T"],
    [{"source": s.number(), "target": s.type_var("T")}],
)
assert solution.status == "solved"
output = evaluate_expression(s.array_of(s.type_var("T")), type_vars=solution.value)
assert output.value == {"kind": "array", "elementSchema": s.number()}
```

The solver infers from supported source-to-target constraints. It is finite,
not a complete host-language generic type checker. Inspect `unknown`
and diagnostics when candidate relations cannot be proved.

## Operations and host callbacks

An operation is a concrete declaration of input, output, errors, and emitted
channels. Structured payloads contain `SchemaNode` values. Functions and
unresolved expressions are not operation payload schemas. See the
[operation model](https://github.com/schemair/schemair/blob/main/spec/protocol/operations.md).

Calculations are synchronous; reference loading supports both execution modes.
The `schemair.resolve` module exposes `resolve_schema_references` and
`resolve_operation_closure` for synchronous resolvers, plus
`resolve_schema_references_async` and `resolve_operation_closure_async` for
synchronous or asynchronous resolvers. Resolvers return schemas or resolution
decisions and receive a tuple of path segments:

```python
from schemair import authoring as s, validate_payload
from schemair.resolve import resolve_schema_references

def resolve_ref(path):
    return s.string() if path == ("UserName",) else None

entry = s.ref(["UserName"])
loaded = resolve_schema_references(entry, resolve_ref=resolve_ref)
if loaded.status != "resolved":
    raise ValueError(loaded.issues)
result = validate_payload("Ada", entry, refs=loaded.graph.refs)
assert result.status == "accepted"
```

For database or network access, use an async loader inside an async host workflow:

```python
from schemair.resolve import resolve_schema_references_async

async def validate_from_catalog(entry, value, fetch_schema):
    loaded = await resolve_schema_references_async(entry, resolve_ref=fetch_schema)
    if loaded.status != "resolved":
        raise ValueError(loaded.issues)
    return validate_payload(value, entry, refs=loaded.graph.refs)
```

Both entry points share traversal, caching, limits, and diagnostics. The sync
loader refuses awaitable results with an `incomplete` diagnostic; the async
loader accepts both immediate and awaitable results. Only reference loading
awaits host I/O; calculations consume the resulting in-memory snapshot.

The loader recursively follows references inside resolved schemas. Its
`ResolvedSchemaGraph` contains `refs` entries shaped as
`{"refPath": ["catalog", "User"], "schema": schema}` and `cycles` as path
arrays. It preserves reference edges rather than inlining schemas. Inspect
`resolved`, `incomplete`, or `rejected` before using the snapshot; loader
limits bound traversal depth, steps, and newly loaded schemas.

Pass the same `refs` snapshot to payload validation, schema and operation
relations, expression evaluation, type-term relations, compatibility reports,
and the solver's `context`. These APIs are synchronous. The expression
interpreter accepts preloaded references only and does not invoke a resolver.

Payload and schema relation APIs also accept synchronous `resolve_ref`,
`semantic_provider`, and `constraint_provider` policies. Modern payload
providers receive `(semantic_path, value, context)` or
`(constraint_path, args, value, context)`; the legacy two-argument forms remain
available. Relation providers receive source and target paths or constraint
lists, with an optional context. Term relations additionally accept
`host_type_relation(source_path, target_path)` for opaque host types.

Prepare external policy data before invoking the core. Awaitable policy
results do not prove validation or assignability. Database uniqueness,
permissions, and other business checks belong in the host workflow.

## Standard vocabulary

`schemair.standard` exposes `STANDARD_SEMANTIC_PATHS`,
`STANDARD_CONSTRAINT_PATHS`, definition maps, `standard_vocabulary_registry`,
and `classify_semantic_path`. Standard validation is opt-in:

```python
from schemair import authoring as s, standard, validate_payload

email = s.string(semantic_path=standard.STANDARD_SEMANTIC_PATHS["string"]["email"])
result = validate_payload("ada@example.com", email, **standard.standard_payload_context)
assert result.status == "accepted"
```

`standard_schema_satisfiability`, `refine_standard_schema`, and
`intersect_standard_schemas` expose the shared conservative standard constraint
profile, including binary64 interval
boundaries, exact float-derived `multipleOf` checks, pattern handling, and
standard argument domains. The relation context implements the shared
semantic narrowing and constraint implication rules; host regex/date behavior
still follows Python's libraries.

Host policies matter: Python uses its date/time, IP, and regex libraries, and
Python integers are unbounded, but
standard `multipleOf` converts numbers to binary64 and compares exact rational
representations. Do not assume arbitrary-precision decimal behavior or
identical regex/date acceptance across hosts. See
[standard vocabulary host policies](https://github.com/schemair/schemair/blob/main/spec/protocol/standard-vocabulary.md).

## Integrations

```python
from schemair import authoring as s, project_json_schema

schema = s.record([s.required_field("name", s.string())])
projection = project_json_schema(schema, target="draft-2020-12")
assert projection.fidelity == "exact"
print(projection.schema, projection.diagnostics)
```

Targets are `draft-2020-12`, `draft-07`, and `openapi-3.0`.
`schema_node_to_json_schema` and `schema_node_to_openapi_schema` return a
diagnostic-bearing projection. `to_json_schema` returns only the dictionary;
prefer the projection API when fidelity matters. Supported options include
both Python `snake_case` names and protocol-compatible `camelCase` aliases
for ref and mapper strategies. Fidelity results are `exact`, `lossy`, or
`unsupported`.

`to_standard_schema(schema, **context)` returns a `~standard` adapter whose
`validate` callable is synchronous and preserves accepted input values. Load
external references before constructing the adapter.
`schemaIR_to_standard_json_schema(schema)` exposes `jsonSchema.input` and
`jsonSchema.output` callables; each takes a target string and raises `ValueError`
for unsupported projections. These are Python callable surfaces.

## Development and verification scope

From `python/`, run:

```sh
python -m unittest discover -s tests -v
python -m compileall -q schemair
```

Tests read shared schema, operation, expression, solver, and vocabulary vectors
from `../conformance/`, alongside local integration tests. Shared expression
vectors assert evaluated term contents and unresolved variables; solver vectors
assert bindings; rejected cases assert diagnostics; local tests cover provider
decision shapes, reference loading, resolver cycles, and interpreter budgets.

The binding does not execute operations, supply a scheduler or reference
catalog, or provide a compile-time inference utility. See the
[roadmap](https://github.com/schemair/schemair/blob/main/ROADMAP.md) for remaining parity and ecosystem work.
