Metadata-Version: 2.4
Name: purrdf
Version: 0.10.0
Classifier: Development Status :: 3 - Alpha
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Rust
Classifier: Typing :: Typed
Requires-Dist: purrdf-rdflib ; extra == 'rdflib'
Provides-Extra: rdflib
Summary: Python bindings for the PurRDF RDF 1.2 kernel, GTS carrier, SHACL, and slice tooling
Home-Page: https://blackcatinformatics.ca/purrdf/
Author-email: "Blackcat Informatics Inc." <paudley@blackcatinformatics.ca>
License-Expression: MIT OR Apache-2.0
Requires-Python: >=3.13
Description-Content-Type: text/markdown; charset=UTF-8; variant=GFM
Project-URL: Homepage, https://blackcatinformatics.ca/purrdf/
Project-URL: Repository, https://github.com/Blackcat-Informatics/purrdf

<!--
SPDX-FileCopyrightText: 2026 Blackcat Informatics® Inc. <paudley@blackcatinformatics.ca>
SPDX-License-Identifier: MIT OR Apache-2.0
-->
# PurRDF for Python

<p>
  <a href="https://pypi.org/project/purrdf/"><img src="https://img.shields.io/pypi/v/purrdf.svg?label=PyPI" alt="PyPI"></a>
  <a href="https://github.com/Blackcat-Informatics/purrdf/blob/main/LICENSING.md"><img src="https://img.shields.io/badge/license-MIT%20OR%20Apache--2.0-blue.svg" alt="License: MIT OR Apache-2.0"></a>
  <a href="https://pypi.org/project/purrdf/"><img src="https://img.shields.io/pypi/pyversions/purrdf.svg" alt="Python versions"></a>
</p>

PurRDF is a from-scratch, dependency-light [RDF 1.2](https://www.w3.org/TR/rdf12-concepts/)
engine — parsers and serializers, SPARQL, SHACL, ShEx, RDFC-1.0 canonicalization, and the
GTS graph-transport container — written in Rust and carried verbatim into Python, JavaScript,
and C. The `purrdf` package is the Python surface of that one engine: the same
byte-identical semantics in every language, including triple terms, reifiers, and
base-direction literals that most incumbent libraries do not carry.

## Install

```sh
pip install purrdf
```

Requires Python 3.13+. Wheels bundle the native extension; no Rust toolchain needed.

## Parse RDF

```python
import purrdf

quads = purrdf.parse(
    '<https://example.org/alice> <http://xmlns.com/foaf/0.1/name> "Alice" .',
    purrdf.RdfFormat.TURTLE,
)
```

`purrdf.parse` accepts Turtle, TriG, N-Triples, N-Quads, TriX, and HexTuples
(`purrdf.RdfFormat`); JSON-LD and RDF/XML travel through the dedicated
`purrdf.from_json_ld` / `purrdf.to_json_ld` and `purrdf.from_rdf_xml` /
`purrdf.to_rdf_xml` converters. All codecs are first-party with
byte-deterministic output.

Configured JSON-LD and YAML-LD use one strict versioned options document. Compile
a reusable context when serializing several datasets:

```python
import json
import purrdf

options = json.dumps({
    "version": 1,
    "mode": "context",
    "prefixes": {"ex": "https://example.org/", "schema": "https://schema.org/"},
})
context = purrdf.CompiledJsonLdContext(options)
jsonld = purrdf.serialize_jsonld(
    nquads,
    format=purrdf.RdfFormat.N_QUADS,
    output_format="jsonld",
    context=context,
)
```

`expanded`, `context`, and deterministic dataset-IRI `derived` modes are
explicit. PurRDF never infers a caller vocabulary or fetches a remote context.

## Project graph, tabular, and research-object carriers

`purrdf.project` and `purrdf.lift` are thin calls into the same Rust projection
engine used by every other surface. Configuration is mandatory, strict JSON:
PurRDF supplies no vocabulary, identity IRI, or resource-limit default.

```python
import json
import purrdf

config = json.dumps({
    "profile": "lpg-csv",
    "config": {
        "rdf_type": "https://example.org/type",
        "scope": {"mode": "all"},
        "limits": {
            "max_artifacts": 16,
            "max_artifact_bytes": 1_000_000,
            "max_total_bytes": 4_000_000,
            "max_archive_bytes": 5_000_000,
            "max_term_depth": 16,
        },
        "execution_limits": {
            "max_input_records": 1_000,
            "max_model_records": 1_000,
            "max_nodes": 1_000,
            "max_edges": 1_000,
        },
    },
})
package = purrdf.project(
    "@prefix ex: <https://example.org/> . ex:alice ex:knows ex:bob .",
    format=purrdf.RdfFormat.TURTLE,
    profile="lpg-csv",
    config=config,
)
lifted = purrdf.lift(package.archive, profile="lpg-csv", config=config)
assert lifted.dataset.quad_count() == 1
print([(loss.code, loss.location) for loss in package.losses])
```

Project profiles are `lpg-csv`, `neo4j-csv`, `open-cypher`, `graphml`,
`csvw-exact`, `csvw-terms`, `okf-terms`, `obo-graphs`, `skos`, `croissant-1.1`,
`ro-crate-1.3`, `datacite-4.6`, `dcat-3`, `dcat-rdf`, `void`, and
`frictionless-data-package-1`. Curated CSVW/OKF terms, OBO Graphs, SKOS, native
DCAT RDF, and VoID are deliberately write-only, ledgered views. Returned
archives are canonical deterministic USTAR bytes and every result carries its always-computed
structured loss records. Research-object contexts, vocabularies, identities,
and profiles are all mandatory caller configuration.

Native RDF dataset descriptions use the same call. The complete JSON names the
output syntax and either a mapped/CONSTRUCT DCAT source or the VoID source
graphs, role vocabulary, dataset-prefix registries, and resource bounds:

```python
from pathlib import Path
import purrdf

source = Path("void-source.trig").read_text()
void_config = Path("void.json").read_text()
description = purrdf.project(
    source,
    format=purrdf.RdfFormat.TRIG,
    profile="void",
    config=void_config,
)
Path("void.tar").write_bytes(description.archive)
```

Portable `void-source.trig`, `void.json`, and `dcat-rdf.json` examples are in
`crates/rdf/tests/fixtures/dataset-description/`.

Attached RO-Crate packaging uses the same call with `assets=` set to a canonical
payload-only USTAR archive and configuration `packaging: "attached"`. The result
contains the exact payloads, deterministic metadata, and self-contained preview;
missing, unowned, reserved, or size-inconsistent members raise `ValueError`.
See the runnable
[`projection_roundtrip.py`](https://github.com/Blackcat-Informatics/purrdf/blob/main/bindings/python/examples/projection_roundtrip.py)
file-producing example.

For large LPG carriers, `purrdf.project_artifacts(...)` invokes a transactional
artifact callback with package/artifact begin, bounded chunk, artifact finish,
commit, and abort events. An optional progress callback receives immutable
`ProjectionProgress` snapshots; callback exceptions abort the package and are
returned unchanged. This path retains the selected canonical LPG model but not
complete artifact bodies or USTAR bytes. See the runnable atomic-directory
[`projection_stream.py`](https://github.com/Blackcat-Informatics/purrdf/blob/main/bindings/python/examples/projection_stream.py)
example.

## Validate with SHACL

The SHACL engine lives at `purrdf.shapes` (mirroring the Rust crate; `purrdf.shacl`
is a back-compat alias):

```python
from purrdf import shapes

report = shapes.validate(shapes_ttl=my_shapes, data_nt=my_data)
print(report["conforms"])
```

Complete SHACL Core, SHACL-SPARQL constraints/targets, and SHACL-AF `sh:rule`
entailment via `shapes.entail(...)`. Reusable parsed shapes are available as
`shapes.Shapes(shapes_ttl).validate_nt(data_nt)`.

## Validate with ShEx

```python
from purrdf import shex

results = shex.validate(
    my_schema_shexc,
    my_data_ttl,
    [("https://example.org/alice", "https://example.org/PersonShape")],
)
print(all(entry["conformant"] for entry in results))
```

The ShEx 2.1 validator passes 1,105/1,105 attempted validation tests of the official
shexTest suite (see the repo's `docs/CONFORMANCE.md`).

## Entailment regimes

The SPARQL entailment regimes live at `purrdf.entail` (mirroring the
`purrdf-entail` Rust crate). It closes a dataset under a regime's own
specification rule table and takes no shapes at all — not to be confused with
`purrdf.shapes.entail(...)`, which applies the SHACL-AF `sh:rule`s a *shapes*
graph declares.

```python
import purrdf
from purrdf import entail

dataset = purrdf.RdfDataset(my_turtle, purrdf.RdfFormat.TURTLE)
closure, report = entail.materialize(dataset, "rdfs", "")
print(closure.to_nquads())
print(report)
```

For callers holding a document rather than a parsed dataset,
`entail.materialize_nt(text, regime, program)` takes N-Triples/N-Quads and returns
`(canonical_nquads, report)`. Both accept the regime as a plain string (`"simple"`,
`"rdf"`, `"rdfs"`, `"owl-rl"`, `"owl-direct"`, `"rif"`, `"d"`) or as
`entail.Regime.RDFS`.

**All seven regimes close; none is refused.** The third argument is the regime's own
rule document. Six regimes take none, so theirs is `""` — and a non-empty one raises
rather than being silently discarded. `"rif"` is the exception: it entails under the
*caller's* rules, which PurRDF does not declare, so its `program` is a normative
RIF-in-XML document:

```python
closure, report = entail.materialize(dataset, "rif", my_rif_xml)
```

`"owl-direct"` takes no program either, and that is a statement rather than an
omission: its extra input is a *query's* class expressions, and this surface closes a
dataset rather than answering a query — so what it runs is the query-independent
tableau augmentation (the classification, the realization, the entailed role
assertions and the `owl:sameAs` identifications the tableau decides about the
ontology's own named terms).

**The report is the second return value and is never optional.** It is a
byte-stable rendering naming which rules fired and how often, which specification
rules did *not* fire, which constructs the run left at a boundary, what it
consumed of the evaluator's fixed ceilings, and the contract hash of the calculus
that ran — so a cached closure minted under a different rule set can be refused
rather than trusted.

The rule tables are readable directly, so coverage is something you measure
rather than something you take on faith:

```python
defined = entail.rules("owl-rl")             # 78 — OWL 2 Profiles §4.3 Tables 4–9
fired = entail.implemented_rules("owl-rl")   # 78
missing = [rule for rule in entail.rules("rdfs") if rule not in entail.implemented_rules("rdfs")]
# [] — RDFS fires 18 of its 18 rules; the gap is legitimately empty
added = entail.extensions("owl-rl")          # ['ext-eq-diff-sym']
```

`extensions(regime)` is a third, disjoint inventory: the rules this build fires
that **no specification table states**. `owl-rl` has one — `ext-eq-diff-sym`,
symmetry of `owl:differentFrom`, sound and shaped exactly like `prp-symp` — and
every other regime has none. It appears in neither `rules()` nor
`implemented_rules()` for any regime, so the 78 above is unaffected by it: those
two are statements about the specification, and firing a sound rule the table
omits does not change what the table says. Asking is a question in its own right
rather than something you learn by materializing a dataset and reading the
report's `extension` line — though the report says the same thing, and the two
cannot drift apart.

`rdfD1`, `rdfD1a`, `rdfs14` and `rdfs14a` are in that fired set and each concludes
about a *fresh* blank node. The restricted chase mints one as a frontier-addressed
Skolem witness and closes under it, so the rules genuinely run — but every
conclusion mentioning a witness is withheld when the closure is materialized back,
because a SPARQL entailment regime draws its answers from the scoping graph and a
minted blank node is not in it. The report says so with a `boundary surrogate`
line rather than with a missing rule, and `completeness` reads
`exact-within-boundaries` rather than `exact`.

**78 / 78 is rule-table coverage, and rule-table coverage is not entailment
conformance.** The two are measured separately and stating only the first is the
overclaim the reasoning report exists to prevent: on W3C's own OWL 2 RL entailment
tests this chase reaches 11 of 27 published positive entailments and correctly
withholds on 23 of 23 negative ones — the latter meaning no unsoundness was found.
Both numbers are true. The full scoreboard, the typed divergence ledger, and every
other suite are in
[`docs/CONFORMANCE.md`](https://github.com/Blackcat-Informatics/purrdf/blob/main/docs/CONFORMANCE.md).

`ValueError` is raised for an unknown regime spelling (the message names the
accepted set), for a `program` that is wrong for the regime — a non-empty one for
any regime but `"rif"`, or one `"rif"` cannot parse as a normative RIF-in-XML
document — and for an exhausted evaluation ceiling. An exhausted ceiling is a
refusal, never a truncated closure handed back as a complete one. Being
`"owl-direct"` or `"rif"` is not itself a refusal: both materialize.

## Description-Logic reasoning services

Materialization is the chase. The OWL 2 Direct-Semantics *reasoner* is a second
lane on the same module — a SHOIQ(D) hypertableau — and every one of its services is on
`purrdf.entail`. Each takes an N-Triples (or N-Quads) document and returns
`(answer, certificate)` as a tuple, so a caller unpacks the evidence rather than
being able to not ask for it:

| Service | Call | Answer |
| --- | --- | --- |
| Consistency | `entail.consistency(data)` | `consistency true` / `false` / `unknown` — `unknown` means the tableau reached its step cap, and is never collapsed to `false` |
| Classification | `entail.classify(data)` | `equivalent`, `subclass` (transitively closed), `direct` (its reduction) and `unsatisfiable` lines |
| Realization | `entail.realize(data)` | `type` lines for the named individuals, then the most specific `direct-type` lines |
| Instance retrieval | `entail.instances(data, class_)` | `instance <term>` lines; `class_` is ONE N-Triples term, angle brackets included |
| Axiom entailment | `entail.entails(data, axiom)` | `entails true` / `false` / `unknown`, then the axiom as it was *read*, so you can see which kind its predicate selected |
| Profile certification | `entail.profile(data)` | `certified <profile>` lines, most restrictive first (`EL`, `QL`, `RL`, `DL`, `Full`) |
| Module extraction | `entail.extract_module(data, signature, method)` | the locality module as canonical N-Quads; `method` is `"bot"`, `"top"` or `"star"` |
| Justification | `entail.justify(data, axiom)` | a minimal subset of the ontology that still entails `axiom`, as canonical N-Quads |
| Proof | `entail.explain_conclusion(data, regime, conclusion)` | `asserted`, `steps`, and one `rule` line per rule the derivation cited |

```python
from purrdf import entail

ontology = (
    "<https://example.org/Cat>"
    " <http://www.w3.org/2000/01/rdf-schema#subClassOf>"
    " <https://example.org/Animal> .\n"
    "<https://example.org/felix>"
    " <http://www.w3.org/1999/02/22-rdf-syntax-ns#type>"
    " <https://example.org/Cat> .\n"
)

answer, certificate = entail.consistency(ontology)
assert answer.strip() == "consistency true"
assert certificate.startswith("purrdf-dl-certificate 1")
assert "completeness decided" in certificate

answer, _ = entail.instances(ontology, "<https://example.org/Animal>")
assert answer.strip() == "instance <https://example.org/felix>"

answer, _ = entail.profile(ontology)
assert answer.splitlines()[0] == "certified EL"
```

The certificate is the point, and there is a different one per kind of evidence.
`consistency`, `classify`, `realize`, `instances` and `entails` render a
`purrdf-dl-certificate 1` block carrying the DL lane's own completeness —
`decided`, `decided-within-boundaries` (an axiom that never became a DL clause,
with each such construct named) or `budget-exhausted`. That is a *different*
notion from the chase report's, which subtracts two rule tables, and it is
rendered under a different banner so neither can be parsed as the other.
`profile` reports no search at all — it is purely syntactic — so it renders a
`purrdf-owl-profile-certificate 1` block ending `one-directional true`: a
certification proves membership, a violation does not prove non-membership.
`extract_module` renders `purrdf-module-extraction 1`, whose `conservative` line
says whether the module is minimal or a sound superset. `justify` renders
`purrdf-justification 1` and `explain_conclusion` renders `purrdf-chase-proof 1`;
both re-check their own answers rather than restating them — `sufficient` and
`minimal` are re-decided over the justification and over each one-axiom-smaller
subset, and a proof's `derived-*` lines are what the checker re-derived from the
proof term, not what the proof claims.

A tableau performs no derivation steps, so `justify` is a *justification* and
deliberately not called a proof; `explain_conclusion` is the chase lane's
genuinely derivational one. They are different kinds of thing rather than two
spellings of one, which is why there is no single `explain`.

Nothing here re-implements the reasoner: every entry point routes through the
same shared boundary the WebAssembly and C hosts call, checked against one
committed golden-vector artifact, so the four hosts return byte-identical results
for the same input.

## rdflib compatibility layer

The package ships an rdflib-shaped API over the native engine:

```python
from purrdf.compat.rdflib import Graph, URIRef

g = Graph()
g.parse(data=my_ntriples, format="nt")
print(len(g), g.serialize(format="turtle"))
```

For a literal, zero-change `import rdflib`, install the opt-in extra:

```sh
pip install purrdf[rdflib]
```

This pulls in the separate [`purrdf-rdflib`](https://github.com/Blackcat-Informatics/purrdf/tree/main/bindings/python-rdflib-shadow)
distribution, whose top-level `rdflib` package re-exports the compat surface, so
existing third-party code doing `import rdflib` / `from rdflib.namespace import RDF`
transparently runs on purrdf. **Caveat:** that shadow claims the `rdflib` import
name and must never be installed alongside the genuine
[`rdflib`](https://pypi.org/project/rdflib/) — the two cannot co-inhabit one
environment. It is a separate distribution (never bundled into the main `purrdf`
wheel) precisely so environments that need the real rdflib simply omit it.

## GTS graph transport and relational exports

GTS is PurRDF's single-file, content-addressed, append-only container for RDF 1.2
graphs. Build one from quads and export it straight to relational stores:

```python
import purrdf

gts_bytes = purrdf.gts_from_quads(my_nquads_bytes, format=purrdf.RdfFormat.N_QUADS)

purrdf.gts_to_sqlite(gts_bytes, "graph.db")
purrdf.gts_to_duckdb(gts_bytes, "graph.duckdb")
files = purrdf.gts_to_parquet(gts_bytes, "out/")
```

The same entry points are grouped under `purrdf.gts` for discoverability.

## Learn more

- Repository: <https://github.com/Blackcat-Informatics/purrdf>
- Project site: <https://blackcatinformatics.ca/purrdf/>
- GTS specification, conformance matrix, and full docs live under
  [`docs/`](https://github.com/Blackcat-Informatics/purrdf/tree/main/docs) in the repo.

Licensed under MIT OR Apache-2.0, at your option.

