Metadata-Version: 2.4
Name: mimir-ontology
Version: 0.2.0
Summary: A lightweight, local-first ontology engine for AI agents.
License-Expression: AGPL-3.0-only
License-File: LICENSE
Requires-Python: >=3.12
Requires-Dist: fastmcp>=3.4.4
Requires-Dist: pydantic>=2.7
Requires-Dist: pyyaml>=6.0
Requires-Dist: typer>=0.12
Description-Content-Type: text/markdown

# Mimir

Mimir is a small, local-first ontology engine for AI agents. You describe your domain in a YAML
schema (types, properties, links), feed the engine claims from your data sources, and it
materializes a typed graph that an agent can query over MCP — with every value traceable to the
claim, and therefore the source, it came from.

The engine is event-sourced. The only source of truth is an append-only log of *claims* ("source
X asserts fact Y"); the queryable graph is a deterministic projection of that log. The same
claims always materialize into the byte-identical graph, nothing is ever mutated in place, and
conflicts between sources are resolved by explicit rules instead of overwrites — the losing
claims stay reachable.

Everything runs in one process on your machine: one SQLite file, one YAML schema, one MCP server.
There is no LLM inside the engine and no network access in the core, by design.

## Install

```
uv add mimir-ontology             # or: pip install mimir-ontology
```

The distribution is named `mimir-ontology`; the package you import and the CLI command are
`mimir`.

## Quickstart

The repository ships a demo world: a fictional plumbing business asserted by two fixture sources
that overlap, disagree and occasionally make mistakes — on purpose. The full walkthrough with
expected output and explanations is in [`examples/plumber/demo.md`](examples/plumber/demo.md);
the short version:

```console
$ uv run mimir init demo-ws
$ uv run mimir schema load examples/plumber/schema.yaml --dir demo-ws
$ uv run mimir ingest examples/plumber/claims.jsonl --dir demo-ws
Ingested claims.jsonl: 88 accepted, 0 duplicates, 0 quarantined.
$ uv run mimir materialize --dir demo-ws
Materialized 88 claims: 43 entities, 35 links, 1 extra properties, 2 quarantined.
Graph hash: grf_42d6a5d8cb69ff50
```

Along the way the engine merged the customers both sources asserted (natural keys), resolved a
phone-number disagreement by observation time, kept an undeclared property in the entity's
`extra` pocket instead of dropping it, and quarantined two unprocessable claims with
machine-readable reasons. Ingest is idempotent: run it again and you get `88 duplicates`.

Inspect the graph from the shell (`--json` for machine output):

```console
$ uv run mimir search Job --filter status=done --limit 3 --dir demo-ws
$ uv run mimir show ent_6f65b3bf159ff099 --dir demo-ws
$ uv run mimir query '{"type": "Job",
    "where": [{"property": "status", "op": "eq", "value": "done"},
              {"link": "invoice", "exists": false}],
    "aggregate": {"function": "sum", "property": "estimated_value"}}' --dir demo-ws
sum(estimated_value) = 40800.0
  (3 matched, 0 missing the property)
```

Serve it to an agent:

```console
$ uv run mimir serve-mcp --dir demo-ws
```

The server speaks MCP over stdio and exposes six read-only tools: `describe_schema`,
`describe_type`, `search`, `query`, `traverse`, `get`. In a clone of this repository,
`.mcp.json` already registers the server for Claude Code — run the quickstart above, start
`claude`, and ask:

> Which finished jobs have no invoice, and what is the total uninvoiced amount?

The agent reads the schema, narrows with a query, and lets the engine compute the sum in the
same call — every figure traceable to a specific claim. When it guesses a link name, an
operator or an enum value that does not exist, the server replies with a structured error
naming the valid options.

## How agents should query

The read surface is designed around how an LLM agent actually works: small working memory, a
tendency to hallucinate structure, and no patience for inventing pagination. The intended
pattern:

1. **Overview first.** `describe_schema` returns the map — every type with its description,
   entity count and links, but no property definitions. It stays small even at a hundred types.
2. **Then the chapter.** `describe_type(type_name)` returns one type in full: properties with
   enums and descriptions, outgoing links, incoming links, entity count.
3. **Narrow.** `search` for simple property filters; `query` for link predicates (both
   directions), ordering and pagination. Results are compact cards; every list response carries
   the true `total`, and a `cursor` continues exactly where the page ended.
4. **Walk.** `traverse` follows declared links from one entity. Expansion is bounded and every
   truncation is reported (`complete: false`) — nothing is ever cut silently.
5. **Let the engine count.** Every count, sum, average, minimum, maximum and group breakdown is
   one `query` call with an `aggregate` block. An agent adding up figures it paged through is
   the failure mode this surface exists to prevent; aggregation walks are never truncated, and
   float sums are bit-stable.
6. **Drill to provenance.** List views show plain values; `get(type_name, entity_id)` returns
   the full dossier — every value with its `observed_at` and `claim_id`. Every aggregate names
   the filter that reproduces its input set, so every number can be re-derived and audited.

Wrong turns fail loudly and teach: an unknown field, an inapplicable operator or a typo'd enum
value returns a structured error listing the valid options instead of silently matching
nothing.

## How it works

```
claims.jsonl ──ingest──▶  claim log            append-only, content-addressed (SQLite)
                             │
                        materialize            pure function: (schema, claims) → graph;
                             │                 same input, byte-identical output
                             ▼
                        typed graph            entities + links, every value with provenance
                             │
                        serve-mcp              read-only MCP tools for agents
```

Five primitives: **Claim** (an atomic assertion by a source), **Schema** (your ontology, loaded
from YAML — the engine hardcodes no domain type), **Entity** and **Link** (the materialized
graph), **Provenance** (who said it, based on what, when). Identity is content-derived
throughout: the same natural keys land on the same entity no matter which source asserted them,
and re-materialization can rebuild the whole graph from the log at any time.

This release adds the agent-facing read surface on top of the engine core: the two-level
schema map, the structured `query` tool with schema-validated predicates and aggregations,
cursor pagination, and bounded reads with loud truncation. Storage and materialization are
unchanged — 0.1 workspaces open as-is.

## Development

Requires Python 3.12+ and [uv](https://docs.astral.sh/uv/).

```
uv sync                  # install dependencies
uv run pytest            # run tests
uv run ruff check .      # lint
uv run ruff format .     # format
uv run mypy .            # strict type check
uv run mimir --help      # CLI
```

## License

AGPL-3.0-only. See `LICENSE`.
