Metadata-Version: 2.4
Name: provem
Version: 0.1.0
Summary: Governed, GDPR-native memory for AI agents: a governance layer (erasure, tenant isolation, prompt-injection defense, audit) over Mem0, Zep, or any store.
Author-email: Bernhard Jackiewicz <bj@techport.io>
License-Expression: MIT
Project-URL: Homepage, https://github.com/BernhardJackiewicz/provem
Project-URL: Repository, https://github.com/BernhardJackiewicz/provem
Project-URL: Issues, https://github.com/BernhardJackiewicz/provem/issues
Keywords: ai-agents,agent-memory,llm-memory,ai-governance,gdpr,prompt-injection,mem0,zep,rag,compliance
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Topic :: Software Development :: Libraries :: Python Modules
Classifier: Operating System :: OS Independent
Requires-Python: >=3.9
Description-Content-Type: text/markdown
License-File: LICENSE
Provides-Extra: graphiti
Requires-Dist: graphiti>=0.1.13; extra == "graphiti"
Provides-Extra: mem0
Requires-Dist: mem0ai>=2.0.0; extra == "mem0"
Provides-Extra: zep
Requires-Dist: zep-cloud>=3.0.0; extra == "zep"
Provides-Extra: letta
Requires-Dist: letta-client>=1.10.3; extra == "letta"
Provides-Extra: all
Requires-Dist: graphiti>=0.1.13; extra == "all"
Requires-Dist: mem0ai>=2.0.0; extra == "all"
Requires-Dist: zep-cloud>=3.0.0; extra == "all"
Requires-Dist: letta-client>=1.10.3; extra == "all"
Dynamic: license-file

# Provem: GDPR-native governed memory for AI agents

[![PyPI](https://img.shields.io/pypi/v/provem.svg)](https://pypi.org/project/provem/)
[![Python](https://img.shields.io/badge/python-3.9%2B-blue.svg)](https://www.python.org/)
[![License: MIT](https://img.shields.io/badge/license-MIT-green.svg)](LICENSE)
[![Tests](https://img.shields.io/badge/tests-493%20passing-brightgreen.svg)](tests/)

**Governed, GDPR-native memory for AI agents.** A governance layer (right-to-erasure, tenant isolation, injection defense, tamper-evident audit) that runs on top of *any* memory store (or as its own). On recall it holds its own: a clear win over Zep and roughly a tie with Mem0 on LoCoMo. Its real job is compliance: it drives violations to zero in a reproducible, self-authored closed-loop benchmark. Every number below is replayable from frozen artifacts at zero cost.

## The problem: recall is solved, governance is the hard part

Take a recruiting agent. It picks up information from everywhere: email, the
CRM, Slack, web pages, PDFs, meetings, the user chat. All of it lands in a
memory, and most agent memories work the same way: store → embed → retrieve.
This works surprisingly well.

Then the memory grows, and the hard question changes. It is no longer
*"can I find it?"*, it becomes ***"am I allowed to use it?"***

A candidate writes: *"Please delete my salary expectation."* Three months later
the agent uses it anyway. The retrieval was perfect. **The memory was legally
wrong.**

That example is deliberately simple. The realistic version is quieter and
harder. To build rapport, a recruiter jots down what a candidate volunteers in
small talk: that they are planning a family, their religion, who they live with.
It goes into the notes, gets embedded, and becomes just another retrievable
memory. Now the agent silently holds **special-category data** (GDPR Article 9
covers health, religious beliefs, and sexual orientation) and traits a hiring
decision may never rest on (Germany's AGG protects gender, religion, and sexual
identity, and even asking about pregnancy or family planning is already
unlawful). The company was never allowed to collect it and is not allowed to act
on it, yet the memory will happily serve it into the next *"is this candidate a
good fit?"* answer. That is prohibited processing plus discrimination liability,
and **better embeddings make it worse**: they surface the sensitive note more
reliably. The same pattern shows up with a patient's offhand health remark, a
customer's political comment, or any sensitive fact a user drops once and the
store keeps forever.

The same shift shows up everywhere once agents hold data about real people and
real companies:

- A scraped page plants a wrong fact, and it quietly becomes "knowledge" and
  re-fires in every future answer (this is how MINJA/AgentPoison-style attacks
  work).
- Customer A's data surfaces in customer B's session.
- Someone asks *"why did the agent say that?"*, and there is no answer.
- The agent states a stale fact confidently instead of saying "I don't know",
  and in a multi-step workflow that error repeats (measured below: one bad
  memory re-fires in 2.12 later steps of the scripted workflow on average).

Better retrieval fixes none of this. These are governance problems, and today
they are mostly "solved" in prompts (unenforceable) or in per-app code
(unauditable).

## What Provem is: a governance layer for AI agent memory

Provem is an open-source research project: a governance layer for AI agent
memory that sits between the agent and whatever actually stores the memories
(Mem0, Zep, or your own database).

```text
   Agent  (any model: GPT, Claude, Llama, Gemini, ...)
     |
     v
   Provem   decides:  may this be stored?   may it be read?
     |                whose data is it?     was it deleted?
     |                do I trust the source?
     |                should I rather say "I don't know"?
     v
   any backend:  SQLite · BM25 · Mem0 · Zep/Graphiti · your own DB
```

Two design decisions make the pieces swappable:

- **Backend-agnostic.** The storage engine is a plug-in behind a small
  `MemoryBackend` protocol. Erasure, scoping, trust, and audit live *above* the
  store and do not change when you swap it. Note the direction: sitting above a
  store, the layer can *filter* what that store returns (serve / refuse /
  abstain) but never *improve its recall*, so wrapping Mem0 or Zep gives you
  their recall plus Provem's governance, not higher recall.
- **Model-agnostic.** The rules have nothing to do with which LLM you run.
  Support on GPT, legal on Claude, internal tools on Llama? The erasure duty is
  the same, tenant isolation is the same, the audit trail is the same. One
  governance layer, instead of re-implementing compliance once per model and
  once per memory vendor.

Concretely, the layer enforces:

| Requirement | Without it | Provem mechanism |
|---|---|---|
| Right to erasure (GDPR Art. 17 & co.) | Agent quotes deleted data months later | Erasure enforced *at recall*, tenant-scoped, with erasure certificates |
| Untrusted sources | A planted fact becomes permanent "knowledge" | Provenance + trust tagging, injection quarantine at write, trust-weighted conflict resolution at read |
| Tenant isolation | Customer A's data in customer B's session | Hard scope isolation per tenant and entity, tested adversarially |
| Auditability | *"Why did the agent say that?"* has no answer | SHA-256 hash-chained audit log; every serve/refuse decision carries reasons and provenance |
| Calibrated uncertainty | Confident stale answers that compound | Abstention as a first-class outcome; retention windows enforced |
| Domain rules | Recruiting ≠ pharma ≠ finance, hardcoded per app | Declarative per-tenant compliance profiles (JSON/YAML) |

This is a research project first: every claim on this page has a reproducible
benchmark behind it, negative results are documented alongside the wins, and
the whole evidence chain replays from frozen artifacts at zero cost.

## The two results, and how to read them

Two results, one system, every number reproducible from frozen artifacts at zero cost:

| Axis | Result |
|---|---|
| **Governance** | Compliance violations **240 → 0**, memory-poisoning success **100% → 0%**, silent compounding errors **72.6% → 0.0%** (paired: governance flips 497 of 960 trajectories, loses 0; deterministic, no API key) |
| **Memory** (recall) | **Provem's own memory vs theirs, head-to-head on full LoCoMo**: 0.614 vs 0.565 vs 0.449 answerable accuracy. That is a **clear win over Zep** (+16.5 pts) and **about a tie with Mem0** (+4.9 pts, p = 2.5×10⁻⁴ over the full set but **not** significant on the held-out split). Ordering holds under an independent Claude judge and under Mem0's own published judge prompt (0.772 / 0.716 / 0.649). Measured in the **dense tier (needs an embeddings key)**; both baselines were re-measured after fixing input bugs that had understated them |

### How to read this: two modes, and what the numbers mean

Provem runs in **two modes**, and the recall numbers apply to only one of them:

1. **As your memory**: Provem's *own* retrieval (the "dense" tier). This is what the **0.614** is: our engine measured head-to-head against Mem0's and Zep's engines, each standalone. Honest margins: clearly ahead of Zep, essentially level with Mem0.
2. **As a governance layer over someone else's store**: plug Mem0, Zep, or your own DB in as the backend. Then you keep **that backend's** recall and Provem adds erasure, tenant isolation, injection defense, and audit *on top*.

**Governance filters; it never invents recall.** Putting Provem in front of Mem0 does **not** raise Mem0's recall: a gate can only serve, refuse, or say "I don't know", never retrieve a memory the store missed. So "better recall" always means mode 1 (our own engine); it never means "we make Mem0 or Zep recall better." What we add to *their* stores is compliance and safety, not more recall.

So the honest one-liner: **on recall we're a peer of Mem0 and ahead of Zep; the reason to run Provem is the governance layer that works on any of them.**

## Quick start

```bash
pip install -e .            # dependency-free core, Python 3.9+
```

Wrap a memory in four lines:

```python
from cognitive_memory.reliability import GovernedMemory, NaiveBackend, Scope

mem = GovernedMemory(NaiveBackend(), policy="recruitment")   # or "pharma", "finance", custom JSON/YAML
mem.remember("cand_1 salary_target 120k", subject="cand_1", relation="salary_target",
             object="120k", tenant="acme", entity="cand_1", source="recruiter", trust=0.9)
mem.remember("cand_1 salary_target 80k",  subject="cand_1", relation="salary_target",
             object="80k",  tenant="acme", entity="cand_1", source="scraper_tool", trust=0.4)  # poisoning attempt
mem.forget("migraine", Scope("acme", "cand_1"))              # GDPR erasure, enforced at recall, certificate issued

mem.recall_value("cand_1 salary_target", tenant="acme", entity="cand_1").answer
# -> "120k"   (trusted source wins; never the poisoned, erased, or cross-tenant value)
```

Run it as a configurable MCP server (one server, many tenants, per-tenant
compliance profiles, hash-chained audit):

```bash
PYTHONPATH=src python3 -m cognitive_memory mcp-serve --config examples/mcp/server_config.json
```

Reproduce the headline numbers:

```bash
PYTHONPATH=src python3 -m cognitive_memory reliability --seeds 1,2,3,4,5,6,7,8,9,10 --scenarios 96
sh scripts/fetch_locomo.sh      # one-time: fetch LoCoMo (CC BY-NC 4.0, not redistributed here), sha256-verified
sh scripts/replay_report.sh     # every three-system number, €0, from frozen caches
```

## Choose your tier (honest numbers)

| Tier | LoCoMo answerable | Governance | Requirements |
|---|---|---|---|
| stdlib retrieval + LLM answerer | 0.388 (below Mem0) | full | LLM API key for answering; no pip dependency |
| fully keyless (extractive answers) | 0.21 (below the 0.24 no-memory baseline) | full | none (no key, no pip dependency, air-gap-safe) |
| **dense (recommended)** | **0.614** (> Mem0 0.565 > Zep 0.449) | full | embeddings API key (cents per conversation, disk-cached) |
| local embeddings | ~0.47 to 0.50 (projected, unbuilt) | full | planned: pip extra, no API key |

The keyless tier is the zero-dependency governance layer and demo path: it
does not compete on recall (its extractive answering scores below a no-memory
abstain baseline on LoCoMo, measured 0.207 to 0.212), and we say so. The dense
tier is the benchmarked configuration.

## Evidence 1: the governance benchmark (deterministic, no API key)

Same recall backend, same 960 trajectories, same seeds; the only difference is
whether governance is on:

| Arm | Task success | Silent (compounding) errors | Poisoning success | Compliance violations | Benign accuracy |
| --- | ---: | ---: | ---: | ---: | ---: |
| ungoverned memory | 0.375 | 72.6% of steps | 100% | 240 | 1.000 |
| **+ Provem governance** | **0.893** | **0.0%** | **0%** | **0** | **1.000** |

Paired per trajectory: governance flips 497 of 960 trajectories from fail to
pass and never loses one the ungoverned arm wins (497/0 discordant). We report
counts, not p-values, here on purpose: the benchmark is a deterministic,
author-designed simulation, so a McNemar p-value would only restate the chosen
scenario count. Benign accuracy stays 1.000, by construction: benign probes
are designed so no governance mechanism can fire, which verifies governance
does not interfere, not that it is calibrated on hard cases. With a stochastic
agent, governance buys +43 to +52 points of end-to-end task success at every
skill level. Attack models are simplified analogs of the MINJA
([arXiv:2503.03704](https://arxiv.org/abs/2503.03704)) and AgentPoison
([arXiv:2407.12784](https://arxiv.org/abs/2407.12784)) mechanisms; see the
attack-family results and their limits, including the same-channel poisoning
boundary governance cannot catch. The agent is a deterministic/
noise-parametrized policy by design: it isolates the memory layer's causal
contribution; it is not an end-to-end LLM claim. Full method and limits:
[`docs/agentic_reliability_benchmark.md`](docs/agentic_reliability_benchmark.md),
[`docs/reliability_results.md`](docs/reliability_results.md).

## Evidence 2: memory quality vs Mem0 and Zep (full LoCoMo, paired, three judges)

All three systems ingested the same 10 LoCoMo conversations (5,882 turns) and
answered the same 1,986 questions with the **identical answerer model**; only
the memory differs. Mem0 ran on its own platform pipeline; Zep ran on Zep Cloud,
configured per **Zep's own published evaluation checklist** (proper user model,
native `created_at` timestamps, parallel edge+node graph searches, chronological
ingestion), with ingestion read-back verified and full graph completion
hard-gated before evaluation. Both baseline arms were **re-measured** after an
audit found two input bugs that understated them (Mem0 got a doubled date
prefix; Zep ingested sessions in lexical order); full v1→v2 disclosure in
[`docs/measurement_changelog.md`](docs/measurement_changelog.md).

| Scoring regime | **Provem (dense)** | Mem0 | Zep |
|---|---|---|---|
| Strict binary judge (gpt-5) | **0.614** | 0.565 | 0.449 |
| Independent cross-vendor judge (claude-opus-5) | **0.502** | 0.368 | 0.304 |
| **Mem0's own published judge prompt** (partial credit, 14-day date tolerance) | **0.772** | 0.716 | 0.649 |
| Abstention on 446 adversarial questions | **0.863** | 0.830 | 0.722 |

**The ordering is invariant under all three judges.** All pairwise differences
in **answerable accuracy** are significant (paired McNemar: Provem>Mem0
p = 2.5×10⁻⁴, Provem>Zep p = 1.1×10⁻³⁴, Mem0>Zep p = 3.1×10⁻¹⁵). Fixing Mem0's
input bug raised it from 0.509 to 0.565, so the Provem→Mem0 strict lead is
**+4.9 pts, not the +10.5 pts v1 reported**, still significant but roughly
half. (The judges disagree on the gap's size: opus scores the corrected Mem0 at
0.368, a wider lead, rejecting more of the extra borderline answers Mem0 now
attempts; strict-vs-opus κ for Mem0 is 0.54.) The abstention edge over Mem0
(0.863 vs 0.830) is a statistical tie (p = 0.12), and Provem's 0.863 comes from
its dev-tuned prompt chain: under the shared neutral prompt Provem's abstention
is 0.693, behind Zep 0.722 and Mem0 0.830 (answerable accuracy still wins under
that neutral prompt: 0.596 vs 0.565). Per category (strict judge): single-hop
**0.717 / 0.672 / 0.566**, temporal **0.614 / 0.523 / 0.290**, multi-hop
**0.411 / 0.394 / 0.330** (narrow), open-domain 0.312 / 0.271 / 0.312 (Provem/Zep
tie, n=96). Holdout conversations (final config chosen on the dev split
only) confirm the ordering (0.617 / 0.582 / 0.469). Full methodology, configs,
and limitations:
[`docs/three_system_benchmark.md`](docs/three_system_benchmark.md).

### The LoCoMo landscape (read before quoting any of it)

LoCoMo scores are **not comparable across papers**: they move ±30 points with
the judge prompt, retrieval depth, and answerer. Merging them into one
leaderboard is exactly the methodological sin the vendors accuse each other of,
so we present three separately-valid rankings instead.

**Ranking 1: measured in this repo, same harness (the only ranking we claim):**

| Rank | System | Strict judge | Claude judge | Mem0's own judge | Abstention |
|---|---|---|---|---|---|
| 1 | **Provem (dense tier)** | **0.614** | **0.502** | **0.772** | **0.863** |
| 2 | Mem0 platform | 0.565 | 0.368 | 0.716 | 0.830 |
| 3 | Zep platform | 0.449 | 0.304 | 0.649 | 0.722 |

Identical questions, answerer, and judges for every row. All pairwise
differences in answerable accuracy are significant (paired McNemar:
Provem>Mem0 p = 2.5×10⁻⁴, Provem>Zep p = 1×10⁻³⁴, Mem0>Zep p = 3×10⁻¹⁵) and the
ordering is identical under all three judges; the abstention column's
Provem-vs-Mem0 gap is a statistical tie and prompt-confounded (see above). Zep
was configured following its own published evaluation checklist (proper user
model, native created_at timestamps, parallel edge+node graph searches) to
pre-empt the misconfiguration critique it raised against Mem0's paper; config
details in the ship report. We rank only what we measured.

**Ranking 2: independent third-party evaluations (quoted verbatim, their
setups; two separate leaderboards, not comparable to each other):**

ENGRAM paper (arXiv 2511.12960)¹, k=20, gpt-4o-mini judge:

| System | Score |
|---|---|
| ENGRAM (academic system¹) | 77.6 |
| MemOS | 73.0 |
| **Mem0** | **64.7** |
| LangMem | 55.3 |
| OpenAI Memory | 52.8 |
| Zep | 42.3 |

LoCoMo-Refined (strict judge, 86% human agreement):

| System | Score |
|---|---|
| MemoraX AI | 82.7 |
| MemOS | 63.6 |
| MemPalace | 58.7 |
| EverMemOS | 58.3 |
| **Mem0** | **48.9** |

¹ Unrelated academic system (arXiv 2511.12960), no relation to this project.

**The anchor that connects the tables:** our strict-judge Mem0 measurement
(0.565, corrected) sits just above LoCoMo-Refined's strict Mem0 (48.9),
consistent with our fully-settled stores and single-date input; our
strict-judge Zep (0.449) lands next to the ENGRAM paper's independent Zep
(42.3), and far from Zep's self-reported 94.7. Under Mem0's own judge our
Mem0 lands at 0.716, inside its published band. Our harness reproduces what
independent evaluations find. MemOS and the remaining systems were not
measured head-to-head, so we make no claims against them.

**Ranking 3: vendor self-reports (marketing conditions, listed for completeness):**
Zep 94.7 (gpt-5.4 CoT reader/judge; after retracting an earlier 84% figure) ·
Mem0 92.5 / 91.6 (top-200 memories, gpt-5 CoT answerer, judge with partial
credit + 14-day date tolerance) · MemMachine 91.7 · Letta 74.0. Each was
produced by the vendor under conditions of its choosing; none is comparable to
any other number on this page.

Sources and the full dispute history (including who retracted what):
[`docs/ship_report.md`](docs/ship_report.md).

## What the governance layer does

- **Write-side:** prompt-injection quarantine, sensitive-without-consent hold,
  provenance + source-trust tagging, natural-language erasure/do-not-use intents.
- **Read-side:** erasure & do-not-use enforcement, tenant/entity scope isolation,
  source-conflict resolution by provenance trust, retention enforcement, and
  **calibrated abstention**: a recoverable "I don't know" instead of a confident
  wrong answer that compounds.
- **Accountability:** SHA-256 hash-chained tamper-evident audit log, GDPR erasure
  certificates, per-tenant compliance profiles (recruitment / pharma / finance /
  custom JSON-YAML), all decisions traceable.

## Honest limitations

- **One recall benchmark** (LoCoMo). LongMemEval port is designed, not run.
- **LLM judges only** (two vendors, κ 0.54 on Mem0 to 0.70 on Provem on the final
  artifacts); no human eval yet.
- Answer prompts were tuned on a dev split: the answerable-accuracy win
  survives with a fully neutral prompt (+3.1 pts vs the corrected Mem0 0.565)
  and on held-out conversations (+3.5 pts, though the holdout gap alone is not
  significant, p=0.074), but the abstention headline does not: under the neutral prompt
  Provem's abstention is 0.693, last of the three systems.
- Mem0 ran with platform defaults; a Mem0 expert might configure it better. Its
  stores kept consolidating between runs (drift favors Mem0).
- Keyless parity is not realistic and we don't claim it: with stdlib retrieval
  and an LLM answerer we measure 0.388; fully keyless (extractive answering)
  measures 0.207 to 0.212, below the 0.237 no-memory abstain baseline.
- The governance benchmark uses a scripted agent by design (causal isolation).
- **Governance defends untrusted-channel attacks, not same-channel ones.** A
  poison delivered through the same fully-trusted channel as the user (equal
  trust, later write) is served by both arms: provenance has no signal there;
  it needs write-side detection/review. Measured and reported in the
  attack-family benchmark (`--attack-families`), not hidden.
- No external security audit; not "production-certified"; self-hosted only.

The complete disclosure list, where Mem0 remains genuinely better (curated
human-readable memories, managed hosting), and the ship recommendation:
[`docs/ship_report.md`](docs/ship_report.md).

## Architecture

```text
Agent / LLM caller
      |
      v
GovernedMemory (backend-agnostic wrapper; MemoryBackend protocol)
      +-- write gate: injection quarantine, trust tagging, consent, retention
      +-- storage:    verbatim episodes + temporal facts (Naive / BM25 / SQLite / Mem0-adapter)
      +-- retrieval:  lexical BM25 + optional dense embeddings + RRF fusion
      +-- read gate:  erasure / scope / trust conflicts / relevance floor / abstention
      +-- audit:      SHA-256 hash-chained log + erasure certificates
      |
      v
MCP server (JSON-RPC/stdio, per-tenant profiles)   or   direct library embedding
```

## Documentation

| Doc | What it contains |
|---|---|
| [`docs/three_system_benchmark.md`](docs/three_system_benchmark.md) | The headline benchmark: Provem vs Mem0 vs Zep, three judges, paired stats |
| [`docs/ship_report.md`](docs/ship_report.md) | Final scoreboard, full limitations, ship recommendation |
| [`docs/lager_optimization_log.md`](docs/lager_optimization_log.md) | Every optimization iteration incl. failures and the bug post-mortem |
| [`docs/reliability_results.md`](docs/reliability_results.md) | Governance benchmark, full statistics |
| [`docs/mcp_server.md`](docs/mcp_server.md) | MCP product guide, profiles, config |
| [`docs/trust_model.md`](docs/trust_model.md) | Security boundaries; what belongs in a gateway |
| [`docs/claim_register.md`](docs/claim_register.md) | Every claim with evidence level and risk |
| [`docs/research_journal.md`](docs/research_journal.md) | Complete MVP history, every synthetic suite, every negative result |
| [`docs/runs/manifest.json`](docs/runs/manifest.json) + `scripts/replay_report.sh` | Bit-exact €0 reproduction of all benchmark numbers |

## Status

Research-grade core with enterprise-ready foundations: 490 tests, deterministic
quality gates, tamper-evident audit, tenant isolation, configurable compliance
profiles. **Not** externally security-audited, no managed hosting, no SLA: the
enterprise wrapper (gateway auth/SSO, hosting, certifications) is deliberately
out of scope for the core and documented in the trust model. Roadmap: LongMemEval
port, local-embeddings tier, cross-session fact rollups, judge-diverse human eval.


## Install

```bash
pip install provem            # dependency-free core, Python 3.9+
pip install "provem[mem0]"    # optional backends: mem0, zep, letta, graphiti, or all
```

## License

MIT, see [LICENSE](LICENSE). Use it, fork it, ship it commercially, no strings.
