Metadata-Version: 2.5
Name: cuimhin
Version: 0.1.0
Summary: Hybrid-retrieval memory layer for AI agents: Qdrant + SQLite FTS5 + RRF, with REST and MCP.
Project-URL: Repository, https://github.com/todd427/cuimhin
Project-URL: Changelog, https://github.com/todd427/cuimhin/blob/main/CHANGELOG.md
Project-URL: Issues, https://github.com/todd427/cuimhin/issues
Author: Todd McCaffrey
License: Apache-2.0
License-File: LICENSE
Keywords: agents,fts5,hybrid-search,mcp,memory,qdrant,rag
Classifier: Development Status :: 3 - Alpha
Classifier: License :: OSI Approved :: Apache Software License
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Software Development :: Libraries
Classifier: Topic :: Text Processing :: Indexing
Classifier: Typing :: Typed
Requires-Python: >=3.11
Requires-Dist: fastapi==0.141.1
Requires-Dist: foxxe-mcp==0.5.4
Requires-Dist: qdrant-client==1.19.0
Requires-Dist: typer==0.27.2
Requires-Dist: uvicorn[standard]==0.52.4
Provides-Extra: dev
Requires-Dist: build==1.6.1; extra == 'dev'
Requires-Dist: httpx==0.28.1; extra == 'dev'
Requires-Dist: mypy==2.3.1; extra == 'dev'
Requires-Dist: pytest==9.1.1; extra == 'dev'
Requires-Dist: ruff==0.16.6; extra == 'dev'
Requires-Dist: sentence-transformers==6.0.1; extra == 'dev'
Requires-Dist: twine==7.0.0; extra == 'dev'
Provides-Extra: sentence-transformers
Requires-Dist: sentence-transformers==6.0.1; extra == 'sentence-transformers'
Description-Content-Type: text/markdown

# <span style="color:#2E86AB">Cuimhin</span>

*Is cuimhin liom* — "I remember."

Cuimhin is a hybrid-retrieval memory layer for AI agents and applications. It stores text chunks with metadata, indexes them twice — dense vectors in Qdrant and keywords in SQLite FTS5 — and fuses the two rankings with Reciprocal Rank Fusion, so proper nouns and project names are found as reliably as paraphrases. It exposes the result as a Python library, a REST API, and an MCP server that plugs straight into Claude, Cursor, or any MCP-capable agent.

It is the open-source core extracted from **Mnemos**, a private memory system in daily production use since early 2026 at ~130k chunks. The retrieval design, the schema discipline and the failure modes documented here were all learned there — including the one that returned HTTP 200 for two days while the keyword leg was silently dead (`docs/SCOPE-SPLIT.md` §4.1).

## <span style="color:#2E86AB">Status</span>

0.1.0. The library, CLI, REST and MCP are here and tested; the cross-encoder reranker and the retrieval eval are 0.2.0. Each milestone's acceptance was run against a real Qdrant and written down as it happened — `docs/M0-RESULT.md` and `docs/M1-RESULT.md`, including the two runs that failed and what was wrong.

## <span style="color:#2E86AB">What it does</span>

- <span style="color:#F18F01">Ingest</span> — chunk text, embed it, write vectors to Qdrant and keywords to FTS5 in one idempotent call. Re-ingesting the same document is a no-op.
- <span style="color:#F18F01">Query</span> — hybrid search with metadata filters, date windows, and source scoping; returns ranked hits with score, source, date, and a reconstruction handle.
- <span style="color:#F18F01">Serve</span> — FastAPI REST (`/query`, `/ingest`, `/document/{id}`, `/export`) and FastMCP tools (`query_memory`, `ingest_document`, `get_document`, `list_sources`, `get_stats`) from one process.
- <span style="color:#F18F01">Export</span> — cursor-paged JSONL export so Cuimhin can be a source of record, not a dead end.

## <span style="color:#2E86AB">What it deliberately does not do</span>

Cuimhin is single-tenant and connector-free by design. Multi-tenant isolation, live connectors (email, drives, chat exports), dedup pipelines, temporal encoding, and enrichment layers are out of scope for this repo. See `docs/SCOPE-SPLIT.md` for the boundary and why it sits where it does.

## <span style="color:#2E86AB">Quick start</span>

```bash
pip install 'cuimhin[sentence-transformers]'    # the library and the local embedder
docker run -p 6333:6333 qdrant/qdrant           # or the Linux binary from Qdrant's releases
```

Two install sizes, and the difference is one dependency:

| | site-packages | what you get |
|---|--:|---|
| `pip install cuimhin` | 195 MB | library, CLI, REST, MCP — you supply an `Embedder` |
| `pip install 'cuimhin[sentence-transformers]'` | 1.5 GB | the above plus the default local model |

The extra brings torch (769 MB) and transformers (115 MB), and on a machine without CUDA a plain resolve pulls another ~2 GB of `nvidia-*` wheels it can never use — install the CPU build first if that matters:

```bash
pip install --index-url https://download.pytorch.org/whl/cpu torch
pip install 'cuimhin[sentence-transformers]'
```

If you embed through an API, or already have a model loaded, skip the extra: `Memory(embedder=...)` takes anything with eight methods (`cuimhin/embed.py`), and without it `Memory()` fails at startup with the command to run rather than at your first query. The model itself (`all-MiniLM-L6-v2`, ~90 MB) downloads once from the Hugging Face hub.

```bash
cuimhin ingest ./notes/                              # any folder of .md/.txt; re-ingest is a no-op
cuimhin query "what did I decide about the auth redesign"
cuimhin query "Ballyvaughan" --leg sparse            # ablate: keyword leg only
cuimhin export > notes.jsonl                         # the corpus, replayable, in (ingested_at, id) order
cuimhin import notes.jsonl                           # restore it exactly — ids, stamps, order; idempotent
```

Serve REST and MCP from one process behind one key:

```bash
export CUIMHIN_API_KEY=$(openssl rand -hex 32)
cuimhin serve                                        # REST on 127.0.0.1:8000, MCP at /mcp

curl -s localhost:8000/health                        # no key needed
curl -s -XPOST localhost:8000/query -H "X-API-Key: $CUIMHIN_API_KEY" \
     -H "Content-Type: application/json" -d '{"q": "auth redesign", "top_k": 5}'
```

`POST /query`, `POST /ingest`, `GET`/`DELETE /document/{id}`, `GET /export?after=&source=&limit=`, `GET /stats`, `GET /health`. Every error is `{"error": ..., "detail": ...}`; a drifted keyword index is a 500 that says `run cuimhin rebuild-fts`, never an empty 200. Rate limit is `CUIMHIN_RATE_LIMIT` per client per minute (default 60, `0` off); `/health` and `/export` are exempt.

### <span style="color:#F18F01">MCP</span>

Five tools: `query_memory`, `ingest_document`, `get_document`, `list_sources`, `get_stats`. Tools return data, not prose, and raise on error rather than returning nothing.

**Locally (Claude Code, Claude Desktop, Cursor):** stdio, no key — the client spawns the process under your own account.

```bash
claude mcp add cuimhin -e CUIMHIN_QDRANT_URL=http://localhost:6333 -- cuimhin serve --stdio
```

```json
{
  "mcpServers": {
    "cuimhin": {
      "command": "cuimhin",
      "args": ["serve", "--stdio"],
      "env": { "CUIMHIN_QDRANT_URL": "http://localhost:6333" }
    }
  }
}
```

**Hosted:** streamable HTTP at `https://your-host/mcp` with `Authorization: Bearer $CUIMHIN_API_KEY` — the same key as REST, the same rate limit. Set `CUIMHIN_ALLOWED_HOST=your-host` so the SDK's Host-header check accepts it (`localhost` and `127.0.0.1` are always accepted). OAuth is not in this layer.

### <span style="color:#F18F01">Python</span>

```python
from cuimhin import Memory

mem = Memory()                       # reads CUIMHIN_* env, or pass Config()
mem.ingest("The auth redesign ships in September.", source="notes", date="2026-08-25")
for hit in mem.query("auth redesign", top_k=5):
    print(hit.score, hit.source, hit.text[:80])

page = mem.export(limit=500)         # records, cursor, has_more — replay with ingest_chunk(ingested_at=...)
```

### <span style="color:#F18F01">Configuration</span>

All `CUIMHIN_*`; nothing else is read. `CUIMHIN_QDRANT_URL`, `CUIMHIN_QDRANT_API_KEY`, `CUIMHIN_COLLECTION`, `CUIMHIN_DB_PATH`, `CUIMHIN_EMBED_MODEL`, `CUIMHIN_CHUNK_TOKENS`, `CUIMHIN_CHUNK_OVERLAP`, `CUIMHIN_API_KEY`, `CUIMHIN_RATE_LIMIT`, `CUIMHIN_ALLOWED_HOST`, `CUIMHIN_RERANKER`, `CUIMHIN_LOG_LEVEL` — defaults and meanings in `docs/ARCHITECTURE.md` §6.

## <span style="color:#2E86AB">Why hybrid, and why RRF</span>

Dense retrieval alone misses the things people actually ask agents about: the name of a project, a person, a file, a variable. FTS5 catches those exactly. RRF merges the two rankings without a tunable weight, which means there is nothing to mis-tune when the corpus changes. A cross-encoder reranker is available as an opt-in hook; on the corpus this was developed against it regressed quality and cost ~90× latency on CPU, so it is off by default and stays that way until you measure otherwise.

## <span style="color:#2E86AB">Documents</span>

| File | What it is |
|---|---|
| `docs/PRD-cuimhin.md` | Product requirements: positioning, scope, API surface, non-goals |
| `docs/SCOPE-SPLIT.md` | Open / licensed / private boundary, and the extraction map from the parent system |
| `docs/ARCHITECTURE.md` | Module layout, the dual-index schema contract, data flow |
| `docs/BRIEF-m0-skeleton.md` | M0 build brief — the walking skeleton; `M0-ISSUES.md` and `M0-RESULT.md` are its findings and acceptance |
| `docs/BRIEF-m1-serve.md` | M1 build brief — REST, MCP, auth, export; `M1-ISSUES.md` and `M1-RESULT.md` likewise |
| `docs/BRIEF-m3-release.md` | M3 build brief — packaging and the 0.1.0 release |
| `docs/BRIEF-m2-eval.md` | M2 build brief — the reranker and the retrieval ablation, for 0.2.0 |
| `CHANGELOG.md` | What changed in each release, and what is deliberately absent |
| `CLAUDE.md` | Working rules for Claude Code in this repo |

## <span style="color:#2E86AB">Licence</span>

Apache-2.0. Copyright 2026 Todd McCaffrey.

Cuimhin is the shell of Mnemos, which stays private. The boundary and the reasoning behind it are in `docs/SCOPE-SPLIT.md` — including the four conditions a feature has to meet to belong here rather than there.
