Metadata-Version: 2.4
Name: slim-llm-memory
Version: 0.1.0
Summary: Slim, fast, persistent memory + retrieval for LLM apps. numpy + httpx, ~1000 LOC, drops into anything.
Author: trbck
License-Expression: MIT
Project-URL: Homepage, https://github.com/trbck/slim-llm-memory
Project-URL: Documentation, https://github.com/trbck/slim-llm-memory/blob/main/docs/IMPLEMENTATION.md
Project-URL: Changelog, https://github.com/trbck/slim-llm-memory/blob/main/CHANGELOG.md
Project-URL: Issues, https://github.com/trbck/slim-llm-memory/issues
Keywords: llm,embeddings,vector,memory,rag,ollama
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Typing :: Typed
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: numpy>=1.24
Requires-Dist: httpx>=0.25
Provides-Extra: gemini
Requires-Dist: google-genai>=1.0; extra == "gemini"
Provides-Extra: graph
Requires-Dist: networkx>=3.0; extra == "graph"
Provides-Extra: rerank
Requires-Dist: sentence-transformers>=3.0; extra == "rerank"
Provides-Extra: anthropic
Requires-Dist: anthropic>=0.40; extra == "anthropic"
Provides-Extra: obsidian
Requires-Dist: watchdog>=4; extra == "obsidian"
Requires-Dist: pyyaml>=6; extra == "obsidian"
Requires-Dist: mcp>=2; extra == "obsidian"
Provides-Extra: test
Requires-Dist: pytest>=7; extra == "test"
Requires-Dist: pytest-asyncio>=0.23; extra == "test"
Requires-Dist: tomli>=2; python_version < "3.11" and extra == "test"
Provides-Extra: dev
Requires-Dist: slim-llm-memory[anthropic,gemini,graph,obsidian,rerank,test]; extra == "dev"
Requires-Dist: ruff>=0.6; extra == "dev"
Requires-Dist: build>=1; extra == "dev"
Requires-Dist: twine>=5; extra == "dev"
Dynamic: license-file

# slim-llm-memory

Slim, fast, persistent **memory + retrieval for LLM apps**. Pure Python
where possible; numpy where it actually helps. Ollama for local
embeddings; cloud LLMs only for hard reasoning. ~1000 LOC, two hard
deps (numpy + httpx), drops into anything.

> **Status:** phase 1 — `Memory` core. See [`docs/IMPLEMENTATION.md`](docs/IMPLEMENTATION.md)
> for the full plan and what comes next (Gemini fallback, Tier router,
> Graph layer, ANN swap).

## Why

Vector DBs and full RAG frameworks are overkill for personal projects
and research code. At < 50k items, a single numpy array, a jsonl file,
and a content-hash for incremental updates is all you actually need.
This library is exactly that — but written carefully enough that you
can build serious things on it without hitting sharp corners.

When you outgrow it, the public API is **swap-compatible** with a real
vector store (faiss / SQLite-vss / Qdrant). The migration is local to
one file.

## Install

```bash
pip install slim-llm-memory           # numpy + httpx only
pip install slim-llm-memory[graph]    # + NetworkX graph layer
pip install slim-llm-memory[rerank]   # + sentence-transformers cross-encoder
```

Working on the library itself:

```bash
pip install -e .                      # then `import slim_llm_memory` works anywhere
```

Install it even for local hacking. Running from the repo root with `PYTHONPATH=.`
works, but it hides packaging bugs — a broken console-script entry survived exactly
that way until the package was first installed for real.

The `[gemini]` and `[anthropic]` extras are declared but **not yet implemented**:
no module imports them. `Embedder` currently offers `noop` and `ollama`, and the
answer path talks only to Ollama.

## 30-second tour

```python
from slim_llm_memory import Memory, Embedder

# Local Ollama for embeddings, persistent index in ./mymemory/
mem = Memory("./mymemory", Embedder.ollama("nomic-embed-text"))

# Add or update items — only changed texts are re-embedded
mem.upsert([
    {"id": "doc1", "text": "how to set up nginx", "meta": {"kind": "note"}},
    {"id": "doc2", "text": "milch kaufen",        "meta": {"kind": "shopping"}},
])

# Top-k semantic search — optional filters
hits = mem.search("nginx tutorial", k=5, kinds={"note"}, min_score=0.55)
for h in hits:
    print(h.id, h.score, h.text)

# Find duplicates by cosine similarity
clusters = mem.find_duplicates(threshold=0.86)

# Atomic persistence — safe to crash mid-anything
mem.flush()
```

`Embedder.noop()` exists for tests and offline development — same
interface, deterministic SHA-256 derived vectors, no network.

## What's in the box (phase 1)

| Module        | Purpose                                                      |
|---------------|--------------------------------------------------------------|
| `index.py`    | `Memory`, `Hit` — public API                                 |
| `store.py`    | Versioned manifest + atomic flush + fcntl lock + tombstones  |
| `embed.py`    | `Embedder.noop` (tests) + `Embedder.ollama` (local)          |
| `obs.py`      | Per-instance ring buffers + counters for `Memory.stats()`    |

Public API surface (the only thing callers see):

```
Memory(path, embedder)
  .upsert(items)              → {added, updated, skipped, embed_calls}
  .search(query, k, kinds, min_score)  → [Hit, ...]
  .neighbours(id, k, kinds)   → [Hit, ...]  (no embed call)
  .search_vector(vec, k, kinds, min_score) → [Hit, ...]  (pre-embedded query)
  .find_duplicates(threshold) → [[id, ...], ...]
  .update_text(id, text)      → bool
  .remove(id)                 → bool
  .stats()                    → dict (JSON-safe)
  .flush(force=False)         → bool
  .close(flush=True)
  context manager: `with Memory(...) as mem: ...`

Embedder.noop(dim=384)
Embedder.ollama(model="nomic-embed-text", base_url="http://localhost:11434", timeout=60)
```

## Persistence model

Files in your index directory:

```
items.vN.jsonl       one record per item: {id, text, hash, meta, ts, deleted?}
vectors.vN.npy       float32 ndarray, shape (N, dim) — row-aligned with items
manifest.json        atomic commit point; loading always honours its version pointer
.lock                advisory exclusive lock (one writer per directory)
```

A crash mid-flush leaves the **previous manifest version intact** — the
old files load cleanly. Garbage versioned files left behind by
crashes are ignored on next load.

## Performance

At p95 on a CPU with prenormalised float32 vectors:

| Items  | Pure-Python cosine | numpy linear scan (this lib) | faiss HNSW (phase 7) |
|--------|--------------------|------------------------------|----------------------|
| 1k     | 5–20 ms            | <1 ms                        | <1 ms                |
| 10k    | 50–200 ms          | 5 ms                         | <1 ms                |
| 50k    | 0.5–2 s            | 30 ms                        | 1–10 ms              |
| 100k+  | dead               | 100–500 ms                   | 1–10 ms              |

Phase 1 ships the numpy linear scan. When you outgrow it, swap the
storage backend behind the same `Memory.search()` signature.

## Topic store: fast context for an LLM working on one topic

`topic()` is the "one numpy store per topic" shape with a `requests`-style
front door: open a store, put text in, get context out.

```python
from slim_llm_memory import topic

t = topic("nginx")                          # ~/.slim-llm-memory/topics/nginx, Ollama nomic-embed-text
t.add("docs/")                              # file, directory, raw text, or {name: text}; saved on return
r = t.ask("how do I enable TLS?")           # one embed call + one numpy scan
r                                           # hits with scores, embed ms, scan ms
r.context                                   # numbered block to prepend to an LLM prompt
t.answer("how do I enable TLS?")            # + a local Ollama chat model, grounded on r.context
```

`t.add` is incremental (unchanged chunks are never re-embedded), `t.forget(name)`
drops a doc, `embedder="noop"` runs offline for tests.

Several topics make a database. `library()` is a folder of topic stores;
`ask` embeds once and scans every topic, archiving is a folder move:

```python
from slim_llm_memory import library

db = library()                              # ~/.slim-llm-memory/topics
db.topic("nginx").add("docs/nginx/")
db.topic("cooking").add({"pasta.md": "..."})
db                                          # table of topics
db.ask("how do I enable TLS?")              # hits labelled by topic, merged by score
db.ask("...", topics=["nginx"])
db.route("how do I enable TLS?")             # stage 1 alone: topics ranked by centroid similarity
db.ask("...", route=True)                    # two-stage: route, then scan only the chosen topics
db.archive("cooking"); db.restore("cooking"); db.delete("cooking")
```

`ask` is exact (one concatenated scan) until the library holds more than
50k chunks, then it routes through topic centroids automatically; topics
within 0.05 of the best centroid are kept, and a prompt that matches no
topic falls back to the exact scan. `examples/03_routing_bench.py` has the
numbers: at 500 topics × 200 chunks, routing cuts the scan from ~40 ms to ~2 ms.

### Accuracy: hybrid retrieval, reranking, evaluation

```python
t.ask(q)                                    # hybrid (default): dense cosine ∪ BM25, fused by normalised score
t.ask(q, mode="dense") / t.ask(q, mode="keyword")
t.ask(q, rerank=True)                       # cross-encoder over the top 4·k  (pip install slim-llm-memory[rerank])
t.ask(q, rerank="auto")                     # ...but only when the top of the ranking is actually contested
t.ask(q, rerank=rr, rerank_margin=0.15)     # same policy with your own reranker; r.rerank_skipped says what happened
t.answer(q, rewrite=True, refuse_below=0.4, stream=False)   # query rewrite, refusal, validated [n] citations

from slim_llm_memory import evaluate
evaluate(t, [("which file is the commit point?", "manifest"), ...], k=5)   # hit@1, hit@k, MRR
```

On the eight doc questions in `notebooks/accuracy_demo.ipynb` (four of them
with the product name in the question, which drags the intro chunks up),
measured over this repo's own docs:

| retrieval | hit@1 | hit@5 | MRR |
|---|---|---|---|
| dense | 0.38 | 0.62 | 0.47 |
| hybrid (default) | 0.38 | 0.88 | 0.56 |
| hybrid + cross-encoder rerank | 0.62 | 1.00 | 0.76 |

Chunks are heading-aware with a 20-word overlap (`topic(..., chunk_words=120,
overlap=20)`); a tuning grid over chunk size, overlap and the fusion weight is
in the notebook and confirms the defaults. Re-tune per corpus with `evaluate()`.

Reranking is the most accurate and by far the slowest step, so `rerank="auto"`
pays for it only when the top of the ranking is contested: it compares the
leader's lead over the runner-up against the pool's spread, and skips the model
when that relative gap is at least `rerank_margin` (default 0.15).
`examples/04_rerank_bench.py` measures the trade on a 14-document corpus and 10
questions — with the real embedder and `bge-reranker-v2-m3` on this CPU box:

| policy | MRR | hit@1 | reranker calls | ms/query |
|---|---|---|---|---|
| off | 1.00 | 1.00 | 0 | 447 |
| auto | 1.00 | 1.00 | 0 of 10 | 749 |
| always | 1.00 | 1.00 | 10 | 3622 |

Same answers, 4.8× faster than reranking everything. On a harder corpus (the
`--offline` run, where dense retrieval alone gets one question wrong) auto
reranks 3 of 10 questions and recovers the full MRR that `always` reaches.
`r.rerank_skipped` reports the decision per query.

### Structure: graph, entities, sessions

```python
t.link("nginx.md", "certbot.md", relation="uses")    # typed edges, graph.json next to the vectors
t.related("nginx.md")                                # 0.6·cosine + 0.4·graph; [[wikilinks]] become edges on add
t.add(text, enrich=True)                             # local LLM extracts entities + relations (slow, opt-in)
t.entities(); t.ask(q, entity="Postgres")            # filter by extracted entity

s = db.session("2026-09-04")                         # conversation memory as a topic store
s.turn("user", "..."); s.recall("what did we decide?"); s.history(5); s.summary(model=...)
```

`notebooks/library_demo.ipynb` walks through it. `notebooks/use_cases_demo.ipynb`
measures four real use cases (grounded answers, paraphrase, languages, agent
session memory) and ends with an honest table of what is missing compared to a
full RAG stack, an ontology, and a vector database.

`notebooks/topic_context_demo.ipynb` (executed, 14 cells) and
`examples/02_topic_context.py` are the proof: this repo's docs as the
topic, live prompts with the latency split into embed vs scan, an
incremental update, an optional grounded LLM answer, and a synthetic scale
run. Measured on an 8-core CPU box (Ollama CPU-only):

| Step | Cost | Where the time goes |
|------|------|---------------------|
| Prompt → context (33 chunks) | 1.2–1.5 s | Ollama embed of the prompt: >99.9 %. Scan: 0.2–0.5 ms |
| Re-index after one edit | 1 embed call | 32 chunks hash-skipped, 1 re-embedded |
| Scan, 1k × 768 | 0.3 ms p50 / 1.3 ms p95 | numpy GEMV + argpartition |
| Scan, 10k × 768 | 2.4 ms p50 / 7.3 ms p95 | |
| Scan, 50k × 768 | 12 ms p50 / 22 ms p95 | |

The retrieval itself is never the bottleneck at this scale; the embedder is.
On a GPU or with a cloud embedder the prompt-to-context time drops to tens
of milliseconds and the scan numbers above are what remains.

```bash
PYTHONPATH=. python examples/02_topic_context.py --fresh                  # cold build + queries + scale
PYTHONPATH=. python examples/02_topic_context.py --llm llama3.2:3b        # + grounded answer
```

## Tests + examples

```bash
pytest                              # 169 tests, no network (Embedder.noop)
python examples/01_minimal.py       # after `pip install -e .`
```

### Notebooks

Start with the four `hello` notebooks — each is about ten lines and answers one
question. They need Ollama running with `nomic-embed-text` pulled.

| Notebook | Shows |
|---|---|
| `notebooks/00_hello_topic.ipynb`   | the three verbs: `topic()` / `.add()` / `.ask()` |
| `notebooks/01_hello_library.ipynb` | many topics behind one handle, and `route()` |
| `notebooks/02_hello_memory.ipynb`  | the low-level `Memory` API this tour uses |
| `notebooks/03_hello_answer.ipynb`  | a grounded answer with citations, and refusal |

The longer notebooks (`topic_context_demo`, `library_demo`, `accuracy_demo`,
`use_cases_demo`) go deeper on measurement.

## Migration paths

When the slim stack stops being enough, swap one file:

| Symptom                              | Replace                                    |
|--------------------------------------|--------------------------------------------|
| Search p95 > 100 ms at your scale    | `index.py` → faiss-cpu HNSW (same API)     |
| Need a 2nd writer process            | `store.py` → SQLite + sqlite-vss extension |
| > 1M items                           | both → Qdrant / Weaviate as a service      |
| Need real multi-hop graph queries    | future `graph.py` → Kùzu (embedded)        |
| Local LLM too slow / quality too low | future `tier.py` → drop L2; route L0 → L3  |

The whole point is: you don't outgrow it gradually. When you do, the
symptoms are obvious and the migration is local.

## License

MIT.
