Metadata-Version: 2.5
Name: goodmem-wandb
Version: 0.2.2
Summary: GoodMem retrieval for Weights & Biases Weave — traced ops, an evaluable weave.Model, and retrieval scorers.
Project-URL: Homepage, https://goodmem.ai/
Project-URL: Source, https://github.com/PAIR-Systems-Inc/goodmem-wandb
Project-URL: Issues, https://github.com/PAIR-Systems-Inc/goodmem-wandb/issues
Author: PairSys
License: MIT
License-File: LICENSE
Keywords: evaluation,goodmem,memory,rag,retrieval,tracing,wandb,weave
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: MIT License
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Topic :: Software Development :: Libraries :: Python Modules
Requires-Python: >=3.10
Requires-Dist: goodmem>=0.1.34
Requires-Dist: pydantic>=2.7
Requires-Dist: weave>=0.51
Provides-Extra: dev
Requires-Dist: httpx>=0.27; extra == 'dev'
Requires-Dist: mypy==1.18.2; extra == 'dev'
Requires-Dist: pytest>=7.0; extra == 'dev'
Requires-Dist: ruff==0.16.8; extra == 'dev'
Description-Content-Type: text/markdown

# goodmem-wandb

GoodMem retrieval for [Weights & Biases Weave](https://weave-docs.wandb.ai/).

Every retrieval is a traced Weave op, so a RAG call shows up in the Weave UI
with its query, its latency, each hit's score **and what kind of score that
is**, and any degradation the server reported. The same retriever can be
wrapped as a `weave.Model` and measured with `weave.Evaluation`.

```bash
pip install goodmem-wandb
```

## Trace a retrieval

```python
import weave
from goodmem_wandb import GoodMemRetriever

weave.init("my-project")

retriever = GoodMemRetriever(space_name="docs")   # credentials from the environment
result = retriever.search("how do I rotate an API key?")

for hit in result["hits"]:
    print(hit["score"], hit["score_kind"], hit["chunk_text"][:80])
```

`search()` returns:

| Key | What it is |
| --- | --- |
| `query` | The query as passed in |
| `hits` | `chunk_id`, `chunk_text`, `memory_id`, `space_id`, `source`, `score`, `score_kind`, `metadata` — in the server's order |
| `score_kind` | `"vector"` or `"reranker"`: what the server actually returned, not what was configured. They are different scales; see below |
| `statuses` | Server statuses that indicate a real problem, `[]` when clean |
| `partial` | `True` when the server reported a real problem during this retrieval — with or without hits. An empty `hits` with `partial=True` is a failed search, not a miss |
| `abstract_reply` | The server-generated summary, only when `llm_id` is set |
| `space_ids` | Which spaces were actually searched |

Credentials come from `GOODMEM_BASE_URL` and `GOODMEM_API_KEY`, or as
constructor keywords:

```python
GoodMemRetriever(space_name="docs", base_url="https://goodmem.example.com",
                 api_key="gm_…")
```

`verify_ssl` defaults to on and exists for a local server with a self-signed
certificate only; no example here turns it off.

They are deliberately **not** Weave fields. Weave publishes an object's
pydantic fields verbatim to the trace server and its redaction helper does not
run on that path, so a field named `api_key` on a `weave.Model` would be
uploaded to W&B in plaintext. This package keeps the connection on a private
attribute; `tests/test_regressions.py` asserts it.

## Evaluate a retrieval configuration

```python
import asyncio

import weave
from goodmem_wandb import GoodMemRetrievalModel, RecallAtK, MRR, FactRecall, RetrievalHealth

weave.init("my-project")

dataset = [
    {"question": "how do I rotate an API key?",
     "expected_memory_ids": ["01a0…"],
     "expected_text": "rotate the key from the console"},
]

baseline = GoodMemRetrievalModel(space_name="docs", limit=5)
reranked = GoodMemRetrievalModel(space_name="docs", limit=5, reranker_id="…")

evaluation = weave.Evaluation(
    dataset=dataset,
    scorers=[RecallAtK(k=5), MRR(), FactRecall(), RetrievalHealth()],
)
asyncio.run(evaluation.evaluate(baseline))
asyncio.run(evaluation.evaluate(reranked))   # compare the two in the Weave UI
```

`Evaluation.evaluate` is a coroutine: called without `asyncio.run` (or
`await` in a notebook) it returns without running anything.

Changing any field on the model versions it, so the two runs are directly
comparable. The credentials are not fields, so they are not part of the
version either.

| Scorer | Measures |
| --- | --- |
| `RecallAtK(k=5)` | Fraction of `expected_memory_ids` in the top *k* **distinct memories** (several chunks of one memory are one document) |
| `MRR()` | Reciprocal rank of the first expected memory; `0.0` if none was retrieved |
| `FactRecall()` | Whether a known fact actually appears in the retrieved text — survives re-chunking and re-embedding, unlike an id metric |
| `RetrievalHealth()` | Whether the retrieval was complete, so a silently-degraded run is visible as its own metric rather than only as a recall drop |

None of them score on the raw relevance number, because that number does not
mean the same thing between two configurations.

## Scores

GoodMem returns two different things in the same field, and this matters:

| | Range observed on a live server | Best match is |
| --- | --- | --- |
| Vector score | negative, e.g. `-0.6154 … -0.3873` | the **lowest** number |
| Reranker score | `0.2105 … -0.1081` — also goes negative | the **highest** number |

Those are real numbers from one capture over the same three memories. So:

* results keep **the server's order** and are never re-sorted here;
* `score_kind` on every hit says which scale you are looking at;
* `min_score` is only applied to reranker scores, and is applied
  client-side where you can see it, never sent as the server's
  `relevance_threshold`.

Even with a reranker the scale is **model-dependent**: on the same documents
Voyage `rerank-2.5` scored `0.27..0.93` and Jina `jina-reranker-v3` scored
`-0.14..0.43`. A `min_score` tuned for one empties the other, so when a
threshold removes every hit the retriever warns and names the observed range.
Calibrate `min_score` for the reranker you use; there is no default.

`score_kind` says what the server did, not what was configured. When
`reranker_id` is set but the reranker fails, the server reports
`RERANKING_FAILED` (and `NOT_FOUND` for a missing reranker) and still returns
the vector-stage hits. Those hits come back as `score_kind: "vector"` with
their vector scores, `min_score` is not applied to them — a reranker
threshold on vector scores would discard every hit the server returned — and
the result is `partial: True` with both codes in `statuses`. Measured on a
live server (v1.0.320) with a missing reranker: the three hits scored
`-0.785, -0.577, -0.111` and all three are returned.

## Filtering

```python
retriever = GoodMemRetriever(
    space_name="docs", metadata_filter={"category": "billing", "archived": False}
)
retriever.search("refunds", metadata_filter={"lang": "en"})   # AND-ed per call
```

Each value is compared as its own type. GoodMem compares a cast of the stored
JSON value, and the cast has to match the type:

| Value | Sent as |
| --- | --- |
| `str` | `CAST(val('$.field') AS TEXT) = '…'`, quoted for the filter grammar (backslash escaping, verified against a live server; control characters are refused rather than mangled) |
| `bool` | `CAST(val('$.field') AS BOOLEAN) = true` or `= false` |
| `int`, `float` | `CAST(val('$.field') AS NUMERIC) = 2026`, finite numbers only, written as a plain decimal |
| `None`, anything else | Refused with `ValueError` before any request is sent |

A mismatched cast is not rejected by the server, it just matches nothing. On a
live server (v1.0.320) a memory with metadata `{"flag": true, "n": 5}` is
found by `CAST(val('$.flag') AS BOOLEAN) = true` and by
`CAST(val('$.n') AS NUMERIC) = 5.0`, but `CAST(val('$.flag') AS TEXT) = 'True'`
and `CAST(val('$.n') AS TEXT) = '5.0'` both return nothing with HTTP 200. A
string stays text: `{"year": "2026"}` compares the text `'2026'`.

For anything more complex, pass an expression directly:

```python
GoodMemRetriever(space_name="docs",
                 filter="CAST(val('$.year') AS NUMERIC) >= 2026")
```

## Attaching to a space by name

```python
GoodMemRetriever(space_name="docs", embedder_id="…")                    # reuse or fail
GoodMemRetriever(space_name="docs", embedder_id="…", create_space=True) # or create it
```

Attach-by-name is idempotent reuse: an existing space whose embedder matches is
reused; one built on a **different** embedder is an error, because retrieving
across mismatched embedders returns plausible-looking nonsense. A name that
matches more than one space is also an error — GoodMem does not require space
names to be unique. This matches the ActivePieces connector.

## Sensitive corpora

`trace_chunk_text=False` keeps chunk ids, scores and statuses in the trace but
leaves the retrieved text out of it.

## Development

```bash
pip install -e ".[dev]"
ruff check src tests examples && mypy && pytest -m "not integration"
```

The offline suite replays NDJSON captured from a live GoodMem server
(v1.0.320) through the real SDK decoders, so the wire format is never
invented. The live suite needs a server and is skipped without one:

```bash
GOODMEM_BASE_URL=… GOODMEM_API_KEY=… GOODMEM_EMBEDDER_ID=… \
  GOODMEM_RERANKER_ID=… GOODMEM_VERIFY_SSL=0 \
  pytest -m integration
```

`GOODMEM_RERANKER_ID` is optional — the one reranker test skips without it.
`GOODMEM_VERIFY_SSL=0` is for a local server with a self-signed certificate.

There is no default credential anywhere in this repository.

## License

MIT
