Metadata-Version: 2.5
Name: goodmem-nlweb
Version: 0.3.0
Summary: GoodMem as an NLWeb retrieval provider — Schema.org retrieval, site filtering, and object lookup.
Project-URL: homepage, https://github.com/PAIR-Systems-Inc/goodmem_nlweb
Project-URL: source, https://github.com/PAIR-Systems-Inc/goodmem_nlweb
Project-URL: issues, https://github.com/PAIR-Systems-Inc/goodmem_nlweb/issues
Author-email: GoodMem <support@goodmem.ai>
License-File: LICENSE
Keywords: goodmem,nlweb,rag,retrieval,schema.org,vector-search
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Topic :: Software Development :: Libraries :: Python Modules
Requires-Python: >=3.10
Requires-Dist: goodmem>=0.1.34
Requires-Dist: nlweb-core>=0.7
Requires-Dist: typing-extensions>=4.7
Provides-Extra: dev
Requires-Dist: httpx>=0.27; extra == 'dev'
Requires-Dist: mypy==1.18.2; extra == 'dev'
Requires-Dist: pytest-asyncio>=0.23; extra == 'dev'
Requires-Dist: pytest>=7.0; extra == 'dev'
Requires-Dist: ruff==0.16.8; extra == 'dev'
Description-Content-Type: text/markdown

# goodmem-nlweb

GoodMem as an [NLWeb](https://github.com/nlweb-ai/NLWeb) retrieval provider.

```bash
pip install goodmem-nlweb
```

## Configure

NLWeb imports a provider by path, so this package never has to be added to
NLWeb's own source tree:

```yaml
retrieval:
  default:
    import_path: goodmem_nlweb
    class_name: GoodMemRetrievalProvider
    base_url: https://localhost:8080
    api_key: gm_…
    space_name: nlweb
```

NLWeb's `ask` handler asks for the retrieval provider named `default`, so
that is the name to give it. Options sit beside `import_path` and
`class_name`: NLWeb passes every other key of the entry to the constructor as
a keyword argument. A key ending in `_env` is read from the environment
instead, so `api_key_env: GOODMEM_API_KEY` keeps the key out of the file.

It implements `nlweb_core.retriever.RetrievalProvider`, so `search()` returns
real `RetrievedItem` objects and `close()` releases the client.

`space_id`, `embedder_id` and `reranker_id` must be GoodMem ids, which are
UUIDs; any other value raises `ValueError` when the provider is built, before
a request is made. (An empty `embedder_id` or `reranker_id` means "not set";
an empty `space_id` is refused.) The GoodMem SDK places ids in the URL path
unencoded, so a value like `../spaces/<id>` would otherwise send the request to
a different endpoint.

## Sites

NLWeb filters every call by *site*; GoodMem has no such concept. A site is
stored as memory metadata and filtered **server-side** with GoodMem's filter
grammar, the same way NLWeb's own Qdrant provider uses a `site` payload field.
Values are escaped, never interpolated — a site called `x' OR '1'='1` matches
nothing rather than everything.

```python
await provider.search("noodle soup", "recipes.example.com")   # one site
await provider.search("noodle soup", ["a.com", "b.com"])      # OR across sites
await provider.search("noodle soup", "all")                   # every site
await provider.search_all_sites("noodle soup")                # same thing
await provider.get_sites()                                    # ['a.com', 'b.com']
```

## Metadata filters

`search()` also takes `metadata_filter`, a mapping of metadata field to
value. Each pair becomes a server-side equality, AND-ed with the others and
with the site filter. GoodMem stores metadata as JSON and compares it under an
explicit cast, and a comparison under the wrong cast is accepted and matches
nothing, so each value is cast by its Python type:

| Value | Compared as | Sent to the server |
| --- | --- | --- |
| `str` | text | `CAST(val('$.cuisine') AS TEXT) = 'Malaysian'` |
| `bool` | boolean | `CAST(val('$.vegetarian') AS BOOLEAN) = true` |
| `int`, `float` | number | `CAST(val('$.servings') AS NUMERIC) = 4` |

`None`, `nan`, `inf` and any other type (a list, a dict, …) raise
`ValueError` before a request is made. Text is escaped exactly as a site is.

```python
await provider.search("noodle soup", "recipes.example.com",
                      metadata_filter={"vegetarian": True, "servings": 4})
```

A filter read from YAML (or JSON) keeps the type the loader gives it, so write
each value the way it is stored: `vegetarian: true` is a boolean, while
`vegetarian: "true"` is the string `"true"` and matches only a stored string.

```yaml
vegetarian: true
servings: 4
cuisine: Malaysian
```

`metadata_filter` is an argument to `search()`, not a provider option: NLWeb's
`ask` handler does not pass one, and a `metadata_filter:` key in the
provider's YAML entry is ignored.

## Ingesting Schema.org documents

`RetrievalProvider` only reads, so ingestion is a helper rather than part of
the interface:

```python
from goodmem_nlweb import GoodMemRetrievalProvider, upload_documents

provider = GoodMemRetrievalProvider(space_name="nlweb", base_url=…, api_key=…)
await upload_documents(provider, [
    {"@type": "Recipe", "url": "https://ex.com/r/laksa", "name": "Singapore Laksa",
     "description": "A coconut curry noodle soup."},
], site="recipes.example.com")
```

| NLWeb | GoodMem |
| --- | --- |
| `url` | `metadata["url"]` — the item's identity, and what NLWeb dedupes on |
| `site` | `metadata["site"]` — what every search filters on |
| `raw_schema_object` | `metadata["schema_json"]` |
| (text to embed) | the memory's content |

The embedded text is **not** the raw JSON. Embedding a JSON blob buries the
words a query would match under punctuation and key names, so the text is
extracted from `name`, `headline`, `description`, `articleBody` and friends,
and the untouched object is kept alongside for NLWeb to return verbatim.

`upload_documents` inspects every item of the batch response: the server
returns HTTP 200 even when an item failed, so per-item `success` is the only
signal. A partial failure raises `GoodMemUploadError` carrying the ids that
did land.

## Looking objects up by URL

```python
from goodmem_nlweb import GoodMemObjectLookupProvider

lookup = GoodMemObjectLookupProvider(space_name="nlweb", base_url=…, api_key=…)
await lookup.get_by_id("https://ex.com/r/laksa")   # the full Schema.org object
```

Implements `ObjectLookupProvider`, so NLWeb can enrich a truncated search
result with the complete object without a second datastore. NLWeb does that
when an `object_storage` provider named `default` is configured:

```yaml
object_storage:
  default:
    import_path: goodmem_nlweb
    class_name: GoodMemObjectLookupProvider
    base_url: https://localhost:8080
    api_key: gm_…
    space_name: nlweb
```

The id NLWeb passes here comes from a web request. It is the item's URL, not
a GoodMem id, so it is not required to be a UUID: it only ever travels as an
escaped value in the `filter` query parameter, never in the URL path.

## Scores, and why there are none

A GoodMem vector score is a negative inner product — the best match is the
*lowest* number — and a reranker score is a different scale that also goes
negative. `RetrievedItem` has no score field, and inventing one would imply a
comparability that does not hold. Results keep the server's ordering, which is
authoritative, and are never re-sorted here.

## Degraded retrieval

If the server reports a problem, whatever it did return is still returned and
the statuses are logged. If it reports a problem *and* returns nothing, the
result is an empty list plus a `UserWarning` — never an exception, because an
exception here would take down an `ask` request that could still answer from
another endpoint. Notices that carry no loss (`FEATURE_DISABLED`,
`LLM_CAPABILITY_INFERRED`) are ignored; a status code this version does not
know is reported as `UNKNOWN` rather than dropped.

## Which NLWeb?

This targets **`nlweb-core`** (the pip-installable package with the
config-driven provider architecture). The `nlweb-ai/NLWeb` reference
implementation has a different interface (`VectorDBClientInterface`, returning
`list[list[str]]`) and a hardcoded provider table, so a third-party package
cannot register with it — that one needs an upstream PR.

## Development

```bash
pip install -e ".[dev]"
ruff check src tests examples && mypy && pytest -m "not integration"
```

The offline suite replays NDJSON captured from a live GoodMem server
(v1.0.320) through the real SDK decoders. The live suite needs a server:

```bash
GOODMEM_BASE_URL=… GOODMEM_API_KEY=… GOODMEM_EMBEDDER_ID=… \
  GOODMEM_RERANKER_ID=… GOODMEM_VERIFY_SSL=0 \
  pytest -m integration
```

`GOODMEM_RERANKER_ID` is optional — the reranker test skips without it.
`GOODMEM_VERIFY_SSL=0` is for a local server with a self-signed certificate.

There is no default credential anywhere in this repository.

## License

MIT
