Metadata-Version: 2.4
Name: moss
Version: 1.15.0
Summary: Python SDK for semantic search with on-device AI capabilities, including local-first session indexing
Author-email: "InferEdge Inc." <contact@usemoss.dev>
Project-URL: Homepage, https://github.com/usemoss/moss-samples
Project-URL: Repository, https://github.com/usemoss/moss-samples
Project-URL: Documentation, https://docs.usemoss.dev/
Keywords: search,semantic,embeddings,vector,usemoss,moss
Classifier: Development Status :: 5 - Production/Stable
Classifier: Intended Audience :: Developers
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3.14
Classifier: Topic :: Software Development :: Libraries :: Python Modules
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE.txt
Requires-Dist: typing-extensions>=4.16.0
Requires-Dist: inferedge-moss-core==0.27.0
Provides-Extra: dev
Requires-Dist: pytest>=9.1.1; extra == "dev"
Requires-Dist: pytest-asyncio>=1.4.0; extra == "dev"
Requires-Dist: tox>=4.56.2; extra == "dev"
Requires-Dist: black<26,>=25.9.0; extra == "dev"
Requires-Dist: isort<8,>=7.0.0; extra == "dev"
Requires-Dist: flake8>=7.3.0; extra == "dev"
Requires-Dist: mypy>=2.1.0; extra == "dev"
Requires-Dist: python-dotenv>=1.2.2; extra == "dev"
Requires-Dist: pdoc>=16.0.0; extra == "dev"
Requires-Dist: pydoc-markdown>=4.8.2; extra == "dev"
Requires-Dist: griffe>=2.1.0; extra == "dev"
Requires-Dist: pyyaml>=6.0; extra == "dev"
Dynamic: license-file

# Moss client library for Python

`moss` enables **private, on-device semantic search** in your Python applications with cloud storage capabilities.

Built for developers who want **instant, memory-efficient, privacy-first AI features** with seamless cloud integration.

## ✨ Features

- ⚡ **On-Device Vector Search** - Sub-millisecond retrieval with zero network latency
- 🔍 **Semantic, Keyword & Hybrid Search** - Embedding search blended with Keyword matching
- ☁️ **Cloud Storage Integration** - Automatic index synchronization with cloud storage
- 📦 **Multi-Index Support** - Manage multiple isolated search spaces
- 🛡️ **Privacy-First by Design** - Computation happens locally, only indexes sync to cloud
- 🚀 **High-Performance Rust Core** - Built on optimized Rust bindings for maximum speed
- 🧠 **Custom Embedding Overrides** - Provide your own document and query vectors when you need full control

## 📦 Installation

```bash
pip install moss
```

## 🚀 Quick Start

```python
import asyncio
from moss import MossClient, DocumentInfo, QueryOptions

async def main():
    # Initialize search client with project credentials
    client = MossClient("your-project-id", "your-project-key")

    # Prepare documents to index
    documents = [
        DocumentInfo(
            id="doc1",
            text="How do I track my order? You can track your order by logging into your account.",
            metadata={"category": "shipping"}
        ),
        DocumentInfo(
            id="doc2", 
            text="What is your return policy? We offer a 30-day return policy for most items.",
            metadata={"category": "returns"}
        ),
        DocumentInfo(
            id="doc3",
            text="How can I change my shipping address? Contact our customer service team.",
            metadata={"category": "support"}
        )
    ]

    # Create an index with documents (syncs to cloud)
    index_name = "faqs"
    await client.create_index(index_name, documents)  # Defaults to moss-minilm
    print("Index created and synced to cloud!")

    # Load the index (from cloud or local cache)
    await client.load_index(index_name)

    # Search the index
    result = await client.query(
        index_name,
        "How do I return a damaged product?",
        QueryOptions(top_k=3, alpha=0.6),
    )

    # Display results
    print(f"Query: {result.query}")
    for doc in result.docs:
        print(f"Score: {doc.score:.4f}")
        print(f"ID: {doc.id}")
        print(f"Text: {doc.text}")
        print("---")

asyncio.run(main())
```

### Scores and `min_score`

Each result's `score` is its relevance to the query from 0 to 1. Results are
ordered by hybrid relevance, so a lower result can show a higher score. Scores
depend on the embedding model: tune `min_score` on your own queries.

`QueryOptions(min_score=0.5)` drops results whose score is below 0.5; a query
returns up to `top_k` results that clear it, in hybrid order. It takes a value
from 0 to 1. Keyword-only queries (`alpha=0.0`) have no embedding to score
against: every result scores 0.0 and `min_score` raises an error.

### Candidate depth

A hybrid query ranks a fixed number of hits on each signal (vector and keyword)
and fuses those two lists, so a document outside both lists is never returned.
By default it looks `2 * top_k` deep. `QueryOptions(candidate_depth=200)` looks
deeper, which raises recall and query time. It takes a value of at least 1 and
is raised to `top_k` when below it. Keyword-only and vector-only queries
(`alpha` 0 or 1) rank one signal, where a deeper list returns the same results,
so they ignore it. `SessionIndex.query` ignores it too.

## 🔥 Example Use Cases

- Smart knowledge base search with cloud backup
- Realtime Voice AI agents with persistent indexes
- Personal note-taking search with sync across devices
- Private in-app AI features with cloud storage
- Local semantic search in edge devices, fully on-device

## Available Models

- `moss-minilm`: Lightweight model optimized for speed and efficiency
- `moss-mediumlm`: Balanced model offering higher accuracy with reasonable performance

## 🔧 Getting Started

### Prerequisites

- Python 3.8 or higher
- Valid InferEdge project credentials

### Environment Setup

1. **Install the package:**

```bash
pip install moss
```

2. **Get your credentials:**

Sign up at [InferEdge Platform](https://platform.inferedge.dev) to get your `project_id` and `project_key`.

3. **Set up environment variables (optional):**

```bash
export MOSS_PROJECT_ID="your-project-id"
export MOSS_PROJECT_KEY="your-project-key"
# Optional: override the manage API host (defaults to https://service.usemoss.dev)
export MOSS_CLOUD_API_BASE_URL="https://service.usemoss.dev"
```

### Basic Usage

```python
import asyncio
from moss import MossClient, DocumentInfo, QueryOptions

async def main():
    # Initialize client
    client = MossClient("your-project-id", "your-project-key")
    
    # Create and populate an index
    documents = [
        DocumentInfo(id="1", text="Python is a programming language"),
        DocumentInfo(id="2", text="Machine learning with Python is popular"),
    ]
    
    await client.create_index("my-docs", documents)
    await client.load_index("my-docs")
    
    # Search
    results = await client.query(
        "my-docs",
        "programming language",
        QueryOptions(alpha=1.0),
    )
    for doc in results.docs:
        print(f"{doc.id}: {doc.text} (score: {doc.score:.3f})")

asyncio.run(main())
```

### Hybrid Search Controls

`alpha` lets you decide how much weight to give semantic similarity versus keyword relevance when running `query()`:

```python
# Pure keyword search
await client.query("my-docs", "programming language", QueryOptions(alpha=0.0))

# Mixed results (default 0.8 => semantic heavy)
await client.query("my-docs", "programming language")

# Pure embedding search
await client.query("my-docs", "programming language", QueryOptions(alpha=1.0))
```

Pick any value between 0.0 and 1.0 to tune the blend for your use case.

### Disk cache

Pass `cache_path` to persist the downloaded index to disk. Later loads reuse the
cached copy while the cloud version is unchanged, so restarts skip the download.
Auto-refresh writes through to the same cache. Each load still contacts the cloud
to check for a newer version.

```python
await client.load_index(
    "my-docs",
    auto_refresh=True,
    polling_interval_in_seconds=300,
    cache_path="/var/cache/moss",
)
```

Set `cache_path` once on the client to make it the default for every load. A
per-call `cache_path` overrides it, and both override the native default
(`~/.moss`). The chosen directory also holds the `.moss-device-id` file that
keys Monthly Active Device billing.

```python
client = MossClient(
    "your-project-id",
    "your-project-key",
    cache_path="/var/cache/moss",
)

# Uses /var/cache/moss.
await client.load_index("my-docs", auto_refresh=True)

# Overrides it for this load only.
await client.load_index("other-docs", cache_path="/tmp/moss")
```

### Offline-tolerant load

With a `cache_path` set and a snapshot already on disk, a load whose cloud
metadata check fails on a transport error (unreachable network, timeout) serves
the cached snapshot instead of failing, so a restart during an outage still
answers queries. The probe uses a few-second budget, not the full retry window,
and the load starts a refresh poller so the index self-heals when the cloud
returns. A real "index not found" (404) still fails, so a deleted index does not
resurrect from cache.

A load served this way opens the index, but a built-in embedding model cannot be
opened until the Moss cloud is reachable, so text queries with `alpha` above 0
fail until then. Keyword-only queries (`alpha=0.0`) and queries that pass their
own `embedding` keep working, and a later refresh or query warms the model once
the cloud returns.

`was_served_stale(name)` reports whether an index is currently served from such a
snapshot; it clears once a refresh reconfirms it against the cloud.

```python
await client.load_index("my-docs", cache_path="/var/cache/moss")
if await client.was_served_stale("my-docs"):
    # Serving a cached copy; the cloud was unreachable at load.
    ...
```

### Multi-index search

Search several loaded indexes in one call and get the global top-K back, with
each result tagged by its source `index_name`. All indexes must be loaded
locally and share the same embedding model.

```python
loaded = await client.load_indexes(["products", "reviews"])
if not loaded.loaded:
    raise RuntimeError(f"no indexes loaded: {loaded.failed}")

results = await client.query_multi_index(
    loaded.loaded,
    "noise cancelling headphones",
    QueryOptions(top_k=5, alpha=0.5),
)
for doc in results.docs:
    print(f"[{doc.index_name}] {doc.id}: {doc.text} (score: {doc.score:.3f})")

await client.unload_indexes(loaded.loaded)
```

`alpha` works exactly as in `query()` (default 0.8): `1.0` is embedding-only,
`0.0` is keyword-only, and anything in between blends both with Reciprocal
Rank Fusion. Keyword scoring runs each index's own BM25 and merges each index's
raw hits before fusion; because BM25 statistics stay per-corpus, cross-index
keyword ranking is approximate. `top_k` caps the merged result, not each index,
and `filter` applies to every index. `load_indexes` is best-effort: names that
fail are reported in `failed` without rolling back the ones that loaded.

### Web sources

POST the existing `/v1/manage` web-source actions and return typed results.
Poll `job_id` with `get_job_status`. Depends on the index-manager release with
several sources per index and source-preserving crawls, and the moss-control
release with the deleteIndex proxy. Set `MOSS_CLOUD_API_BASE_URL` to send
manage calls to a non-default host (defaults to `https://service.usemoss.dev`).

```python
created = await client.create_web_source(
    "https://docs.example.com",
    "knowledge-base",
    max_pages=200,
    refresh_cadence="weekly",
)
await client.get_job_status(created.job_id)

sources = await client.list_web_sources(index_name="knowledge-base")
source = await client.get_web_source(created.id)
updated = await client.update_web_source(created.id, refresh_cadence="daily")
# updated.job_id is set when resync=True enqueued a crawl.
resync = await client.resync_web_source(created.id)
deleted = await client.delete_web_source(created.id)
# deleted.purge_job_id is set when a purge was enqueued.
```

### Metadata filtering

You can pass a metadata filter directly to `query()` after loading an index locally:

```python
results = await client.query(
    "my-docs",
    "running shoes",
    QueryOptions(top_k=5, alpha=0.6),
    filter={
        "$and": [
            {"field": "category", "condition": {"$eq": "shoes"}},
            {"field": "price", "condition": {"$lt": "100"}},
        ]
    },
)
```

For a complete runnable example, see `python/user-facing-sdk/samples/metadata_filtering.py`.

## 🧠 Providing custom embeddings

Already using your own embedding model? Supply vectors directly when managing
indexes and queries:

```python
import asyncio

from moss import DocumentInfo, MossClient, QueryOptions


def my_embedding_model(text: str) -> list[float]:
    """Placeholder for your custom embedding generator."""
    ...


async def main() -> None:
    client = MossClient("your-project-id", "your-project-key")

    documents = [
        DocumentInfo(
            id="doc-1",
            text="Attach a caller-provided embedding.",
            embedding=my_embedding_model("doc-1"),
        ),
        DocumentInfo(
            id="doc-2",
            text="Fallback to the built-in model when the field is omitted.",
            embedding=my_embedding_model("doc-2"),
        ),
    ]

    await client.create_index("custom-embeddings", documents)  # Defaults to moss-minilm
    await client.load_index("custom-embeddings")

    results = await client.query(
        "custom-embeddings",
        "<query text>",
        QueryOptions(embedding=my_embedding_model("<query text>"), top_k=10),
    )

    print(results.docs[0].id, results.docs[0].score)


asyncio.run(main())
```

Leaving the model argument undefined defaults to `moss-minilm`.
Pass `QueryOptions` to reuse your own embeddings or to override `top_k` on a per-query basis.

## Telemetry

The SDK reports usage to the Moss cloud at `$MOSS_INDEX_URL/telemetry`
(`https://service.usemoss.dev/index/telemetry` by default). Telemetry is always
on, because the device id and the query count drive usage billing. Requests are
sent in the background. A failed send is dropped and never fails the call that
caused it.

Queries do not send a request each. The SDK buffers their events, with the same
fields as below, and sends them every 3 seconds while queries run, plus once
more when the client or session is released. A send puts up to 500 query events in
one JSON array per request. The buffer holds up to 10,000 events between sends,
and events past that are dropped. Other events are sent one per request, when
the operation happens.

### Fields on every event

| Field | Value |
|---|---|
| `action` | Always `"telemetry"`. |
| `eventType` | The event type, from the table below. |
| `projectId` | Your project id. |
| `clientId`, `sessionId` | The same random id, one per client or session object. |
| `sdkVersion`, `sdkPackageVersion` | Version of the native binding package. |
| `nativeCoreVersion` | Version of the Rust core. |
| `modelCacheSchemaVersion` | Layout version of the on-disk model cache. |
| `deviceId`, `mossDeviceId` | The Moss device UUID stored in `.moss-device-id` under `cache_path` or `~/.moss`. `deviceId` is the billable one. |
| `indexName` | The index, when the event concerns one index. |
| `modelId`, `modelArtifactVersion`, `modelManifestSha256` | The embedding model and its exact artifact, when known. |

### Event types and their extra fields

| Event | Sent when | Extra fields |
|---|---|---|
| `index.load` | An index loads | `docCount`, `autoRefresh`, `refreshIntervalSecs` (with auto refresh), `cached` (a cache path was given), `servedStale` (served from the cache while the cloud is unreachable), `modelCacheGeneration` |
| `index.auto_refresh` | An auto refresh poll finishes | `changed`, `modelCacheGeneration` |
| `index.unload` | An index unloads or the client closes | `modelCacheGeneration` |
| `index.query` | A single-index query runs, sent in a batch | `latencyMs`, `modelCacheGeneration` |
| `index.query_multi` | A multi-index query runs, sent in a batch | The `index.query` fields, plus `indexCount` and `indexNames`. `modelCacheGeneration` when every index shares one, else `modelCacheGenerationsByIndex`. |
| `index.usage.periodic_flush`, `index.usage.final_flush` | Every 3 seconds while there is usage, and at close | `queryCount`, `docsEmbedded`, `tokensEstimated` |
| `session.load` | A session loads an index from the cloud | `docCount` |
| `session.load_from_disk` | A session loads a local snapshot | `docCount`, plus `rebuilt` and `persisted` when a legacy snapshot is rebuilt |
| `session.add_docs`, `session.delete_docs` | Documents are added or deleted | `docCount` |
| `session.push_index` | A session pushes its index | `docCount` |
| `session.auto_refresh_staged` | A newer cloud version is downloaded | `updatedAt` |
| `session.auto_refresh` | A downloaded version is installed | `changed`, `docCount`, `updatedAt` |
| `session.unload` | A session is released | None |
| `session.query` | A session query runs, sent in a batch | `latencyMs` |
| `session.usage.periodic_flush`, `session.usage.final_flush` | Every 3 seconds while there is usage, at release, and before a push | `queryCount`, `docsEmbedded`, `tokensEstimated` |

`latencyMs` on a query event is the time that query took, in milliseconds.
`index.query` and `index.query_multi` time the whole call, including the query
embedding. `session.query` times the search lookup only, without embedding, so
the two are not directly comparable.

No document text, query text, metadata or embedding is ever sent.

## 📄 License

This package is licensed under the [PolyForm Shield License 1.0.0](./LICENSE.txt).

- ✅ Free for any use, including production and commercial use.
- ❌ Not permitted: providing a product that competes with Moss.
- 📩 For commercial licenses, contact: <contact@usemoss.dev>

## 📬 Contact

For support, commercial licensing, or partnership inquiries, contact us: [contact@usemoss.dev](mailto:contact@usemoss.dev)
