Metadata-Version: 2.4
Name: navia-sdk
Version: 2.0.0
Summary: Official Python SDK for the Navia platform.
Project-URL: Homepage, https://moxoff.com
Author: Moxoff
License: Proprietary
Keywords: chat,connector,data-extraction,mcp,navia,ocr,sdk
Classifier: Intended Audience :: Developers
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Requires-Python: >=3.10
Requires-Dist: httpx<1.0,>=0.27
Description-Content-Type: text/markdown

# navia-sdk

Official Python SDK for the [Navia](https://moxoff.com) platform.

The SDK lets external services talk to an Navia backend over HTTP. A
single class — `Navia` — wraps every endpoint the backend publishes
under `/sdk/*`:

| Capability          | Token permission | What it does                                   |
| ------------------- | ---------------- | ---------------------------------------------- |
| External connectors | `connector`      | Register a connector, run sync cycles, push files, query checksums. |
| Artifacts           | `mcp`            | Publish generated files for delivery in chat.  |
| OCR                 | `ocr`            | Extract Markdown/JSON/PDF from a document.      |
| Data extraction     | `data-extraction`| Extract structured data from a document against a saved template or inline schema. |
| Chat                | `chat`           | OpenAI-style chat completions *(not implemented yet)*. |

A token's permissions decide which methods are callable; the backend rejects
calls made with a token that lacks the required permission. The chat methods
raise `NotImplementedError` until the backend implementation lands.

## Installation

```bash
pip install navia-sdk
```

Imports use the top-level `navia` package:

```python
from navia import Navia
```

## Authentication

All calls authenticate with an API token sent through the `X-Api-Key` HTTP
header. Tokens are created from the Navia admin panel and carry one or
more permissions (`connector`, `mcp`, `ocr`, `chat`).

The client reads two environment variables by default (or accepts them as
constructor arguments):

| Variable               | Purpose                             |
| ---------------------- | ----------------------------------- |
| `NAVIA_API_URL`   | Base URL of the Navia backend. |
| `NAVIA_API_TOKEN` | API token (`omni_...`).             |

```python
from navia import Navia

# Reads NAVIA_API_URL / NAVIA_API_TOKEN from the environment.
with Navia() as client:
    print(client.connector_id)        # connector bound to this token, if any
    print(client.last_sync_started_at)  # checkpoint of the last sync run

# …or pass them explicitly:
client = Navia(api_url="https://navia.example.com", api_token="omni_…")
```

When the token is bound to a connector, the client resolves the connector ID
on construction. Pass `resolve_connector=False` to skip that lookup when you
only use the MCP or OCR features.

## External connectors

A connector pushes documents from an external source into Navia. The
`run_sync` helper drives a full cycle — optional auto-register → notify start
→ your sync function → notify end (with stale-item cleanup):

```python
import os
from pathlib import Path

from navia import Navia


def sync(client: Navia) -> list[str]:
    source_dir = Path(os.environ["NAVIA_SOURCE_DIR"])
    existing = set(client.get_existing_checksums())
    active: list[str] = []

    for path in sorted(source_dir.rglob("*")):
        if not path.is_file():
            continue
        checksum = Navia.compute_checksum(path)
        active.append(checksum)
        if checksum not in existing:
            client.push_file(path, source_id=str(path.relative_to(source_dir)))

    # Returning the active checksums lets the backend delete stale items.
    return active


if __name__ == "__main__":
    with Navia() as client:
        client.run_sync(sync, name="local-files", description="Local files")
```

`item_exists(source_id=...)` / `item_exists(checksum=...)` query the backend
for a single item without fetching the whole checksum list. See
[`examples/external_connector/`](examples/external_connector/) for a runnable
connector and Dockerfile.

## Artifacts

Publish a file produced by an MCP tool as an **Artifact**, and fetch it back
by id:

```python
from navia import Navia

with Navia(resolve_connector=False) as client:
    info = client.artifact_upload("./report.pdf", display_name="Q4 report")
    client.artifact_download(info["artifactId"], "./downloaded.pdf")
```

Uploading alone only *stages* the file. To deliver it to the chat user, the
MCP **tool result** must declare the id under the reserved
`navia_artifacts` key of its structured content:

```json
{
  "navia_artifacts": [
    {"artifact_id": "<artifactId>", "filename": "report.pdf",
     "content_type": "application/pdf", "size": 12345}
  ]
}
```

Only `artifact_id` is required — the other fields are hints. The chat backend
then shows the file as a downloadable card on the assistant's message and in
the user's Artifacts list. Identical re-uploads by the same token are
deduplicated server-side (`deduped: true` in the response).

> `mcp_upload_attachment` / `mcp_download_attachment` are **deprecated**
> aliases of the old `/sdk/mcp/*` routes and will be removed in 0.5.0.

## OCR

Run OCR on a single document. By default the structured result is returned as
a dict; pass `output_format` together with `dest` to download the rendered
file instead:

```python
from navia import Navia

with Navia(resolve_connector=False) as client:
    result = client.ocr_extract("./document.pdf", mode="STRUCTURED")
    print(result["markdown"])

    # Render and download a file:
    client.ocr_extract(
        "./document.pdf", output_format="MARKDOWN", dest="./document.md"
    )
```

`mode` accepts `"PLAIN"`, `"STRUCTURED"` (default) or `"VLM"`;
`output_format` accepts `"MARKDOWN"`, `"JSON"` or `"PDF"`. Result keys are
snake_case (`markdown`, `regions`, `page_count`, …).

## Data extraction

Extract structured data from a document against a saved **template** or an
inline **definition** (provide exactly one). Non-PDF inputs are converted to
PDF server-side. The structured result is returned as a dict; pass
`output_format` (`"csv"`/`"xlsx"`) together with `dest` to download a table:

```python
from navia import Navia

with Navia(resolve_connector=False) as client:
    # Against a saved template:
    result = client.extraction_extract("./invoice.pdf", template_id="<id>")
    for field in result["fields"]:
        print(field["field_path"], field["value"], field["confidence"])

    # Against an inline definition + a hint:
    client.extraction_extract(
        "./invoice.pdf",
        definition={"fields": [{"key": "total", "label": "Total", "type": "currency"}]},
        hints="the grand total is bottom-right",
    )

    # Download a CSV of the extracted fields:
    client.extraction_extract(
        "./invoice.pdf", template_id="<id>", output_format="csv", dest="./out.csv"
    )
```

Result keys are snake_case: `data` (the structured record), `fields` (each
with `field_path`, `value`, `value_type`, `confidence`, `source_spans`) and
`usage`.

## Chat *(not implemented)*

```python
from navia import Navia

with Navia(resolve_connector=False) as client:
    # Raises NotImplementedError today.
    client.chat_completions({"messages": [{"role": "user", "content": "Hi"}]})
```

## Development

The project uses [uv](https://docs.astral.sh/uv/) and
[ruff](https://docs.astral.sh/ruff/) (pinned in `pyproject.toml`).

```bash
uv sync
uv run ruff check src tests
uv run ty check
uv run pytest
```

## Versioning & release

The package version lives in `[project].version` of `pyproject.toml` (a single
source of truth; `navia.__version__` is read from the installed package
metadata). Releases are driven by CI:

- pushes to **`develop`** publish a dev build to TestPyPI (the version is
  suffixed with `.devN` so each build is unique);
- pushes to **`main`** publish the exact `pyproject.toml` version to PyPI as
  `navia-sdk`.

Bump `version` in `pyproject.toml` before promoting a release to `main`.
