Metadata-Version: 2.4
Name: bespokelabs-nimble
Version: 0.1.0
Summary: Typed Python clients for hosted Bespoke Nimble models
Author: Bespoke Labs
Classifier: Development Status :: 3 - Alpha
Classifier: Programming Language :: Python :: 3
Classifier: Typing :: Typed
Requires-Python: >=3.10
Requires-Dist: httpx<1,>=0.27
Requires-Dist: pydantic<3,>=2.9
Provides-Extra: dev
Requires-Dist: build<2,>=1; extra == 'dev'
Requires-Dist: pytest-asyncio<2,>=0.24; extra == 'dev'
Requires-Dist: pytest<10,>=8; extra == 'dev'
Requires-Dist: ruff>=0.8; extra == 'dev'
Requires-Dist: twine<7,>=6; extra == 'dev'
Description-Content-Type: text/markdown

# Bespoke Nimble Python SDK

A lightweight Python client for hosted [Bespoke Nimble](https://github.com/bespokelabsai/nimble): ask typed questions, get decisions and candidate probabilities.

- Package: `bespokelabs-nimble`
- Import: `from bespokelabs.nimble import Nimble, AsyncNimble, Choice, Noul, Score`
- Python 3.10+; runtime dependencies are HTTPX and Pydantic. No model weights or GPU libraries.
- Uses Nimble's existing `POST /v1/systemone` endpoint.

Configure the client with your Nimble deployment URL. The package does not assume a public production API URL.

## Install

```bash
uv pip install bespokelabs-nimble
# Or add it to a project:
uv add bespokelabs-nimble
# Or use pip:
pip install bespokelabs-nimble
```

From a source checkout:

```bash
uv pip install -e ./nimble-sdk
```

## Quickstart

Point the client at your hosted Nimble server. The URL may be the service root or end in `/v1`; gateway path prefixes are preserved.

```bash
export BESPOKE_NIMBLE_BASE_URL="https://your-nimble-host.example"
export BESPOKE_API_KEY="your-api-key"  # Only if your gateway uses bearer authentication.
```

```python
from bespokelabs.nimble import Choice, Nimble, Noul, Score

with Nimble() as client:
    result = client.system_one(
        state="I was charged twice. Please refund the duplicate payment.",
        questions={
            "refund": Noul(instructions="Does the customer request a refund?"),
            "department": Choice(
                instructions="Which department should handle this request?",
                criteria={
                    "billing": "Charges, payments, and refunds",
                    "technical": "Software bugs and outages",
                },
            ),
            "urgency": Score(
                instructions="Assess operational urgency.",
                criteria=[
                    "Service works normally",
                    "Some functionality unavailable",
                    "Complete outage",
                ],
            ),
        },
    )

print(result.nouls["refund"].noul)  # P(true), from 0 to 1
print(result.choices["department"].choice)  # e.g. "billing"
print(result.choices["department"].probabilities)  # All candidate probabilities
print(result.scores["urgency"].score)  # Expected level, from 0 to 2
print(result.usage.input_tokens)
```

`state` and `instructions` can also be JSON objects or lists. The server serializes them as text for the model. Raw question dictionaries with the same shape as the HTTP API are accepted.

### Question and answer types

| Question | Input | Answer |
| --- | --- | --- |
| `Noul` | Yes/no instructions; optional definitions of true and false | `noul`: probability of true |
| `Choice` | 2–26 option keys mapped to descriptions or `None` | `choice`, `probabilities`, `confidence` |
| `Score` | 2–26 ordered rubric descriptions, lowest first | `score`, `legend`, `probabilities`, `confidence` |

`Noul` preserves the API's name for a yes/no probability. To customize its criteria:

```python
from bespokelabs.nimble import Noul, NoulCriteria

supported = Noul(
    instructions="Is the claim supported by the evidence?",
    criteria=NoulCriteria(yes="Supported", no="Unsupported"),
)
```

`Score.score` is `sum(index * probability[index])`, using zero-based indices. It is not the most likely level, and it is not normalized to 0–1 unless there are two levels. `confidence` measures distribution concentration, not calibrated correctness. Probability thresholds need evaluation on your own data.

Answers are accessible through `result.answers` or the typed `result.nouls`, `result.choices`, and `result.scores` dictionaries. Responses retain extra server fields through Pydantic's `model_extra`. `result.raw_http_response` exposes the HTTPX response, including timing headers. Unknown answer types or missing/mismatched answers raise `APIResponseError`.

## Async

```python
import asyncio
from bespokelabs.nimble import AsyncNimble, Noul


async def main():
    async with AsyncNimble() as client:
        result = await client.system_one(
            "The checkout service is down.",
            {"incident": Noul(instructions="Is there an active service incident?")},
        )
        print(result.nouls["incident"].noul)


asyncio.run(main())
```

Reuse one client for multiple calls. `asyncio.gather` works for concurrent requests; keep concurrency within `await client.limits()`. This SDK does not add a batch endpoint or a GPU inference runtime.

## Configuration and authentication

```python
client = Nimble(
    base_url="https://your-nimble-host.example",
    api_key="your-api-key",
    model="nimble-latest",
    timeout=120.0,
    max_retries=2,
)
# Use a context manager or call client.close() when done.
```

| Setting | Environment fallback | Default |
| --- | --- | --- |
| `base_url` | `BESPOKE_NIMBLE_BASE_URL` | Required |
| `api_key` | `BESPOKE_API_KEY` | No authorization header |
| `model` | — | `nimble-latest` |
| `timeout` | — | 120 seconds per HTTPX operation |
| `max_retries` | — | 2 retries, up to 3 attempts |

An explicit argument takes precedence over its environment variable. `api_key` sends `Authorization: Bearer ...`. The SDK does not create API keys or add authentication to the upstream server. Your gateway must enforce bearer authentication. Public deployments can omit the key (and leave `BESPOKE_API_KEY` unset).

For a private Modal proxy, pass the headers required by that deployment:

```python
import os
from bespokelabs.nimble import Nimble

with Nimble(
    default_headers={
        "Modal-Key": os.environ["MODAL_KEY"],
        "Modal-Secret": os.environ["MODAL_SECRET"],
    }
) as client:
    print(client.health())
```

`default_headers` also supports other gateway authentication schemes. Redirects are not followed. A custom `httpx.Client` or `httpx.AsyncClient` can be passed as `http_client`; the caller retains ownership and must close it. Requests use the SDK's timeout and headers, plus the injected client's defaults.

Retries cover transport failures, HTTP 408, HTTP 429, and HTTP 5xx (including upstream overload status 529). Other errors fail immediately. Retries use exponential backoff with jitter and honor numeric or HTTP-date `Retry-After` values, capped at 60 seconds. Set `max_retries=0` to disable retries. A retry may run inference again; the upstream API does not advertise idempotency support. Async cancellation propagates immediately.

Cold starts can exceed the default timeout. Increase `timeout` for scale-to-zero hosting, or keep the service warm. The configured timeout applies to individual network operations; it is not a deadline for all attempts combined.

## Discovery and errors

```python
from bespokelabs.nimble import APIStatusError, Nimble

with Nimble() as client:
    print(client.models())  # GET /v1/models
    print(client.limits())  # GET /v1/limits
    try:
        print(client.health())  # GET /health; unavailable services can return HTTP 503
    except APIStatusError as error:
        print(error.status_code, error.request_id)
```

`AuthenticationError` and `RateLimitError` specialize `APIStatusError`. Transport errors raise `APIConnectionError` or `APITimeoutError`; incompatible success responses raise `APIResponseError`. All inherit from `NimbleError`. Local validation raises `ValueError` (including Pydantic `ValidationError`) before sending a request. Error messages omit request bodies and credentials; HTTP failures expose `error.response` for explicit inspection.

The SDK validates the hosted API's limit of 64 questions and 26 candidates per question. Token limits and aggregate request budgets are enforced by the server; use `client.limits()` to inspect your deployment.

## Hosting contract

Use the server already provided by [Nimble's serving module](https://github.com/bespokelabsai/nimble/blob/main/nimble/serving/server.py). Its [Modal deployment guide](https://github.com/bespokelabsai/nimble/blob/main/docs/MODAL_SERVING.md) describes deploying the model. This SDK talks to that service directly.

Request:

```http
POST /v1/systemone
Content-Type: application/json
```

```json
{
  "model": "nimble-latest",
  "state": "Please refund the duplicate payment.",
  "questions": {
    "refund": {"type": "noul", "instructions": "Does the customer request a refund?"}
  }
}
```

Response:

```json
{
  "model": "nimble-latest",
  "answers": {"refund": {"type": "noul", "noul": 0.99}},
  "usage": {"input_tokens": 200, "output_tokens": 2}
}
```

Compatibility was checked against Nimble revision `f136b3f75721fda4ea961f73993cc50b08488835`, its pinned `openjev-sglang` revision `7f84bedc169439f03379c2fa8d00ada220af2295`, and its recorded deployment smoke response. Tests replay that recorded response without invoking a model. A successful local test run is not a live GPU deployment test.

The Python interface follows [TypeSafe's SDK](https://docs.typesafe.ai/sdk/python). Package/import naming follows [Curator](https://github.com/bespokelabsai/curator), and `BESPOKE_API_KEY` follows the [MiniCheck SDK convention](https://docs.bespokelabs.ai/models/bespoke-minicheck/api). The `bespokelabs` directory is an implicit namespace so this package does not add a competing root `__init__.py`.

## Develop and release

From `nimble-sdk/`:

```bash
uv venv
uv pip install -e '.[dev]'
uv run --no-project pytest
uv run --no-project ruff check .
uv run --no-project ruff format --check .
uv build
uv run --no-project twine check dist/*
```

Before the first public release, select the repository/license metadata, confirm control of the PyPI project name, and configure PyPI publishing credentials or Trusted Publishing. Then publish the reviewed artifacts:

```bash
uv publish dist/*
```

Publishing is a separate release step; nothing in this checkout automatically uploads packages. After publication, verify installation in a clean environment with `uv pip install bespokelabs-nimble` and point it at your running deployment.
