Metadata-Version: 2.5
Name: open-decision-ai
Version: 0.2.0
Summary: An open-weight, local decision model for software: typed probabilistic decisions from arbitrary state.
Project-URL: Homepage, https://github.com/open-decision/open-decision
Author: OpenDecision contributors
License-Expression: Apache-2.0
License-File: LICENSE
Keywords: calibration,classification,decision,fastapi,inference,local
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: Apache Software License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Python: >=3.11
Requires-Dist: fastapi>=0.116
Requires-Dist: httpx>=0.28
Requires-Dist: huggingface-hub>=0.34
Requires-Dist: numpy>=2.0
Requires-Dist: pydantic-settings>=2.10
Requires-Dist: pydantic>=2.10
Requires-Dist: safetensors>=0.5
Requires-Dist: torch>=2.8
Requires-Dist: transformers>=4.56
Requires-Dist: typer>=0.16
Requires-Dist: uvicorn[standard]>=0.35
Description-Content-Type: text/markdown

# OpenDecision

**An open-weight, local decision model for software.**

LLMs return strings. Software needs a *branch*: which path, how sure, act or escalate.
OpenDecision takes arbitrary state — text, JSON, logs, key/value dumps, conversations — plus a
**typed question**, and returns a **typed, probabilistic decision** with calibrated confidence.

```text
State + Typed Question  ->  Probabilistic Decision
```

```bash
curl -s localhost:8000/v1/decisions -H 'content-type: application/json' -d '{
  "state": {"customer_message": "When will my new card arrive?"},
  "question": {"type": "choice", "question": "What is the customer'\''s intent?",
               "options": ["card_arrival", "card_not_working", "exchange_rate", "top_up_failed"]}}'
```

```json
{"type": "choice", "decision": "card_arrival",
 "probabilities": {"card_arrival": 0.978, "card_not_working": 0.018,
                   "exchange_rate": 0.002, "top_up_failed": 0.003},
 "confidence": 0.978, "abstained": false, "model": "open-decision-150m"}
```

It runs entirely on your machine: ~150M parameters, ~10 ms per decision on a GPU, no API key,
no telemetry, no data leaving the host.

## Why this exists

A general LLM can answer that question, but you pay for it in latency, cost, and — most
importantly — you get prose you must parse, with a confidence number that means nothing.
For the narrow, repetitive decisions that make up real workflows:

```text
which queue does this ticket go to?        -> choice
is this incident customer-impacting?       -> binary
should we auto-approve, review, or reject? -> choice + confidence threshold
```

…a small encoder is the right tool. OpenDecision gives you:

* **a schema you can trust** — the decision is always one of your declared options, the
  probabilities are valid and sum to 1, and no free-form text is ever returned;
* **calibrated confidence** — ECE 0.047, so a 0.9 really is about 90% (see Benchmarks);
* **abstention** — below your threshold it returns `null` instead of guessing;
* **order invariance by construction** — shuffling your options changes nothing
  (measured flip rate: 0.0);
* **local-first** — after the model is downloaded, no internet is required at all.

OpenDecision explores the same broad problem as recent "machine-native" / "System One"-style
decision models, from an open, local-first perspective. It is an independent implementation,
not a reproduction of any proprietary system.

---

## Get the model

The weights are **not** in this repository. Download the archive and point the server at it —
it is fetched once, verified, extracted and cached under `~/.cache/open-decision/`.

> **Model:** `open-decision-150m` v0.2.0 · ModernBERT-base · Apache-2.0 · 558 MB zip
> **Download:** <!-- replace with your Google Drive share link -->
> `https://drive.google.com/file/d/<FILE_ID>/view`
> **sha256:** `<paste the checksum here>`

### Option A — let the server download it (recommended)

```bash
open-decision serve --model 'gdrive:<FILE_ID>#sha256=<checksum>'
```

`gdrive:<FILE_ID>` works with a bare id or a full Drive share link. Any plain URL works too, so
you can host the archive anywhere:

```bash
open-decision serve --model 'https://example.com/open-decision-150m-v0.2.0.zip#sha256=<checksum>'
```

The `#sha256=` fragment is optional but recommended — the download is rejected on a mismatch.
Re-running is instant; the cache is reused and the network is never touched again.
Set `OPEN_DECISION_CACHE` to move the cache, and `--offline` to forbid any download.

### Option B — download it yourself

Unzip it anywhere and pass the directory:

```bash
unzip open-decision-150m-v0.2.0.zip -d ./models
open-decision serve --model ./models/model
```

The directory must contain `config.json`, `model.safetensors`, `open_decision.json` and the
tokenizer files. Only `safetensors` weights are loaded and `trust_remote_code` is never
enabled, so nothing in a model directory can execute code.

---

## Install

> **Heads-up before you run anything:** `pip install open-decision` /
> `uv tool install open-decision` **will not work yet** — the package is not on PyPI. And a
> default install pulls the **CUDA build of PyTorch: ~5.4 GB**, which downloads for minutes with
> almost no output and looks like a freeze. Every command below pins the CPU build (**~950 MB**,
> ~1 minute) unless you actually need a GPU.

### Step 1 — install `uv`

```bash
curl -LsSf https://astral.sh/uv/install.sh | sh     # macOS / Linux
# Windows: powershell -c "irm https://astral.sh/uv/install.ps1 | iex"
```

Restart the shell, then check: `uv --version` (0.8 or newer).

### Step 2 — get the code

Until the package is published, install from a checkout:

```bash
git clone https://github.com/<gh-user>/open-decision.git
cd open-decision
```

### Step 3 — install

Pick **one**.

**A. As a command-line tool** (a global `open-decision` binary, isolated environment):

```bash
uv tool install --index pytorch-cpu=https://download.pytorch.org/whl/cpu \
                --index-strategy unsafe-best-match .
```

Note `uv tool install` ignores `UV_TORCH_BACKEND`, so the `--index` flags above are how you get
the small CPU build. Without them you get the 5.4 GB CUDA build.

**B. As a project** (for development, or to use it as a library):

```bash
uv sync --torch-backend cpu       # ~950 MB, about a minute
uv run open-decision --version
```

**C. In your own virtualenv:**

```bash
uv venv && uv pip install --torch-backend cpu .
```

**With a GPU** — drop the CPU pins and let uv detect your CUDA version. Expect a multi-GB
download; run it with `-v` so you can see progress:

```bash
uv sync --torch-backend auto -v
```

Verify whichever you chose:

```bash
open-decision --version            # or: uv run open-decision --version
```

### Step 4 — get the model and serve it

See [Get the model](#get-the-model) above for the download link and checksum.

```bash
open-decision serve --model 'gdrive:<FILE_ID>#sha256=<checksum>' --abstain-threshold 0.7
```

First start downloads ~558 MB and prints progress; later starts are instant from the cache.
Wait for `{"status":"ready"}` on `/ready`, then:

```bash
curl -s localhost:8000/v1/decisions -H 'content-type: application/json' -d '{
  "state": {"customer_message": "My card is not working at the ATM"},
  "question": {"type": "choice", "question": "What is the customer'"'"'s intent?",
               "options": ["card_arrival", "card_not_working", "exchange_rate"]}}'
```

### Troubleshooting installs

| symptom | cause and fix |
| --- | --- |
| `uv tool install open-decision` hangs, then fails | not on PyPI yet — install from a checkout (Step 3) |
| install seems frozen for minutes | downloading the CUDA PyTorch build (~5.4 GB). Ctrl-C and add the CPU pins above, or add `-v` to watch progress |
| `open-decision: command not found` | `~/.local/bin` is not on `PATH`; run `uv tool update-shell` and restart the shell |
| server returns HTML error on model download | the Drive file is not shared as "Anyone with the link", or the daily quota is exhausted — download it manually and pass the directory |
| `/ready` returns 503 for a long time | still downloading or loading the model; watch the server log |
| slow decisions on CPU | expected — CPU is roughly 10x slower than GPU. Use `--device cuda` with a CUDA install, or Docker |

### No-install option: Docker

Nothing to install but Docker itself — the model is fetched on first start and cached in the
mounted volume:

```bash
# GPU
docker run --gpus all -p 8000:8000 -v ~/.cache/open-decision:/root/.cache/open-decision \
  -e OPEN_DECISION_MODEL='gdrive:<FILE_ID>' ghcr.io/<gh-user>/open-decision:latest

# CPU
docker run -p 8000:8000 -v ~/.cache/open-decision:/root/.cache/open-decision \
  -e OPEN_DECISION_MODEL='gdrive:<FILE_ID>' ghcr.io/<gh-user>/open-decision:cpu
```

---

## Using it

### Python SDK

```python
from open_decision import Client

client = Client("http://localhost:8000")
result = client.decide(
    state={"customer_message": "My card is not working at the ATM"},
    question={"type": "choice", "question": "What is the customer's intent?",
              "options": ["card_arrival", "card_not_working", "exchange_rate"]},
)
print(result.decision, result.confidence, result.probabilities)
```

`AsyncClient` has the same methods; concurrent calls are batched into one forward pass.

### In-process, no server

```python
from open_decision import DecisionEngine

engine = DecisionEngine.load("gdrive:<FILE_ID>")     # or a local directory
decision = engine.decide(
    state="status=failed retry_count=4 region=eu-west",
    question={"type": "binary", "question": "Should this job be retried again?"},
)
print(decision.decision, decision.probability)
```

### Several decisions about one state

Workflows usually need more than one narrow decision. They are all batched into a single
forward pass:

```json
{"state": {"...": "..."},
 "questions": [
   {"id": "intent", "type": "choice", "options": ["refund", "reship", "investigate"]},
   {"id": "urgent", "type": "binary", "question": "Is this urgent?"}
 ],
 "abstain_threshold": 0.7}
```

```json
{"decisions": {"intent": {"...": "..."}, "urgent": {"...": "..."}},
 "model": "open-decision-150m"}
```

### Acting on confidence

The whole point is that you can branch on the number:

```python
d = client.decide(state=ticket, question={"type": "choice", "options": [...]})
if d.abstained or d.confidence < 0.60:
    route_to_human(ticket)
elif d.confidence < 0.80:
    queue_for_review(d.decision, ticket)
else:
    execute(d.decision, ticket)
```

---

## Benchmarks

v0.2.0, on a held-out test set of **4,660 examples** (synthetic + Banking77 + CLINC150 + TREC).
Reproduce with the training repo's `evaluate` command; full numbers in `runs/v0.2/eval.json`.

### Against baselines

| | accuracy | ECE ↓ | NLL ↓ |
| --- | --- | --- | --- |
| **OpenDecision 150M** | **0.670** | **0.047** | **0.718** |
| TF-IDF + logistic regression | 0.367 | 0.095 | 1.423 |
| majority class | 0.350 | 0.105 | 1.515 |

**+30 accuracy points over a bag-of-words cross-encoder** on the same data.

### By dataset

| dataset | n | accuracy | ECE | notes |
| --- | --- | --- | --- | --- |
| CLINC150 | 562 | **0.932** | 0.022 | 150-way intent (8 options sampled) |
| Banking77 | 596 | **0.923** | 0.027 | 77-way banking intent |
| synthetic | 1712 | 0.685 | 0.102 | 20 domains, 3 languages |
| TREC | 451 | 0.333 | 0.055 | **zero-shot** — never trained on (chance 0.167) |

The headline 0.670 blends supervised intent classification with zero-shot transfer. For
in-domain intent classification the model is in the low 90s.

### Confidence is the point

Selective accuracy — the model answers only when confident enough:

| threshold | coverage | accuracy on answered |
| --- | --- | --- |
| 0.6 | 50.5% | **91.5%** |
| 0.7 | 41.5% | **95.3%** |
| 0.8 | 34.9% | **97.2%** |
| 0.9 | 28.2% | **98.2%** |

Confidence separates answerable from unanswerable input: feeding it garbage
(`"asdf qwer"`, `{"a": 1}`, `"ok"`) yields 0.29–0.59 confidence versus ~1.00 on a clear intent
— an AUC of 1.00 on banking intents and 0.84 on open-ended routing questions.

### Other properties

| property | value |
| --- | --- |
| option-order invariance | flip rate **0.0**, mean JSD 3e-16 (500 examples) |
| held-out domains | 0.600 vs 0.702 on seen domains |
| languages | en 0.698 · fr 0.609 · ar 0.539 |
| state formats | text 0.716 · logs 0.668 · kv 0.628 · conversation 0.560 · json 0.517 |
| throughput | 326 decisions/s (RTX 5090, bf16) |
| memory | 0.96 GB GPU at inference |

### Suggested policy

```text
confidence >= 0.80  ->  execute automatically   (97.2% accurate, ~35% of traffic)
0.60 - 0.80         ->  execute with review
confidence <  0.60  ->  route to a human
```

Calibrate these on your own labelled sample — the numbers above come from this test set.

### Known limitations — please read

* **`score` questions are not functional in v0.2.** The score head is effectively a constant
  predictor (0.47 spread on a 1–10 scale). The API accepts them; do not rely on the output.
* **`binary` is weaker than `choice`** — 0.638 accuracy and ECE 0.136 (choice: 0.045). Usable
  with a high threshold and review, not for unattended automation.
* **Arabic (0.539) lags English (0.698)**; JSON and conversation-shaped states are the weakest
  formats.
* **Confidence is not comparable across question types.** A "real" banking intent scores ~1.00
  while an open-ended routing question scores ~0.63. Pick the threshold per workflow.
* Trained largely on synthetic data distilled from one open-weight teacher, so it inherits that
  teacher's biases. Validate on your own data before trusting it.

---

## Decision types

| type | request | response |
| --- | --- | --- |
| `choice` | `{"type":"choice","options":["A","B","C"],"question":"optional"}` | `decision` (one of the options), `probabilities`, `confidence` |
| `binary` (alias `noul`) | `{"type":"binary","question":"Should this be escalated?"}` | `decision` (bool), `probability` (P(yes)), `confidence` |
| `score` | `{"type":"score","question":"How urgent?","min":1,"max":10}` | `score`, `std`, `confidence` — **not functional in v0.2** |

Below `abstain_threshold` the decision is `null` and `abstained` is `true`; the probabilities
are still returned so your application can apply its own policy.

## API

| Method | Path | Purpose |
| --- | --- | --- |
| `POST` | `/v1/decisions` | decisions for one state and one or more questions |
| `GET` | `/v1/models` | model id, version, device, dtype, calibration |
| `GET` | `/health` | process is up |
| `GET` | `/ready` | model loaded (503 while loading / downloading) |
| `GET` | `/docs`, `/openapi.json` | interactive docs and OpenAPI schema |

Full schema in [docs/api.md](docs/api.md); runnable samples in [examples/](examples/).

## CLI

```bash
open-decision serve --model <dir | gdrive:ID | https://...> [--port 8000] [--abstain-threshold 0.7]
open-decision health
open-decision model-info
echo '{"state":"...","question":{"type":"binary","question":"..."}}' | open-decision decide
echo '{...}' | open-decision decide --model ./models/model     # no server needed
```

## Architecture

```text
                 ┌──────────────────── OpenDecision model ─────────────────────┐
 state ─────┐    │ per option: "[CLS] question / option [SEP] state [SEP]"     │
 question ──┼──▶ │       ModernBERT encoder ──▶ [CLS] ──▶ choice head ──▶ logit│──▶ softmax
 options ───┘    └────────────────────────────────────────────────────────────┘    over options
                                    │
                temperature scaling (optional) ──▶ abstention threshold ──▶ typed decision
```

Choice and binary questions are scored **one option at a time** by a cross-encoder, which is
why the model is invariant to option order by construction and accepts any option strings
rather than a fixed label set. See [docs/model-format.md](docs/model-format.md).

## Local-first / privacy

* Inference runs entirely on your machine. After the model is cached, `--offline` guarantees no
  network access.
* No telemetry, no external API calls, no request-body logging by default.
* Optional API key, CORS allow-list, body-size limit, max state length and question count.

Details in [docs/privacy.md](docs/privacy.md).

## Configuration

Every setting is a CLI flag or an `OPEN_DECISION_*` environment variable:

| Setting | Default | Notes |
| --- | --- | --- |
| `OPEN_DECISION_MODEL` | – | directory, Hub id, `gdrive:<id>`, or archive URL (required) |
| `OPEN_DECISION_CACHE` | `~/.cache/open-decision` | where downloaded models are stored |
| `OPEN_DECISION_HOST` / `PORT` | `127.0.0.1` / `8000` | Docker images use `0.0.0.0` |
| `OPEN_DECISION_DEVICE` / `DTYPE` | `auto` | `cpu`, `cuda`, `mps` |
| `OPEN_DECISION_ABSTAIN_THRESHOLD` | `0.0` | never abstain by default |
| `OPEN_DECISION_CALIBRATION` | `true` | apply stored temperature scaling |
| `OPEN_DECISION_BATCHING` | `true` | dynamic batching across requests |
| `OPEN_DECISION_API_KEY` | – | enables `Authorization: Bearer` / `X-API-Key` |
| `OPEN_DECISION_OFFLINE` | `false` | never download anything |
| `OPEN_DECISION_LOG_REQUESTS` | `false` | request bodies are never logged unless enabled |

## Development

```bash
uv sync
uv run pytest -q
uv run ruff check src tests && uv run mypy src
```

Tests build a tiny random model on the fly; no downloads required.

## Roadmap

* fix the score head; strengthen binary
* better Arabic and structured-state (JSON) performance
* larger encoder option, quantization (int8/fp8), ONNX
* multi-label and ranking decision types

## License

Apache-2.0 for the code. Model weights are released under Apache-2.0 and carry their own model
card; third-party datasets used in training keep their own licenses.
