Metadata-Version: 2.4
Name: solvi
Version: 0.6.1
Summary: Decision systems from a catalog of Python functions and checks: a strategist plans the flow, every answer comes with a quote or formula and a verifiable trace
Keywords: decision systems,rules,document extraction,ModernBERT,explainability,verification,audit trail
Author: mxkuzn
Author-email: mxkuzn <hi@mxkuzn.dev>
License-Expression: Apache-2.0
License-File: LICENSE
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: Programming Language :: Python :: 3
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Topic :: Software Development :: Libraries
Requires-Dist: numpy
Requires-Dist: scipy
Requires-Dist: pydantic>=2
Requires-Dist: mcp>=2 ; extra == 'mcp'
Requires-Dist: torch ; extra == 'model'
Requires-Dist: transformers ; extra == 'model'
Requires-Dist: onnxruntime ; extra == 'onnx'
Requires-Dist: tokenizers ; extra == 'onnx'
Requires-Dist: huggingface-hub ; extra == 'onnx'
Requires-Dist: fastapi>=0.100 ; extra == 'serve'
Requires-Dist: uvicorn ; extra == 'serve'
Requires-Python: >=3.10
Project-URL: Homepage, https://github.com/solvi-ai/solvi
Project-URL: Documentation, https://github.com/solvi-ai/solvi/blob/main/docs/guide.md
Project-URL: Issues, https://github.com/solvi-ai/solvi/issues
Project-URL: Models, https://huggingface.co/solvi-ai
Provides-Extra: mcp
Provides-Extra: model
Provides-Extra: onnx
Provides-Extra: serve
Description-Content-Type: text/markdown

# solvi

[![PyPI](https://img.shields.io/pypi/v/solvi.svg)](https://pypi.org/project/solvi/)
[![CI](https://github.com/solvi-ai/solvi/actions/workflows/ci.yml/badge.svg)](https://github.com/solvi-ai/solvi/actions/workflows/ci.yml)
[![License: Apache-2.0](https://img.shields.io/badge/license-Apache--2.0-blue.svg)](LICENSE)
[![Hugging Face](https://img.shields.io/badge/%F0%9F%A4%97%20Hugging%20Face-solvi--ai-yellow)](https://huggingface.co/solvi-ai)

Build decision systems from a catalog of Python functions and checks plus typed questions, and get answers you can verify.

## Why

You describe a task with plain Python functions (computations, checks, answer rules) and questions with typed answers
(yes/no, a choice, a score, "not stated", a span of the text, a ranking, a number range). For each request, a strategist
plans which functions and checks to run for the asked questions. Every answer comes with:

- a **confidence** (calibratable per question);
- a **reason you can check**: the rule inputs, a formula over computed facts, or a quote with character offsets in the
  source document;
- a **hash-chained trace** of every step, which can be re-executed later to confirm the answer or pinpoint the step that
  was altered.

When something cannot be computed, a function fails, or a rule returns an answer outside the allowed options, solvi
abstains instead of guessing. A failed hard check always overrides any model confidence.

**Grounded decisions.** Fuzzy proposes, deterministic decides, everything is in the trace. Each fact and answer records its
provenance — `given`, `computed`, `quoted`, `decided` (a model's choice among options, with probabilities) or `learned` —
and model-backed steps record the model's id and fingerprint, so a replay can tell when the model changed since a decision.
A model's quote that is not literally the text at its offsets, or a choice outside its options, is rejected and counted
(`system.stats`); a fallback producer runs or the question abstains. `print(res.audit())` shows what each answer rests on,
which safeguards fired, and how much of its support is deterministic. A decision without models and one with models are the
same system ([examples/12_grounded_audit.py](examples/12_grounded_audit.py)).

And it is fast. The strategist plans a flow over a 10 000-part catalog in about 6 ms and runs only the parts the questions
need (2.4% of that catalog). Hard checks run first, so a failing one skips the expensive rest; independent slow parts (API
calls, model inference) run in parallel. On an insurance-claim desk with slow services
([examples/09_strategy_at_scale.py](examples/09_strategy_at_scale.py)) a full decision takes 463 ms instead of 1 122 ms for a
script that computes everything, and 152 ms when an expired policy settles the claim first.

## Try it

- [solvi playground](https://huggingface.co/spaces/solvi-ai/playground): write a decision task in Python and run it, watch the
  strategist's plan, tamper with a trace and see the replay catch it, learn rules from examples.
- [solvi documents](https://huggingface.co/spaces/solvi-ai/documents): cited, typed answers from contracts, invoices, receipts,
  leases and more — the ModernBERT extractor (ONNX) and the decisions both run in your browser; add a field by describing it.
- [solvi arcade](https://huggingface.co/spaces/solvi-ai/arcade): game agents that explain every move — tic-tac-toe, maze,
  minesweeper, 20 questions, Mafia detective, a bot arena, and "hack the trace".
- [solvi realms](https://huggingface.co/spaces/solvi-ai/realms): an endless strategy game whose factions are solvi systems —
  tested for 100 000 turns: flat decision time (~0.3–0.6 ms), bounded memory and state, every sampled trace replay OK.
- All run **entirely in your browser** (Pyodide): no server, no GPU, nothing you type leaves the page.
- Models: [solvi-ai/solvi-large](https://huggingface.co/solvi-ai/solvi-large) (typed decisions, 396M, preview),
  [solvi-ai/solvi-base](https://huggingface.co/solvi-ai/solvi-base) (the same answers on a CPU / in ONNX, 150M, preview),
  [solvi-ai/extract-base](https://huggingface.co/solvi-ai/extract-base) (fields by description) and
  [solvi-ai/extract-receipts](https://huggingface.co/solvi-ai/extract-receipts). Each model card states what the model was
  measured on, how well it does, and its limits; all models are listed at [huggingface.co/solvi-ai](https://huggingface.co/solvi-ai).

## Gallery

[gallery/](gallery) — twelve decision tasks across directions (support triage, email routing, content guard, security alerts,
AI-agent audit, release rollout, KYC/AML, clinical screening, credit with adverse-action reasons, procurement 3-way match,
double-charge refunds, predictive maintenance), each with scenarios, a runner and a side-by-side against an answer-only model.

## Install

```bash
pip install solvi              # core: rules, checks, learned answer heads (numpy, scipy, pydantic)
pip install "solvi[model]"     # + torch, transformers: ModernBERT field extractors for documents and the decider
pip install "solvi[onnx]"      # + onnxruntime, tokenizers: the decider (solvi.decide) on CPU without torch
pip install "solvi[serve]"     # + fastapi, uvicorn: `solvi serve app.py:system` — the questions over HTTP (also --mcp)
```

Requires Python 3.10+.

## Quickstart (core only, no model)

```python
from datetime import date
from solvi import Answer, Catalog, Question, System

cat = Catalog()

@cat.fn                                            # argument names = facts it reads; function name = fact it sets
def days_requested(start, end):
    return (end - start).days + 1

@cat.fn
def remaining_after(balance, days_requested):
    return balance - days_requested

@cat.check(hard=True, then={"approve": "reject"})  # if this check is False, "approve" is forced to "reject"
def enough_balance(remaining_after):
    return remaining_after >= 0

@cat.check
def enough_notice(start, today, days_requested):
    return days_requested < 5 or (start - today).days >= 14

@cat.rule("approve")
def approve(enough_notice):
    return "approve" if enough_notice else "needs_manager"

system = System(cat, [Question("approve", "Approve the leave?",
                               Answer.choice(["approve", "needs_manager", "reject"]),
                               checkpoints=["enough_balance"])])
res = system.ask({"start": date(2026, 10, 19), "end": date(2026, 10, 23),
                  "today": date(2026, 9, 25), "balance": 14})
print(res["approve"].answer, res["approve"].confidence, res["approve"].why)
print(res.computed_state)
print(res.trace.replay(cat))
```

Output:

```
approve 1.0 enough_notice = True
days_requested           = 5
remaining_after          = 9
enough_balance           = True
enough_notice            = True
{'ok': True, 'steps': 5, 'mismatches': [], 'models': []}
```

With `"balance": 3` the hard check fails and the answer is `reject` with `status == "forced"`, whatever the rule says.
`solvi.show.show(res, cat)` prints answers, the planned flow, the computed state and the replay result in one go.

## How it works

- **Catalog.** `@cat.fn` (computation), `@cat.check` (bool), `@cat.extract` (value from text, returned as a `Quote` with
  offsets) and `@cat.rule(question)` (answer rule). A part's contract is its signature: argument names are the facts it
  reads, the function name is the fact it sets. Type hints are optional and become the facts' types
  (`def risk_score(risk_points: dict[str, float]) -> float`): producer and consumer types are checked when a part is
  registered, values are validated / coerced with pydantic at run time, and a value that fails is rejected like an
  ungrounded quote (safeguard `type_rejected`). Untyped parts cost nothing.
- **Questions.** `Question(name, text, Answer.yes_no() | Answer.choice([...]), checkpoints=[...])`. Questions without a
  rule get a small answer head trained from labeled examples (`system.fit`) or a readable learned rule list
  (`system.learn_rule`).
- **Strategist.** For each question it walks backwards from the rule's arguments (or the learned features) through the
  catalog signatures to the keys of `init_state`, adds the question's checkpoints and every check that touches a computed
  fact. Everything else in the catalog is not executed; the flow records why each part was taken or skipped.
- **Execution.** Each part runs once, even if several questions need it. Hard checks and their inputs run first; when one
  fails, the steps only the settled questions needed are skipped (`res.trace.skipped`). With `System(..., workers=8)`
  independent steps run in parallel threads as soon as their inputs are ready. Results go into `computed_state` with their
  provenance; extracted values keep their quote. Each step record is hashed and chained to the previous one in flow order,
  so the trace does not depend on scheduling.
- **Answers and trace.** `res[q].answer / .confidence / .why / .status` (`ok`, `forced`, `abstain`), plus
  `res.trace.replay(catalog)`, which recomputes every step from recorded inputs and reports mismatches, broken hash links
  and quotes outside the text.

## Typed decisions with a model

Types declare questions; the model proposes; checks decide. The fields of a pydantic model are the questions, their types
the kinds (one option, several, an ordered score, yes/no); a decider (`solvi.decide`, `solvi[onnx]` or `solvi[model]`)
answers them about a text or a JSON / pydantic state — several in one forward pass when the checkpoint can — with
probabilities, a calibrated confidence and act / escalate. Hard checks, constraints and rules still decide.

```python
from typing import Literal
from pydantic import BaseModel, Field
from solvi import Catalog, Scale, System
from solvi.decide import DecideModel

class Triage(BaseModel):
    team: Literal["billing", "technical", "shipping"] = Field(description="Which team should handle this ticket?")
    urgency: Scale[Literal["low", "medium", "high", "critical"]] = Field(description="How urgent is it?")
    angry: bool = Field(description="Is the customer angry?")
    topics: list[Literal["refund", "delay", "bug"]] = Field(description="What does the ticket mention?")

model = DecideModel.load("solvi-ai/solvi-base")         # or a local folder; solvi_decide.json says what it can do
cat = Catalog()
questions = model.questions(cat, Triage, text_fact="ticket", escalate_below=0.6)

@cat.check(hard=True, then={"urgency": "critical"})         # a legal threat is critical, whatever the model says
def no_legal_threat(ticket) -> bool:
    return "lawyer" not in str(ticket).lower()
questions[1].checkpoints.append("no_legal_threat")

res = System(cat, questions).ask({"ticket": {"subject": "Charged twice", "body": "Refund my double payment!",
                                             "customer": {"tier": "pro"}}})
print({q: (r.answer, r.status) for q, r in res.results.items()})
print(res.audit("team"))           # probabilities, the model's fingerprint, the shared pass, what escalated and why
```

A state is read as key paths (`customer.tier: pro`), the format the decider is trained on; an unsure or escalated answer
abstains with the reason (`system.stats["model_escalated"]`, `["low_confidence"]`). The checkpoint contract is in
[docs/decide_format.md](docs/decide_format.md); [examples/15_typed_decisions.py](examples/15_typed_decisions.py) runs the
whole story with a stand-in model. Published deciders are previews: read the model card before relying on one, and fit
it on 30–60 labelled examples of your task (`part.fit`, `part.calibrate_for`) — checks, constraints and escalation are what
make the answers safe to act on, not the model alone.

Every answer is a value and a confidence, and the types also declare answer primitives: `Maybe[T]` ("not stated" —
`solvi.Unknown` — is a real answer, unlike an abstention), `Span[float]` (an exact piece of the text, parsed), `Rank[...]`
(the top k, in order), `Estimate[0, 7, 14]` (a number with an interval), and evidence quotes on any answer
(`Claim(value, evidence=[...])`, `Question(require_evidence=True)`) — each checked in the text, from rules or a model
([guide](docs/guide.md#answer-primitives-not-stated-evidence-spans-rankings-estimates),
[examples/16_primitives.py](examples/16_primitives.py)).

## Escalation with a guarantee, several models, serving

- **A guaranteed risk.** `part.act_guard(examples, risk=0.10)` calibrates on a few hundred labelled examples of your stream
  so that P(answered alone and wrong) ≤ 10% for inputs like them (conformal risk control); the audit shows the promise
  behind every answer, or says there is none. `part.conformal(examples)` gives the person who takes an escalation a short
  list of candidates. Near ties escalate (`min_margin=`), and the answer does not depend on the order the options are listed
  in (sorted by default).
- **Any decision model.** `solvi.systemone.systemone(url, model)` puts any `POST /v1/systemone` service (Jev, Kev, Von,
  Laya-serve, …) behind your rules; `Cascade`, `Vote` and `Route` (`solvi.multi`) combine models — small first, a larger
  one only when the small one escalates, or answer only when models of different families agree — under one guarantee.
- **Serving and operations.** `solvi serve module:system` exposes the questions over HTTP (OpenAPI from the same types),
  MCP and the System One API; `await system.aask(...)` runs async parts concurrently with timeouts; `costs="measured"`
  lets the planner pick the fastest equivalent source and switch when it slows down. `TraceStorage` keeps decisions with a
  hash chain across them; `solvi diff` shows which stored decisions a rule or model change would flip; `solvi test`,
  `solvi check` and the honesty suite (`solvi honesty`) belong in CI.

## Planning around dead ends and costs (code strategist)

The default strategist needs the inputs of every producer of a fact. `solvi.strategy.ModelStrategist()` plans around
producers whose inputs are never given, and with `producers="equivalent"` picks the cheapest verified plan by declared
`cost=` (an exact 0/1 program; hard checks that govern a question always stay in the plan). No model is involved; the plan
is one hashed record in the trace and replay re-verifies it.

```python
from solvi.strategy import ModelStrategist
system = System(cat, questions, strategist=ModelStrategist(producers="equivalent"))
```

A segment model that proposes producers when costs are not declared, and `solvi.aliases` (wiring parameter names that match
no fact), ship as **experimental**; their weights are not published. See [docs/strategist.md](docs/strategist.md) and
[examples/17_model_strategist.py](examples/17_model_strategist.py).

## Extract from documents

With `solvi[model]`, fields are found by a fine-tuned ModernBERT extractor. The extracted value is always a span of the
document, so every answer built on it can be cited.

```python
from solvi.extract_multi import MultiSpanExtractor

ex = MultiSpanExtractor(["total", "date"])          # ModernBERT-large, one pass per document for all fields
ex.fit(train_docs, train_spans, epochs=4)           # train_spans: [{"total": (start, end), "date": (start, end) | None}]
cat.extract(ex.field("total"))                      # doc -> Quote(text, start, end, confidence)
cat.extract(ex.field("date"))

@cat.fn
def amount(total):                                  # extracted values are strings; parse them in ordinary functions
    return float(total.replace(",", ""))
```

For long documents (contracts), `solvi.extract_long.LongSpanExtractor` reads the whole text in overlapping 1024-token
windows, takes a field description instead of a fixed field list, and supports "no answer" with a per-field threshold.
See [docs/guide.md](docs/guide.md#extracting-fields-from-documents).

## Results

Same training documents for both sides. The baseline, Laya, is a ModernBERT-large model that answers the typed questions
directly, fine-tuned with its authors' recipe. Details and caveats: [docs/benchmarks.md](docs/benchmarks.md).

| Task (test set) | solvi | Baseline |
|---|---|---|
| SROIE receipts, 6 questions (361 receipts) | **97.8%** | 92.6% |
| SROIE, only 100 labeled training receipts | **95.9%** | 80.0% |
| CORD receipts, 4 questions (100 receipts) | **98.5%** | 96.0% |
| CORD, share of questions answered at >= 99% precision | **99.7%** | 14.2% |
| CUAD contracts, 5 questions (102 contracts, median 33k chars) | **94.8%** | 77.5% (sees first 1024 tokens only) |

- Calibrated confidence: ECE 0.008 (SROIE), 0.011 (CORD), 0.027 (CUAD).
- On the document benchmarks, 100% of answers are backed by a quote at stated offsets, or the system abstains.
- Speed: about 0.4 ms per decision when no model is involved; about 39 ms per receipt with the one-pass extractor on an
  A100 GPU.

## Speed

Strategist on random layered catalogs ([benchmarks/strategist_scale.py](benchmarks/strategist_scale.py), one CPU core):

| catalog parts | plan | plan + run + trace | parts run | share of catalog |
|---:|---:|---:|---:|---:|
| 100 | 0.2 ms | 1.3 ms | 31 | 31% |
| 1 000 | 0.7 ms | 2.6 ms | 115 | 12% |
| 10 000 | 6 ms | 10 ms | 236 | 2.4% |

Insurance claim desk with six slow services of 100-300 ms ([examples/09_strategy_at_scale.py](examples/09_strategy_at_scale.py)):

| questions asked | script computing everything | solvi, one by one | solvi, `workers=8` |
|---|---:|---:|---:|
| fast track? | 1 122 ms | 503 ms | 313 ms |
| full decision + payout | 1 122 ms | 773 ms | 461 ms |
| all five questions | 1 122 ms | 1 025 ms | 463 ms |
| expired policy (hard check settles it) | 1 122 ms | 152 ms | 153 ms |

A rule-only decision on a small catalog takes well under a millisecond; on documents the extractor dominates (about 39 ms per
receipt with the one-pass extractor on an A100).

## When to use it

- The answer is a computation or a rule over a few values found in a document or a record: receipts, invoices,
  contracts, requests, orders.
- You need to show why: auditors, compliance, or a human reviewing low-confidence cases.
- Some rules are non-negotiable (hard checks), and the rest can be learned from about 100 labeled examples.

## When not to use it

- Open-ended free-text questions or generated answers. solvi answers typed questions only: yes/no, choices, scores,
  multi-label, "not stated", exact spans of the text, rankings and number ranges.
- New fields with no labeled examples. Extracting a field from its description alone does not work yet (14% and 66% on
  two held-out fields); a universal extractor is coming.
- No labels at all. Plan on roughly 100 labeled documents (field positions) per task.
- CPU-only deployment with a quantized model: the int8 ONNX extractor loses up to 12 points on amounts, company
  names and addresses. fp32 on CPU keeps accuracy but takes about 0.7 s per receipt on 2 cores.

## Examples

| File | What it shows |
|---|---|
| [examples/01_leave_request.py](examples/01_leave_request.py) | Leave request from a plain dict: rules, hard checks, parts the strategist skips |
| [examples/02_shop_order.py](examples/02_shop_order.py) | Shop order: two rule-based questions plus a "suspicious?" question learned from history with `fit` |
| [examples/03_invoices.py](examples/03_invoices.py) | Invoice approval from text: regex extractors with quotes, four questions, one learned |
| [examples/04_refunds.py](examples/04_refunds.py) | Refund e-mails: yes/no learned from examples under a hard "within 30 days" check |
| [examples/05_tic_tac_toe.py](examples/05_tic_tac_toe.py) | Tic-tac-toe agent: each move is an answer with its reason (win, block, fork, ...); a hard check rejects invalid boards |
| [examples/06_learned_rules.py](examples/06_learned_rules.py) | `learn_rule`: route parcels to delivery zones from free-form addresses with a readable if-then list learned from labels |
| [examples/09_strategy_at_scale.py](examples/09_strategy_at_scale.py) | Insurance claim desk: the strategist generates a different plan per question set, hard checks first with early exit, slow services in parallel; timed |
| [examples/10_learn_in_milliseconds.py](examples/10_learn_in_milliseconds.py) | `fit_fast`: a new question learned in milliseconds, then corrected one example at a time (each correction ~0.2 ms, nothing retrained) |
| [examples/11_answer_types_and_constraints.py](examples/11_answer_types_and_constraints.py) | Multi-label and ordinal answers tied by constraints between answers; contradictions in learned answers are repaired by joint decoding |
| [examples/12_grounded_audit.py](examples/12_grounded_audit.py) | One catalog with and without models: provenance, `res.audit()`, a hallucinated quote caught by grounding, a decision outside its options, a model changed since the decision, lifetime safeguard stats |
| [examples/13_decide_model.py](examples/13_decide_model.py) | Support-email routing by a decider model as a catalog part: bias correction on unlabelled emails, 16 labelled examples, abstention, a constraint with a rule-based question, a hard check, the audit, `teach`, escalation for a target error rate, a JSON ticket (the real model with `SOLVI_DECIDE_MODEL`, a stand-in otherwise) |
| [examples/14_typed_catalog.py](examples/14_typed_catalog.py) | Typed facts: a pydantic request, type hints as fact types, answer types from the rules' return types (Enum, Literal, bool), a mismatch caught at registration, rejected values → fallback / abstention, a response as JSON that loads back and replays |
| [examples/15_typed_decisions.py](examples/15_typed_decisions.py) | Typed decisions: a pydantic ticket, the questions as a pydantic model's fields (choice, ordinal score, yes/no, multi-label), four answers from one forward pass, a hard check, a constraint and a rule over the model, an escalation in the audit and stats (the real model with `SOLVI_DECIDE_MODEL`, a stand-in otherwise) |
| [examples/16_primitives.py](examples/16_primitives.py) | Answer primitives: "not stated" vs abstain, spans parsed into numbers, evidence quotes checked in the text (`require_evidence`), a ranking with scores, an estimate with an interval — from rules and from a decider with the L14g contract; confidence per kind, JSON round trip, replay |
| [examples/17_model_strategist.py](examples/17_model_strategist.py) | The code strategist: dead ends dropped, the cheapest verified plan by declared costs, a model's proposal checked and rejected; aliases for names that match no fact (experimental; stand-ins without weights) |
| [examples/18_several_models.py](examples/18_several_models.py) | Several models, one decision: a cascade small → large, a vote of two model families, a route by code — each under one `act_guard` guarantee, with cost per question; every stage in the audit and the trace |
| [examples/07_receipts_model.py](examples/07_receipts_model.py) | Expense check on a scanned receipt: a receipts-tuned extractor cites each field, rules and a hard check decide (needs `solvi[model]`) |
| [examples/08_contracts_by_description.py](examples/08_contracts_by_description.py) | Contract review with fields defined only in words: the general extractor reads the whole contract, cites clauses or says "absent" (needs `solvi[model]`) |

Run them from a clone: `python examples/01_leave_request.py`.

## More

- [docs/guide.md](docs/guide.md): full API walkthrough.
- [docs/decide_format.md](docs/decide_format.md): the decider checkpoint contract (for training your own).
- [docs/strategist.md](docs/strategist.md): the code strategist and the experimental model strategist and name matching.
- [ROADMAP.md](ROADMAP.md), [CHANGELOG.md](CHANGELOG.md), [CONTRIBUTING.md](CONTRIBUTING.md), [SECURITY.md](SECURITY.md).
- [docs/benchmarks.md](docs/benchmarks.md): setups, per-question numbers, caveats.
- [benchmarks/](benchmarks/): dataset loaders and benchmark scripts (SROIE, CORD, CUAD, Kleister-NDA).
- Tests: `pytest`.

## License

Apache-2.0. See [LICENSE](LICENSE).

## Citation

A paper is in preparation. Until then, please cite the repository:

```bibtex
@software{solvi,
  title  = {solvi: verifiable decision systems from catalogs of functions and checks},
  author = {mxkuzn and solvi contributors},
  year   = {2026},
  url    = {https://github.com/solvi-ai/solvi}
}
```
