Metadata-Version: 2.4
Name: strands-decider
Version: 0.1.0
Summary: An experimental typed classifier: calibrated yes/no, choice and score answers from an LLM torso with a pointer readout.
License-Expression: Apache-2.0
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: torch>=2.7
Requires-Dist: transformers<6,>=5.15
Requires-Dist: peft>=0.21
Requires-Dist: accelerate>=0.34
Requires-Dist: huggingface_hub>=1.5
Requires-Dist: pydantic>=2.7
Requires-Dist: typer>=0.12
Requires-Dist: rich>=13.7
Requires-Dist: pyyaml>=6.0
Requires-Dist: fastapi>=0.115
Requires-Dist: uvicorn[standard]>=0.30
Requires-Dist: numpy>=1.24
Requires-Dist: safetensors>=0.4
Provides-Extra: train
Requires-Dist: datasets>=2.20; extra == "train"
Provides-Extra: dev
Requires-Dist: pytest>=8.0; extra == "dev"
Requires-Dist: ruff>=0.6; extra == "dev"
Requires-Dist: mypy<3,>=1.15; extra == "dev"
Requires-Dist: pre-commit>=3.7; extra == "dev"
Requires-Dist: commitizen>=3.28; extra == "dev"
Dynamic: license-file

<div align="center">
  <div>
    <a href="https://strandsagents.com">
      <picture>
        <source media="(prefers-color-scheme: dark)" srcset="https://strandsagents.com/latest/assets/wordmark-github-dark.svg">
        <img src="https://strandsagents.com/latest/assets/wordmark-github-light.svg" alt="Strands" width="320">
      </picture>
    </a>
  </div>

  <h1>Strands Decider</h1>

  <h2>A small, fast decision model for agentic AI.</h2>

  <div align="center">
    <a href="https://github.com/strands-labs/strands-decider/graphs/commit-activity"><img src="https://img.shields.io/github/commit-activity/m/strands-labs/strands-decider" alt="Commit Activity"></a>
    <a href="https://github.com/strands-labs/strands-decider/issues"><img src="https://img.shields.io/github/issues/strands-labs/strands-decider" alt="Open Issues"></a>
    <a href="https://github.com/strands-labs/strands-decider/pulls"><img src="https://img.shields.io/github/issues-pr/strands-labs/strands-decider" alt="Open PRs"></a>
  </div>

</div>

<hr>

Strands decider is one of a new class of **decision models**, or "system one" models. Unlike an
LLM, which can generate arbitrary text, a decision model picks between sets of options and
rates things on a scale. This class of model works best for problems that fall between LLMs
and traditional classification models: it is a general-purpose classification and scoring
model that responds faster than an LLM and does not take the time and expertise to train
that a traditional classifier does. That makes it a natural fit for the decisions inside
agentic workflows built with the [Strands Agents SDK](https://github.com/strands-agents/sdk-python).

- **Pick from options** — "Are there three r's in strawberry? Yes or no." "What language is
  the phrase 'sihamba ngokushesha' in? English, Zulu, or Dutch."
- **Rate on a scale** — "How positive is 'this is the best doc I've ever read'? Between 0 and 1."
- **Calibrated reliability scores** — every decision carries a confidence. On short
  classification tasks it has never seen, answers at a confidence of 0.9 or more are right
  about 95% of the time ([evaluation/results.md](evaluation/results.md#summary)); below that,
  confirm or ask a person. Frontier LLM inference APIs do not expose anything equivalent.

## Getting Started

The easiest place to get started is through the `strands-decider` cli:

```bash
pip install strands-decider
```
### Choice question
You can ask the model to choose based on some state and a question:
```bash
strands-decider ask StrandsAgents/strands-decider-2B-hobson-v19 \
  --state "Help! My payouts have been failing for 3 days! " \
  --choice "Which team should handle this?=billing,sales,retail" 
```
Example Output:

```bash
choice_0 -> billing (confidence 0.768)
  billing                  0.845
  retail                   0.091
  sales                    0.064
```

### Noul question
You can also ask the model a Yes/No question:

```bash
strands-decider ask StrandsAgents/strands-decider-2B-hobson-v19 \
  --state "Help! My payouts have been failing for 3 days! " \
  --noul "Does this convey urgency?" 
```

Example Output:

```bash
noul_0 noul = 0.828
```
Closer to 1 is leaning more toward Yes

### Score question
Or give a question a score:
```bash
strands-decider ask StrandsAgents/strands-decider-2B-hobson-v19 \
  --state "Help! My payouts have been failing for 3 days! " \
  --score "How frustrated is the writer?=calm,frustrated,depressed"
```
Example Output:

```bash
score_0 score = 1.10 (confidence 0.518)
  0: calm                                     0.163
  1: frustrated                               0.573
  2: depressed                                0.265
```

### Multiple question types

You can combine multiple questions into a single command. This is more efficient as you only need to load the state for the model once:

```bash
strands-decider ask StrandsAgents/strands-decider-2B-hobson-v19 \
  --state "Help! My payouts have been failing for 3 days! " \
  --choice "Which team should handle this?=billing,sales,retail" \
  --noul "Does this convey urgency?" \
  --score "How frustrated is the writer?=calm,frustrated,depressed"
```
<details>
  <summary>Example Output</summary>

  ```bash
  noul_0 noul = 0.829
  choice_0 -> billing (confidence 0.769)
    billing                  0.846
    retail                   0.090
    sales                    0.064
  score_0 score = 1.10 (confidence 0.519)
    0: calm                                     0.163
    1: frustrated                               0.574
    2: depressed                                0.263
  ```
</details>

### Running as a server

You can also run the model as a server, and ask questions via http requests:
```bash
strands-decider serve StrandsAgents/strands-decider-2B-hobson-v19 --port 8000
```

```bash
curl -s localhost:8000/v1/systemone \
  -H 'content-type: application/json' \
  -d '{
    "state": "Help! My payouts have been failing for 3 days!",
    "questions": {
      "is_urgent": {"type": "noul", "instructions": "Does this convey urgency?"}
    }
  }'
```

<details>
  <summary>Example Output</summary>

  ```bash
  {
    "model": "strands-decider-2B-hobson-v19",
    "answers": {
      "is_urgent": {
        "type": "noul",
        "noul": 0.8277
      }
    },
    "usage": {
      "input_tokens": 86,
      "output_tokens": 1
    },
    "latency_ms": 140.03
  }          
  ```
</details>

## About the model

`strands-decider-2B`, the first model of the family, has 1.9 billion parameters. It answers a
question in a median 115 ms on an RTX 3090 and also serves on an Apple-silicon Mac or on CPU,
and many questions about one text are cheap, because the text is read once and each question
adds only its own tokens
([docs/inference.md](docs/inference.md#asking-many-questions-is-nearly-free)). Its
accuracy and calibration are measured on JevBench, a third-party benchmark for this class of
model ([Performance](#performance)).


## Model Architecture

The core idea: take a pretrained decoder LLM torso ([Qwen3.5-2B-Base](https://huggingface.co/Qwen)),
**discard its language-modelling head** — taking away its ability to generate text — and
replace it with a small **pointer head** of about a million parameters. The head scores each
option by comparing the hidden state at the `<answer>` position against the hidden state at
that option's own last token. One forward pass, no generation, no decoding loop. The torso is
adapted with a rank-16 LoRA adapter, and the head runs in fp32.

Because the head holds no per-option parameters, nothing can learn that "the first option is
usually right", nothing caps how many options a question may carry, and label sets are
defined by the request rather than baked into the weights. Three question types come out of
the same masked softmax, read back differently: `noul` (yes/no), `choice` (one of N) and
`score` (an ordered rubric).

As you browse the research, you will find that this is the second major iteration of the
architecture. The first used a slot head, which mapped the final hidden state to a fixed set
of slots and performed significantly worse. Every change since is captured in the research
so you can follow along with the work. The reference model today is **v19**; v20 is the most
recent experiment and did not displace it ([The research](#the-research)).
[docs/architecture.md](docs/architecture.md#the-architecture) has the full design, the
training objective, and the decisions behind them.

## Performance

Three targets matter: **accuracy**, **calibration** and **latency**. v19 measures:

| Measurement | v19 | Source |
| --- | --- | --- |
| JevBench v1 public set, 231 tasks, accuracy | 0.723 (167 of 231) | [evaluation/README.md](evaluation/README.md) |
| JevBench Brier score / expected calibration error | 0.342 / 0.052 | [evaluation/README.md](evaluation/README.md) |
| Tiers (this repository's split of the public tasks): easy / standard / hard | 1.000 / 0.875 / 0.505 | [evaluation/jevbench.md](evaluation/jevbench.md#board-position-v142-25-september-2026) |
| Latency per JevBench question, RTX 3090 under WSL2, median / 95th percentile | 115 ms / 299 ms | [evaluation/results.md](evaluation/results.md#summary) |
| Latency per question, M3 Pro (Apple silicon), warm median, under 300 tokens / all tasks | 153 ms / 234 ms | [evaluation/results.md](evaluation/results.md#serving-on-a-mac-accuracy-and-latency) |

Every task in the easy tier is answered correctly, and the nearest comparison is
`decider-2b`, which shares v19's torso with a different recipe
([evaluation/jevbench.md](evaluation/jevbench.md#against-the-nearest-open-systems)).

The JevBench figures are at the 3072-token window the run was preregistered at; the recipe
saves a 4096-token window, at which v19 scores 168. And 231 tasks are few: six retrains of
the v17 recipe had a standard deviation of 3.2 tasks, so treat a difference under about 10
tasks between two single runs as unresolved. [evaluation/README.md](evaluation/README.md)
has the scripts and the limitations, and links the results by version and the board
position with its caveats.

## Why 2B?

Two reasons. First, **experimentation**: at 1.9 billion parameters you can serve the model,
and retrain the whole recipe, on hardware you already have — about 11 hours on one RTX 3090,
and serving works on an Apple-silicon Mac. That makes trying an idea fast and low-risk.
Second, ~2B parameters looks like a sweet spot: small enough to experiment with, large enough
to do meaningful work.

## What can I do with it?

The decisions inside an agentic workflow are the natural target:

- **Model routing** — pick the right LLM for a task
- **Tool selection** — decide which tool an agent should call next
- **Argument checking** — verify a tool call's arguments before it runs
- **Triage** — route an incoming request to the team or queue that should own it
- **Guardrails** — grounding checks, safety classification, policy classification
- **Evaluations** — score model outputs, and check whether an answer is adequate, at low cost
- **Hybrid agents** — let the LLM make the hard decisions and a decider make the rote ones,
  reducing cost and latency

## Trying it in an Agent

The repository includes a worked example of strands decider inside a Strands agent, under
[`examples/strands/`](examples/strands/README.md): a `before_tool_call` intervention that gates a
weather-tool call on two yes/no decisions, so the agent asks which city instead of guessing. See
the example's [README](examples/strands/README.md) for setup and the walk-through.

## Training

Two entry points, both on a Linux or WSL2 host with NVIDIA GPUs and the training
environment of [training/README.md](training/README.md#setup) (the `train` extra, the
pinned torch and `flash-linear-attention`):

- `training/recipe.sh all`: one host, local, 1 to 8 GPUs (`NGPU=8` for eight). It builds
  the corpora, trains, calibrates and evaluates.
- `training/run_recipe.sh all` with [training/aws/](training/aws/README.md): a distributed
  8-GPU host. It adds per-stage timing and logs, row-count checks, the S3 copy of the
  outputs (`PY` and `S3_PREFIX` must be set), and `FAST=1`, which trades 24 GB
  compatibility for speed on 80 GB GPUs.

About 11 hours on one RTX 3090 (24 GiB), or 1 hour 10 minutes on eight H100s with `FAST=1`.
[training/README.md](training/README.md) has the setup, the stages, the settings and the
hardware notes; [data/sources.md](data/sources.md) lists every source and its licence, and
[data/README.md](data/README.md#reproduction-contract) states the reproduction contract.

## The research

The record of the work is in the repository. Since v9, every training run states its
predictions and its failure conditions before training, and the outcome is appended after
the run without editing what came before. By that rule, a run that misses its bar does not
replace the reference model; five runs were promoted by the maintainers' decision anyway,
and the record says so. Most runs missed their bar, v20 among them: 169 of 231 at the
4096-token window against v19's 168, inside the retrain noise, with four predictions
failed. [research/README.md](research/README.md) lists each run and its outcome, and
[research/history.md](research/history.md) tells what moved the benchmark and what did
not. To propose an experiment, open an issue with a preregistration ([CONTRIBUTING.md](CONTRIBUTING.md)).

## License

This project is licensed under the Apache License 2.0 - see the [LICENSE](LICENSE) file for details.

## Security

See [CONTRIBUTING](CONTRIBUTING.md#security-issue-notifications) for more information.
