Metadata-Version: 2.5
Name: judgements
Version: 0.1.0
Summary: TypeSafe System One for pydantic users: declare questions as fields of a model, get a model back.
Author: Nagarjuna Kumarappan
License-Expression: MIT
License-File: LICENSE
Keywords: classification,judgements,probabilities,pydantic,system-one,typesafe
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Software Development :: Libraries
Requires-Python: >=3.11
Requires-Dist: msgspec>=0.18
Requires-Dist: pydantic>=2
Requires-Dist: typesafe-sdk<0.7,>=0.6
Requires-Dist: typing-extensions>=4
Description-Content-Type: text/markdown

# judgements

TypeSafe's System One model answers narrow, typed questions about a piece of state and returns
calibrated probabilities instead of text. `judgements` wraps that API so it looks like pydantic:
you declare the questions as fields of a model, and you get a model back.

`hello.py` is the whole idea in forty lines. This README explains what each line does.

## Install

From this directory:

```
uv sync
```

This installs `judgements` in editable mode. Plain `pip install -e .` works too if you prefer your
own environment.

## Run hello.py

Put your key in the environment:

```
export TYPESAFE_API_KEY=...
```

Then:

```
python hello.py
```

```
Triage(billing=True p=0.98, tone=frustrated p=1.00, urgency=today p=0.98)
route to billing
```

## The model is the questions

```python
class Triage(Judgements):
    billing: bool = question("Is this ticket about billing?")
    tone: Tone = question("What is the customer's tone in `body`?")
    urgency: Urgency = score("How urgent is this ticket?")

triage = Triage.ask(ticket)
```

Each field is one question. The annotation is the answer type, the marker holds the instructions,
and `Triage.ask(state)` sends every field in a single request and returns a `Triage` whose fields
hold the plain answers.

There are three kinds of question, and the annotation picks the kind:

| Annotation | Kind | What comes back |
| --- | --- | --- |
| `bool` | noul, "does this hold?" | `True` or `False` |
| an `Enum` or `Literal["a", "b"]` | choice, "which one?" | the chosen member |
| an `Enum` with `score(...)` | score, "how much, on this scale?" | the most probable level |

`question(...)` infers noul or choice from the annotation. Nothing in a type says "ordered rubric",
so a score is always declared with `score(...)`. `noul(...)` and `choice(...)` exist when you want
to be explicit, and `noul` takes options: `true=` and `false=` describe the two outcomes, and
`threshold=` sets where the `bool` flips, 0.5 by default.

For an Enum, member names are the labels the model chooses between and member values are their
descriptions. Write the descriptions as the concrete situations you mean:

```python
class Urgency(Enum):
    can_wait = "No deadline is implied; handle in the normal queue"
    this_week = "The customer expects a resolution within a few days"
    today = "The customer is blocked or demands immediate action"
```

For a score the order matters: first member is level 0, the lowest. For a `Literal`, the strings
are undescribed labels in a choice and the level descriptions themselves in a score.

## Reading the answer

The fields are plain values, so `triage.tone == Tone.angry` and `if triage.billing:` just work.
The probabilities behind them are one attribute away:

```python
triage.p.billing            # 0.98, probability that the answer is yes
triage.p.tone               # {Tone.calm: 0.0, Tone.frustrated: 1.0, Tone.angry: 0.0}
triage.confidence.tone      # 1.0, the model's reported confidence in the chosen option
triage.expected.urgency     # 1.97, see below
triage.results["urgency"]   # ScoreResult(can_wait 0.01, this_week 0.01, today 0.98; expected 1.97)
triage.usage                # Usage(requests=1, input_tokens=312, output_tokens=48)
```

`p` is short for probability: one number for a `bool` field, a distribution for the others.

`expected` exists for score fields only. Levels have positions, `can_wait` is 0, `this_week` is 1,
`today` is 2, and `expected` is the probability-weighted average of those positions. With the
probabilities above that is 0 × 0.01 + 1 × 0.01 + 2 × 0.98 = 1.97. It falls between levels and is
the number to use when averaging or ranking many items. The field itself holds the single most
probable level.

Use the probabilities to make policy explicit rather than trusting the top answer:

```python
if triage.billing and triage.p.billing > 0.9:
    route_to_billing()
elif triage.confidence.tone < 0.6:
    escalate_to_human()
```

`triage.model_dump()` gives `{'billing': True, 'tone': 'frustrated', 'urgency': 'today'}`, with
Enum names rather than descriptions.

Because the answers and their probabilities share one object, a few field names are reserved:
`p`, `confidence`, `expected`, `results`, `usage`, `questions`, `ask` and `from_answers`. Using one
raises a `TypeError` at class definition.

## Clients

`Triage.ask(ticket)` uses a default client that reads `TYPESAFE_API_KEY`. For anything beyond a
script, make a client:

```python
ts = TypeSafe()                                   # or AsyncTypeSafe(), then `await ts.ask(...)`

triage = ts.ask(ticket, Triage)
triage, refund = ts.ask(ticket, Triage, wants_refund)      # several things, one request
triages = ts.map(tickets, Triage)                          # one request per ticket, in order
ts.usage                                                   # requests and tokens so far
```

`ask` accepts a pydantic model, a dict, a list or a string as state. It is sent as JSON exactly as
it is, so backticked paths in instructions, like `` `body` ``, are relative to the state itself.
Put related material together in one object when a judgement needs to compare parts of it.

`ask` also takes `model=`, `retry=` and `timeout=` for one call. The async `map` takes
`concurrency=`, eight by default.

## Questions on their own

A question does not need a model. On its own it returns the full result object:

```python
wants_refund = question("Does the customer explicitly ask for money back?", bool)
r = ts.ask(ticket, wants_refund)      # NoulResult(no, p=0.08)
bool(r), r.probability
```

Options can be decided per request, which is how you rerank or select among candidates:

```python
best = choice("Which of `candidates` best answers `query`?", candidates)   # a list of labels
r = ts.ask({"query": query, "candidates": candidates}, best)
r.choice, r.ranked                    # the winner, and every candidate by probability

relevance = score("How relevant is `text` to `query`?", {"none": "Off topic", "partial": "Related", "direct": "Answers it"})
r = ts.ask({"query": query, "text": text}, relevance)
r.level, r.score, r.at_least("partial")
```

A dict gives each label a description. A list gives labels only.

## Writing good questions

- Ask one narrow judgement per field. Split independent dimensions into separate fields; they are
  answered in parallel in the same request at no extra latency.
- Put the judgement in the instructions and the possible answers in the type. The field name is for
  your code and is not shown to the model.
- Include a way out when nothing may fit, such as an `unclear` or `other` member.
- Check the exact request before spending tokens:

```python
request(ticket, Triage)               # {"state": {...}, "questions": {"Triage.billing": {...}, ...}}
```

## Testing without a key

```python
from judgements.testing import FakeTypeSafe

fake = FakeTypeSafe(billing=0.9, tone=Tone.angry, urgency={Urgency.today: 0.7, Urgency.this_week: 0.3})
triage = fake.ask(ticket, Triage)     # same parsing path as the real client, no network
fake.requests[0]["state"]             # what would have been sent
```

Answers are matched by field name. A `bool` field takes a probability or a bool; a choice or score
field takes the chosen option or a dict of option to probability. A missing answer raises. There is
an `AsyncFakeTypeSafe` too, and `tests/test_judgements.py` shows both in use:

```
python -m unittest discover -s tests
```

## Further reading

The live docs are the reference for the model itself: [System One](https://docs.typesafe.ai/concepts/system-one),
[state](https://docs.typesafe.ai/concepts/state), [the three primitives](https://docs.typesafe.ai/primitives)
and [confidence](https://docs.typesafe.ai/confidence).
