Metadata-Version: 2.5
Name: hunch-jev
Version: 0.7.0
Summary: Plain function verbs on Jev: classify, score, check, pick, rank, generate.
Project-URL: Homepage, https://github.com/steven-shoemaker/hunch
Project-URL: Repository, https://github.com/steven-shoemaker/hunch
Author: Steven Shoemaker
License: MIT
License-File: LICENSE
Requires-Python: >=3.10
Requires-Dist: pydantic>=2.0
Requires-Dist: tqdm>=4.66
Requires-Dist: typesafe-sdk>=0.7.0
Provides-Extra: dev
Requires-Dist: pandas>=2.0; extra == 'dev'
Requires-Dist: pytest>=8.0; extra == 'dev'
Description-Content-Type: text/markdown

# hunch

![hunch — lists in, lists out](https://raw.githubusercontent.com/steven-shoemaker/hunch/main/docs/banner.png)

Plain functions on Jev. Lists in, lists out.

[Jev](https://docs.typesafe.ai) is TypeSafe's System One model. You send it some state and a typed question, and it sends back a label, a score, or a yes/no probability instead of a paragraph. I think it's the most useful thing to happen to "AI in a for loop" in a while. Calling it raw is fiddly, though: build a state object, build a question object, dig the answer out of the response. hunch is the version I wanted, where each of those is one function call and you can hand it a list or a pandas column instead of one thing at a time.

The rule the whole library follows: Jev decides, your code owns the workflow, and if an LLM is involved at all it only gets to propose candidates.

```python
import hunch

hunch.classify("This product is amazing!", ["positive", "negative", "neutral"])
# 'positive'

hunch.classify(df["JOB_TITLE"], ["Sales", "Engineering", "Marketing"])
# Series of labels, same index

hunch.score("Critical system failure", ["cosmetic", "degraded, workaround exists", "down for everyone"])
# 1.87  (position on the scale, 0 .. n-1)

hunch.check("BUY NOW!!!", "is unsolicited advertising")
# True

drafts = hunch.generate(str, n=20, instructions="tweets introducing hunch")   # an LLM writes
hunch.pick(drafts, "most likely to make a Python developer install it")      # Jev chooses
# 'Most of my "AI" code was a for loop around a prompt and a JSON parser. ...'
```

## Install

```sh
pip install hunch-jev
```

Python 3.10+. Set `TYPESAFE_API_KEY` in your environment, or call `hunch.configure(api_key=...)` at startup. Keys don't belong in source files.

## Where it fits

Anywhere a person is reading rows and making a call. A few that come up:

**Engineering: ticket triage, and a verifier step for a code review agent.**

```python
tickets = tickets.join(tickets.hunch.ask({
    "kind":     Classify(["bug", "feature request", "question"]),
    "severity": Rate(["cosmetic", "degraded, workaround exists", "blocked", "outage"]),
}))

# Keep only the review-bot comments Jev agrees describe a real defect in the diff
real = comments.hunch.where("describes an actual defect present in the diff",
                            columns=["comment"], context={"diff": diff}, threshold=0.7)
```

**GTM: ICP fit.** Pass the ideal customer profile as context and let Jev read every row against it.

```python
ICP = "B2B SaaS, 200 to 2,000 employees, sells to mid-market, has a RevOps or sales ops function, US or UK."

prospects = prospects.join(prospects.hunch.ask({
    "fit":   Rate(["not our buyer", "partial fit", "good fit", "textbook ICP"], "How well does this account match the ICP in context?"),
    "buyer": Check("this person could sign or sponsor a purchase for their team"),
}, context={"icp": ICP}))

outreach = prospects[prospects.buyer].nlargest(50, "fit")
```

**SEO: intent and thin content, then a title tag Jev picks from LLM drafts.**

```python
pages = pages.join(pages.hunch.ask({
    "intent": Classify(["informational", "commercial", "transactional", "navigational"]),
    "thin":   Check("the page is thin content that adds nothing over the top results for its query"),
}))
titles = hunch.generate(str, n=10, instructions="title tags for this page", context=page)
best   = hunch.pick(titles, "most likely to earn the click for the target query", context=page)
```

**Finance: anomalies and categorization.**

```python
suspect = txns.hunch.where("looks like a duplicate or erroneous charge",
                           columns=["merchant", "amount", "date", "memo"])
txns["account"] = txns.hunch.classify(GLAccount, columns=["merchant", "memo"])   # your Enum of GL accounts
```

## Verbs

| Verb | Jev primitive | Returns |
| --- | --- | --- |
| `classify(data, labels, multi_label=False, instructions=None)` | Choice, or one Noul per label | label, Enum member, or list of labels |
| `score(data, levels, instructions=None)` | Score | float position on the scale, or `{dim: float}` |
| `check(data, statement, criteria=None, threshold=0.5)` | Noul | bool, or `{name: bool}` |
| `pick(candidates, instructions)` | Choice over the candidates | the winning candidate |
| `rank(candidates, dimensions, levels, weights=None)` | Score per dimension | `Ranked` rows, best first |
| `generate(target, n=1, instructions=None)` | your LLM, validated by pydantic | `target` or `list[target]` |
| `ask(data, {name: Classify(...) \| Rate(...) \| Check(...)})` | all of the above, one request per item | dict per item, or a DataFrame for a Series |
| `where(data, statement, columns=None, threshold=0.5)` | Noul per row, then filter | the rows that match, strongest first |

Hand any verb one item and you get one answer back. Hand it a list, a tuple, or a pandas Series and you get the same container back, same length, same index. Hand it a DataFrame and each row is the thing being judged, so Jev sees every column, and the answers come back on the frame's index ready to `join`. Pass `columns=` to any verb to limit which columns of a DataFrame Jev reads; the answers still line up with the whole frame. Duplicate values are only asked once, and the distinct ones run in parallel across `max_workers` threads. All of them take `context=` for extra state that should ride along with the input, and `client=` if you don't want the default. Each has an `_async` twin that runs its requests on your event loop through the async SDK client, up to `max_concurrency` at a time, which is the one to use inside a service.

`labels` can be a plain list, an `Enum` class (you get members back, not strings), or a dict of label to description when the names alone are ambiguous. On `score` and `check`, `instructions` can be a dict of name to question. Those go out as one request per item and you get a dict back per item, which is how you score five dimensions without five round trips.

### Semantic WHERE

`where` is the filter you wish SQL had. It keeps the rows for which a statement holds and returns them strongest match first. On a DataFrame, Jev reads every column unless you pass `columns=`; the whole row comes back either way.

```python
df.hunch.where("probably likes cats")
df.hunch.where("is a decision-maker at a company that sells to enterprises", columns=["title", "company"], threshold=0.7)
```

`detail=True` returns all rows with `match` and `match_p` columns so you can draw your own line. Statements about evidence in the row filter well. Predictions about behavior cluster near 0.3 to 0.4 when the row says nothing either way, so rank those instead of thresholding them:

```python
cats  = df.hunch.where("probably likes cats", columns=["name", "age", "bio"], threshold=0.7)
scary = df.hunch.where("might yell at a waiter for getting their order wrong",
                       columns=["name", "age", "bio"], detail=True).nlargest(5, "match_p")
```

### `df.hunch`

Every verb is also on a `.hunch` accessor for DataFrames and Series, so it reads left to right in a notebook:

```python
df["review"].hunch.classify(["positive", "negative", "neutral"])
df.hunch.ask({"vibe": Classify([...]), "red_flag": Check("...")})
df.hunch.rank({"hook": "...", "clarity": "..."}, levels=[...])
best = df.hunch.pick("the best first date for the person in context", context={"looking_for": ME})
```

### Several questions, one request

When you want more than one thing about the same data, `ask` sends every question in a single Jev request per item. Each question is a small spec with the same arguments as its verb. A Series comes back as a DataFrame on the same index, so it joins straight onto your frame.

```python
from hunch import ask, Classify, Rate, Check

answers = ask(prospects["JOB_TITLE"], {
    "function": Classify(functions, FUNCTION_INSTRUCTIONS),
    "seniority": Classify(seniorities, SENIORITY_INSTRUCTIONS),
    "urgent": Check("this person should be contacted this week"),
    "fit": Rate(["poor", "okay", "strong"], "How well does this title fit an enterprise sales motion?"),
})
prospects = prospects.join(answers)
```

Context on the call rides along with every question. When only one question should see something, put it on that question instead: `Rate([...], "How compatible is this profile with the person in context?", context={"looking_for": ME})`. Questions whose context differs can't share a request, so `ask` groups them and sends one request per group.

With a Series or DataFrame, `detail=True` spreads each answer into columns instead of handing you objects: `fit`, `fit_level`, `fit_confidence`, `fit_shape` for a score; `label`, `label_p`, `label_confidence`, `label_shape` for a classify; `check`, `check_p` for a check. No lambdas to unpack anything.

`pick` on a Series or DataFrame returns the winner's index label, so `df.loc[best]` is the row. `rank` returns a DataFrame with `composite` and one column per dimension, sorted best first, on the same index.

### `detail=True`

The bare return is the answer. `detail=True` returns the whole distribution:

| Verb | Detail type | Fields |
| --- | --- | --- |
| `classify` | `Answer` | `.label .p .probabilities .confidence .shape .top2 .on()` |
| `classify(multi_label=True)` | `MultiAnswer` | `.labels .probabilities .threshold` |
| `score` | `Rating` | `.score .level .normalized .probabilities .legend .confidence .shape .on()` |
| `check` | `Feeling` | `.p .threshold`, truthy at threshold |
| `pick` | `Pick` | `.winner .ranked .confidence .shape .on()` |

`.shape` is a judgment about the distribution, and the cutoffs are yours, not Jev's:

| Shape | Meaning |
| --- | --- |
| `sure` | One option dominates |
| `split` | Two options are close |
| `unsure` | Flat or weak evidence |

The two common policies are arguments on `classify` (and on `Classify(...)` inside `ask`):

```python
seniority = hunch.classify(df["title"], ["IC", "Manager", "Director"], split="rematch", unsure="review")
```

`split="rematch"` re-asks between the top two labels, only for the rows that were split, batched and cached like everything else. Any other value is used as the label for those rows. `unsure="review"` does the same for flat distributions. Both default to keeping the first answer. For anything more custom, `detail=True` gives you the `Answer` and `.on(sure=, split=, unsure=)` branches on it; pass a callable for a branch that costs a call.

Cutoffs live on `ShapePolicy` and work on the probabilities alone: `sure_peak`, `unsure_peak`, `split_margin`, `split_mass`. Jev's `confidence` is derived from the top probability, so it carries no extra information and the policy ignores it. Neither says whether the label is correct. A confidently wrong answer is still confident, which is why `evaluate` below exists. Changing the policy never re-runs inference, because the cache stores the raw distribution and the shape is computed on the way out.

## Check it before you trust it

Label 50 to 100 rows by hand, then measure:

```python
pred = hunch.classify(sample["title"], LEVELS, detail=True)
hunch.evaluate(pred, sample["true_level"])
# Evaluation(accuracy=91.0% on 100 rows, by shape: sure: 98% of 71, split: 79% of 19, unsure: 60% of 10)
```

`by_shape` tells you whether "sure" really means right on your data, and so which rows to send for review. `.errors` lists every miss, and `.table()` gives the confusion matrix.

For `check` and `where`, pick the cutoff from data instead of by feel:

```python
p = hunch.check(sample, "is an economic buyer", detail=True)
cut = hunch.tune_threshold(p, sample["is_buyer"], precision=0.9)
# Threshold(threshold=0.71, precision=0.92, recall=0.64, ...)
buyers = df.hunch.where("is an economic buyer", threshold=cut.threshold)
```

`precision=` gives the lowest cutoff that keeps that share of matches correct. `recall=` gives the highest cutoff that still catches that share of true matches. With neither, it maximizes F1.

## Generate, rank, pick

This is the part where an LLM is allowed in the room. It writes the candidates. Jev scores them and picks. The weights stay in your code, so re-ranking after you change your mind costs nothing.

```python
import hunch

hunch.configure(llm=hunch.openrouter(), cache="~/.cache/hunch")

tweets = hunch.generate(str, n=20, instructions="Tweets introducing hunch to Python developers", context=README)

ranked = hunch.rank(
    tweets,
    {"hook": "How strong is the first line?", "clarity": "How clearly does it say what hunch does?", "specific": "How concrete, not generic, is it?"},
    levels=["weak", "okay", "strong", "excellent"],
    weights={"hook": 2, "clarity": 1, "specific": 1},
)
finalists = [row.item for row in ranked[:5]]

winner = hunch.pick(finalists, "the tweet most likely to make a Python developer install hunch")
```

`generate` accepts `str`, `int`, dataclasses, `TypedDict`s, pydantic models, `list[str]`, and any other type pydantic can validate. Large `n` is drawn in batches of 25 that avoid repeating earlier items, and the result is cached, so re-running a notebook cell returns the same items. Pass `fresh=True` for a new draw. `hunch.openai`, `hunch.cerebras`, and `hunch.openrouter` are OpenAI-compatible adapters; pass `llm=` on `configure()` or on `generate()`.

## Client

```python
jev = hunch.Client(
    api_key=..., model="jev-latest", cache="~/.cache/hunch",
    max_workers=8,          # threads for sync calls
    max_concurrency=64,     # requests in flight for _async calls
    max_rps=None,           # cap requests per second, e.g. 20
    errors="raise",         # or "skip": failed rows come back None, with a warning
    policy=ShapePolicy(...),
)
hunch.classify(x, labels, client=jev)
jev.usage   # calls, cache hits, tokens, model
```

`hunch.configure(...)` takes the same arguments and sets the default used when `client=` is omitted. `cache=` writes raw Jev answers to disk keyed by state and question, so re-running a script over the same data is free.

On a big column, `errors="skip"` means one bad request doesn't sink the other 49,999. Good answers are cached as they arrive, so running the same call again only re-sends the rows that failed. The SDK already retries 429s and 5xx with backoff before anything counts as failed.

To see what a call would cost before running it:

```python
with hunch.dry_run() as plan:
    df.hunch.ask({...})
plan   # Plan(requests=8214, questions=16428, items=50000)
```

Nothing is sent inside the block and nothing is cached. Verbs return placeholder answers so the rest of your code keeps running. Rematches from `split="rematch"` aren't counted, since they depend on real answers.

Big columns get a progress bar. Any call that needs 10 or more requests shows one, counting requests rather than rows, so it already reflects dedupe and cache hits. `generate` shows an elapsed timer while it waits on the LLM. `progress=True` forces it on, `progress=False` turns it off.

## Examples

Each is a single file with the data inline, so you can run it as-is. Three need only `TYPESAFE_API_KEY`. The ones that `generate` also want an OpenRouter key.

| File | Shows |
| --- | --- |
| [`find_angry_reviews.py`](examples/find_angry_reviews.py) | `check` over a column as a boolean mask, ranking by probability, several checks in one request |
| [`dating_profiles.py`](examples/dating_profiles.py) | `generate` typed profiles, `ask` two questions per row, `score` against a described person, `pick` a date, then semantic `where` for cat people and waiter-yellers |
| [`classify_job_titles.py`](examples/classify_job_titles.py) | `classify` with Enums, `detail=True`, and `.on()` routing sure / split / unsure with a rematch |
| [`triage_tickets.py`](examples/triage_tickets.py) | `score` on two scales in one request, multi-label `classify`, paging policy kept in code |
| [`introduce_hunch.py`](examples/introduce_hunch.py) | `generate` 20 tweets with an LLM, `rank` them on weighted dimensions, `pick` the winner |
| [`organize_downloads.py`](examples/organize_downloads.py) | An LLM proposes a folder taxonomy, `classify` assigns every file, the script moves them. `--dry-run` prints the plan |

## What this is not

Jev does not invent labels. Whatever you pass as `labels` is the entire set of allowed answers, and that constraint is the point. `generate` is the one place invention happens, and it has no tools and takes no actions. If you want open-ended writing or a multi-step agent, this is the wrong library, on purpose.

## License

MIT. Jev and TypeSafe are [typesafe.ai](https://typesafe.ai); this library is not affiliated.
