Metadata-Version: 2.4
Name: paperllm
Version: 0.1.0
Summary: Calling a local model about a paper, and keeping what the call cost: tool-submitted answers, truncation discipline, and a row per call.
Author: Bilal Zonjy, MD
License-Expression: MIT
Project-URL: Source, https://github.com/BilalZonjy/paperllm
Keywords: ollama,llm,pubmed,extraction,observability
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Science/Research
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Scientific/Engineering :: Bio-Informatics
Requires-Python: >=3.12
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: ollama<1,>=0.5.1
Requires-Dist: pydantic<3,>=2
Requires-Dist: sqlalchemy<3,>=2
Provides-Extra: migrations
Requires-Dist: alembic<2,>=1.13; extra == "migrations"
Provides-Extra: dev
Requires-Dist: pytest>=8.0; extra == "dev"
Dynamic: license-file

# paperllm

Calling a local model about a paper, and keeping what the call cost.

- **Call shapes that survive a small model** — an agentic call whose answer arrives through a tool
  rather than a `format=` grammar, a truncation check at every point a reply is returned, and a
  thinking-versus-content split that finds the answer where a reasoning model actually put it.
- **A row per call** — tokens in and out, the window and budget it ran under, why generation
  stopped, how long it took, and whether it worked.

Extracted from a SUDEP literature-review pipeline, where it runs a three-call extraction cascade over
~2,500 papers. **v0.1.0, and honestly 0.x**: one consumer so far, which is why the public surface is
deliberately small (see *What it does not do*). Expect the API to move before 1.0; pin a version.

## Two boundaries worth stating before you read further

**Ollama-shaped.** `done_reason`, the thinking channel, the retry semantics and the shape of a tool
round are Ollama's. The name says `llm`; the code says Ollama. A second backend would need more than
a new client object, and pretending otherwise here would be the kind of premature generality this
library was extracted specifically to avoid.

**Paper-scoped.** `llm_call.pmid` is a PubMed identifier, not a generic subject key. That is
deliberate — a sibling to `pubmedcorpus`, not a general-purpose LLM client.

## Installing

```bash
pip install paperllm
```

Take the extra if you intend to run the migration that creates the call log's tables:

```bash
pip install "paperllm[migrations]"
```

**No database driver is installed for you.** Nothing here is Postgres-specific, so which driver to
use is your decision — but read the `now()` note under *The recorder* before choosing a primary-key
strategy on Postgres.

## Using it

Everything the library needs to know about your deployment is one object:

```python
from paperllm.caller import Caller
from paperllm.config import CallConfig

caller = Caller(CallConfig(
    host="http://localhost:11434",
    # (model, purpose) -> context window. A judgement about *your* ensemble: this model's VRAM
    # ceiling, this kind of call's needs. The library only needs to be able to ask.
    num_ctx_for=lambda model, purpose: 32768,
    # Whether this model may emit reasoning tokens.
    think_for=lambda model: True,
    # Optional. Without it, nothing is written down and no database is needed.
    recorder=None,
))

answer = caller.extract(model, messages, MySchema, num_predict=16384)
```

`purpose` is one of `paperllm.config.EXTRACTION | PROSE | AGENTIC` — the three kinds of call this
library makes. It is **opaque to the library**: handed to `num_ctx_for`, stored on the call record,
never branched on. You decide what each is worth in context tokens.

### The calls

| | |
|---|---|
| `chat_raw` | one constrained-decoding call → raw content |
| `classify` / `extract` | the same, parsed and validated into a Pydantic model |
| `chat_prose` | unconstrained generation |
| `chat_agentic` | the model may call tools, you run them, it continues |
| `reason` | one agentic call whose answer arrives through a tool; returns `(scratchpad, was_truncated)` |

Two exceptions, and they are **siblings, not a hierarchy**: `ExtractionError` means the model
produced nothing usable; `Truncated` means generation stopped on `num_predict`. A caller that
swallows the first must decide about the second separately, because a budget bug filed as a
judgement about a paper is how a number ends up in the record as evidence.

`Truncated` carries both `content` and `thinking`. On a reasoning model the second is usually the
only one with anything in it — measured over one run, 31 of 48 tool-calling turns had empty
`content` — so a cut-off call is a shorter scratchpad rather than a lost paper.

### Budgets are yours

`num_predict` defaults to `MIN_NUM_PREDICT` (2048), which is a **floor, not a working budget**.
Thinking tokens count against the budget, so a number sized for the answer alone leaves a reasoning
model with nothing left to answer with — 2048 is what "you forgot to choose" looks like, small
enough to catch.

The right number depends on how your prompt was built and how much room it reserved for the reply.
Only you know both. The pipeline this came from keeps its measured budgets in
`sudep/analysis/budgets.py`; every one of them cost a failed run to learn.

## The call log

Two tables. `llm_call` is one row per call, append-only; `call_stage` is the vocabulary naming which
call in a cascade a row is.

```
pmid  model  stage  created_at        <- the key
purpose  label                        <- what kind of call, and which step
success  error_type  error_message
num_ctx  num_predict                  <- what it ran under
prompt_tokens  eval_tokens            <- what it used
done_reason  duration_ms
```

**`stage` names which call this was, not how the paper ended.** A paper that fails call 1 and
succeeds on call 2 has no single outcome, but each of its calls has one. The questions worth asking
are `stage × success`, and a pmid with no successful row at any stage is the failed paper. Seed
`call_stage` with your own cascade; the shipped rows describe a three-call one.

`success` means the call produced what its stage was asked for — a well-formed reply carrying
unusable JSON is a failure, because the cascade treats it as one and the log has to agree with the
cascade rather than with HTTP.

### The recorder, and the one trap

```python
def recorder(call):
    with your_own_session_scope() as session:   # NOT the caller's session
        paperllm.record.record(session, call)
```

**It must not join the transaction of the work it is describing.** A failed extraction rolls back,
and a record of the failure written in that transaction rolls back with it — leaving a log that
holds exactly the calls that went well. Open a short-lived session per record and commit it.

That also makes the primary key safe: Postgres's `now()` is the *transaction* timestamp, so one row
per transaction is what keeps `created_at` distinct. If you batch, use `clock_timestamp()`.

A recorder that raises does not break the call — `record.emit` swallows it and logs at WARNING. A run
that finishes with an incomplete log beats a run that died protecting its bookkeeping.

## It grows, and nothing prunes it

One row per call, and calls are never updated or replaced — a full pass over a few thousand papers
with a handful of models is tens of thousands of rows. **There is no retention policy, no TTL and no
cleanup job, deliberately.** A log that deletes on its own is a log you cannot trust to answer a
question about last month, and the only honest default for "how long is a call worth keeping" is
however long its answer stays interesting.

So deletion is an operator's decision, taken deliberately:

```sql
-- What you are about to remove, before removing it.
SELECT date_trunc('month', created_at) AS month, count(*)
FROM llm_call GROUP BY 1 ORDER BY 1;

DELETE FROM llm_call WHERE created_at < '2026-01-01';
```

`ix_llm_call_created_at` exists for exactly that scan.

**Age is the only axis available, and that is worth knowing before you rely on it.** The table
carries no schema or prompt version, so "delete everything from before the prompt changed" has to be
expressed as a date — look up when the version bumped and cut there. Storing the versions here was
rejected on the grounds that a call log should not have to be told what question the caller was
asking; if that turns out to be wrong, it is an added column and a migration, not a redesign.

Nothing downstream reads this table, so a delete cannot break a pipeline — only an analysis you had
not run yet.

## Migrations

The library owns its schema and ships an Alembic branch labelled `paperllm`, creating both tables and
seeding the vocabulary in one revision — an unseeded lookup table would make the foreign key reject
every insert. Add it to your `version_locations`, then `alembic upgrade heads` (**`heads`**, not
`head`: with more than one branch, `head` is ambiguous and errors).

`paperllm.db.Base` has its own `MetaData`, deliberately: nothing here has a foreign key crossing into
your schema — `llm_call.pmid` names a pmid and carries none, so a deleted paper does not take the
record of its failures with it.

## What it does not do

- **No CLI.** Deciding which environment variables must be present before touching a database is the
  application's call.
- **No sessions.** `record` takes one you opened.
- **No environment reading.** Everything arrives on `CallConfig`.
- **No re-exports** from `__init__.py`. Import submodules, so the public surface stays small enough
  to reshape once a second consumer shows where the joints actually belong.


Offline by construction: no Ollama, no database. `_ollama` is the single seam every call goes
through, and `CallConfig` is built in the test rather than patched onto a global — which is the point
of the window and thinking policies being callables.
