thunc()
Guide

Caching, tracing and profiling

Save answers on disk so the model is asked once per input, clear them from Python or the command line, record every call to a file, see where a program's time goes, and replay recorded answers in tests.

Caching answers: cache=True

cache=True saves each answer on disk and reuses it when the same inputs come again, so the model is asked once.

@thunc.function(cache=True)
def category(ticket: str) -> Literal["bug", "billing", "other"]:
    """Classify this support ticket."""
    ...

It's off by default, because it only suits some functions:

When a saved answer is reused

Only for the exact same function, prompt, backend and model, so changing the docstring, the return type or the model asks again. A saved answer is checked against the return type and ensure= before it's reused, and failed calls are never saved. The function's name is part of the key, so renaming a function starts its cache fresh.

Where answers go

In .thunc_cache/ in the working directory; change it with configure(cache_dir=...) or THUNC_CACHE_DIR. Each call is one JSON file holding the full prompt in plain text, inputs included, so treat the folder like the data you send.

Clearing the cache

Clear everything, or one function's answers, from Python:

thunc.clear_cache()                                # everything
thunc.clear_cache(urgency)                         # one function
thunc.clear_cache("urgency")                       # the same, by name
thunc.clear_cache(older_than=timedelta(days=30))   # answers saved more than 30 days ago
thunc.cache_info()                                 # what's saved, one group per function

Or from the command line:

thunc cache list                                   # saved answers per function
thunc cache clear                                  # everything
thunc cache clear --function urgency               # one function (repeat for several)
thunc cache clear --older-than 30d --dry-run       # what would go, without deleting

A name is the function's name (urgency, or Triage.urgency for a method), optionally with its module (support_inbox.urgency). For thunc.call, pass name="..." to group its answers the same way; unnamed calls are cleared only with everything or by age. Ages count from when the answer was saved. clear_cache returns how many answers it deleted.

Clearing deletes only cache entries, never other files in the folder, and it's safe while another process is using the cache. The thunc command (also python -m thunc) reads THUNC_CACHE_DIR, or takes --cache-dir; it can't see a configure(cache_dir=...) in your code.

Tracing every call

thunc.configure(trace="calls.jsonl")

Every call is appended to the file as one JSON line, or set THUNC_TRACE instead. An agent run is one line too, with every model reply in it. log_triage.py and repo_guide.py read their traces back.

Profiling: thunc run --profile

Run your program through the thunc command to see where the time went when it ends:

thunc run --profile support_inbox.py --limit 20   # a script and its arguments
thunc run --profile -m myapp.triage               # a module, as with python -m

The report goes to stderr. Per function: the calls, cache hits, retries and failures, the total, mean, p95 and slowest time, and how much of it was the model and how much thunc's own work (building the prompt, parsing, the cache). Agent runs get their steps, model time and time in each tool. It also says what share of the program's wall time was spent in thunc, and how much calls overlapped under thunc.map.

stderr
CALLS
FUNCTION  CALLS  CACHED  RETRIES  FAILED  TOTAL   MEAN    P95    MAX  MODEL  LOCAL
urgency      11       1        1       0  1.70s  155ms  309ms  309ms  1.69s   12ms

In thunc:      774ms of 980ms wall time (79%); the rest was the program's own code
Model time:    1.69s, 99% of the time in calls (anthropic/default model 1.69s)
Concurrency:   calls overlapped 2.2x on average (thunc.map or threads)
Slowest:       urgency took 309ms

Without --profile, thunc run just runs the program and records nothing. The program's exit code is passed through.

To watch the same numbers while the program runs, with each call and agent step as it happens, use thunc watch.

Testing code that calls a model: record and replay

Record the answers once, commit them, and replay them in tests and CI with no model, backend or API key. Point thunc at a folder:

conftest.py
import thunc

thunc.configure(recordings="tests/recordings")   # or set THUNC_RECORDINGS=tests/recordings

Then fill it once, and replay from then on:

THUNC_RECORD=1 pytest              # ask the model, and save every answer in tests/recordings
pytest                             # replay: no model is asked, and a call that wasn't recorded fails
THUNC_RECORD=missing pytest        # replay what's there; ask for, and save, only what's missing

With a recordings folder set, every thunc.call, @thunc.function and agent run is answered from it. A call that isn't there raises ThuncError saying how to record it, instead of reaching a model, so CI never spends money or waits on a login. THUNC_RECORD=1 asks the model and saves (or replaces) every answer; THUNC_RECORD=missing keeps the answers the folder has and asks only for the rest. THUNC_RECORD without a folder is an error, not a silent no-op.

What's matched

A call's function name, system prompt and request (instructions, inputs and return type), so changing any of them needs a new recording. The backend and model are saved with each answer but not matched: CI needs no backend, and a recording made on one model replays under another, so re-record after changing models. The same call twice gets the same recorded answer.

A recorded answer is parsed into the return type and run through ensure= before it's used; one that no longer fits fails with a message saying to record it again. Failed calls are never recorded, only the answer that passed. Replaying never reads or writes the cache, and recording saves a cache hit's answer too.

Agent runs

A run is matched on its agent, task, request, tools and system=, and the model's replies are replayed in order. The tools run for real, so files are written and commands run as when it was recorded: give each test a fresh copy of its workspace. A run whose tools go differently can run out of replies, which fails it with a message saying to record it again. Memory and followed files aren't matched. Durable runs aren't recorded; while replaying, they fail instead of reaching a model.

The recordings folder

Recordings use the cache's format, one JSON file per call or run, holding the full prompt in plain text, so they're easy to review in a diff. thunc cache list --cache-dir tests/recordings lists them per function, and thunc cache clear --cache-dir tests/recordings --function urgency deletes one function's. Answers no test asks for any more stay until you delete them: to prune, delete the folder and record again. The trace marks replayed calls with "replayed": true (and "cached": true: no model was asked), and thunc watch shows them as replayed (from the next thunc-watch release).

Edit this page on GitHub