thunc()
Blog · Development log

Eight days of thunc

thunc went from an empty repository on 3 October 2026 to its twelfth release on 10 October. This is the whole story so far: what was built each day, what the benchmarks showed, what went wrong, and what changed because of it.

8days
12thunc releases, plus thunc-watch
61pull requests
61 → 765offline tests

The releases

  1. 0.1.0 First release: typed functions on Claude, Claude Code and Codex
  2. 0.1.1 The OpenAI backend and local models
  3. 0.1.2 cache=True, system= and a parser hardened over twelve rounds
  4. 0.1.3 The Jev backend
  5. 0.2.0 Agents
  6. 0.2.1 Durable agents with Temporal
  7. 0.2.2 Native tool calls on Claude Code, after a benchmark
  8. 0.2.3 Native tool calls everywhere agents run; the last of 0.2
  9. thunc-watch 0.1.0 A live dashboard in the terminal
  10. 0.3.0 thunc watch, and thunc write as an experiment
  11. 0.3.1 The changes that missed 0.3.0's package
  12. 0.3.2 A stability pass

Each release has its full notes in the changelog. This post is the story around them.

3 October · 0.1.0 and 0.1.1

A function that is a prompt

The idea behind thunc is that calling a model from code shouldn't need a prompt template, a JSON schema and a page of parsing code. You write the function you want. The docstring is the prompt, the parameters are the inputs, and the return type is the contract: thunc checks the reply against it, sends a wrong answer back to the model to try again, and raises a ThuncError if it still doesn't fit.

import thunc


@thunc.function
def urgency(ticket: str) -> int:
    """Rate how urgent this ticket is,
    from 1 (can wait) to 5 (customer is blocked)."""
    ...


urgency("I was charged twice!")  # 4, a checked int

The name is think + function. Two constraints were set on the first day and haven't moved since: thunc has no dependencies beyond the standard library, and it runs on whatever you already have. 0.1.0 shipped with three backends, the Claude API and the Claude Code and Codex command-line tools, so anyone logged in to one of those could try it without an API key. It also had thunc.call for prompts built in code, thunc.map to run many calls at once, ensure= for your own checks, and JSONL tracing.

The first commit with code landed at 13:33. 0.1.0 was on PyPI at 13:49, through PyPI's Trusted Publishing: creating a GitHub release builds and uploads the package, with no tokens to manage. By the end of the afternoon 0.1.1 added an OpenAI backend on the Responses API, and with it local models through OPENAI_BASE_URL, tested with LM Studio and gpt-oss-20b. The same afternoon added a CONTRIBUTING guide, GitHub Discussions and this website.

3–4 October · 0.1.2

Twelve rounds of trying to break the parser

Models don't always follow the contract. They wrap JSON in a code fence, open with a <think> block, answer {"rating": 4} when asked for an int, or reply NaN. The parser has to read the harmless near-misses, send the ambiguous ones back, and never return a wrong value as if it were right.

To find out where it failed, independent agents attacked it in rounds. Each round wrote 25 to 37 new tests against the contract, offline with scripted backends, on Python 3.10 and 3.11. Every finding was fixed with a regression test or written down as a known issue. Rounds 11 and 12 were regression hunts that compared the branch with main over more than 150,000 (reply, type) pairs. Against a fixed set of 102 realistic bad replies:

OutcomeBeforeAfter
Wrong value returned, no error185
Crash that skipped the retry loop20
Near-miss read correctly724

The last rounds taught the most useful lesson. They mostly found bugs in the branch's own cleverest rules: an exactness check for whole-number floats, unwrapping under a field name, an elaborate rule for a code fence's opening line. Instead of patching them again, the final commit removed them. The test suite went from 61 tests to 205, and only ThuncError can now escape a call because of a bad reply. The same release added cache=True, which asks the model once per input and keeps the answer on disk, and system= for your own system prompt.

The first release mistake. The 0.1.2 GitHub release was published about a minute before the version-bump pull request merged. The tag pointed at the old commit, the build produced 0.1.1 again, and PyPI rejected it: a version number can never be uploaded twice. The fix the next morning was to turn the release back into a draft, move the tag to the bump commit and publish again. The rule since then: merge the bump, check the version on main, then publish.

4 October · 0.1.3 and 0.2.0

Agents that hand back a typed answer

The morning's 0.1.3 added a backend for Jev, a small judgment model that answers yes/no, labels and ratings in about 0.3 seconds. It also made the Codex backend ignore your own Codex config and tools, so a call behaves the same on every machine.

Then the big one. A function answers from what you pass it; an agent can look around first. 0.2.0 introduced thunc.Agent: a typed task, written like a function, that lists, searches and reads files in a working directory and, where its permissions allow, edits them and runs commands, before it finishes with a checked value of its return type.

fixer = thunc.Agent(
    "fixer",
    workdir=".",
    permissions=["write:src/**", "run:pytest"],
)


@fixer.task
def fix_failing_tests() -> bool:
    """Run the tests and fix any problems that they surface."""
    ...

The pieces landed as seven stacked pull requests and merged within two minutes of each other:

The agent's default system prompt was a deliberate choice. Copying the prompts of Claude Code or Codex was ruled out early: their tools differ from thunc's, they assume a person at the keyboard, and they change between versions. thunc has a short prompt of its own, and live_tests/eval_prompts.py compares no prompt, the default and each preset on fixture repositories; 90 of 90 runs passed on the Claude API and Claude Code. CI started running the offline tests on Windows the same day.

4 October · 0.2.1

Durable agents, the same afternoon

An agent that edits files and runs commands for minutes at a time shouldn't lose its work when a process dies. 0.2.1 added an optional runtime on Temporal, installed with thunc[temporal]. Each model turn and tool call is recorded in a workflow, so a run survives a worker restart and can be picked up from another process. File writes go through an intent and a receipt, so a retried step doesn't write twice, and a command that may or may not have run before a crash waits for you to say what happened with resolve() rather than guessing.

To make that possible, the agent loop's decisions moved into one engine shared by local and durable runs. Local thunc stays dependency-free; the Temporal SDK is only needed by those who ask for it. The evening went to the README (quickstart first, with demo GIFs) and a new look for this site, with the docs section you can read today.

4 October · measuring

Measuring thunc's own cost

Before tuning anything, thunc got two ways to measure itself. thunc run --profile app.py runs a program and reports where its time went: per function, calls, cache hits, retries and model time against thunc's own; per agent task, steps and time in each tool. And benchmarks/ is a suite that times thunc against stand-in models that answer at once, so the numbers are thunc's work, not a model's. It found three speedups:

ChangeMeasuredBefore → after
The API backends reuse their connections20 calls, with 60 ms of simulated connection setup1.47 s → 69 ms
Text-protocol agents can send several actions per replyMean run time over 3 tasks on Codex, 30/30 passing both ways37 / 27 / 27 s → 29 / 19 / 21 s
Codex answers return when the turn ends, not when the CLI exits8 pairs of real callsfaster in 8 of 8, ~0.67 s each
4–5 October · 0.2.2

The benchmark that changed the plan

The next question was harder: when a thunc agent fails or runs slowly, how much of that is thunc rather than the model? The tool-use benchmark (live_tests/bench_tooluse.py) answers it with eight small repositories, each pressing on one part of a harness: a value hidden two hops from its call site behind a stale build copy, a bug near line 1,900 of a 2,400-line file, a rename across 14 files, a module that is mostly backslashes and quotes, tab-indented near-duplicates, one real error in 4,000 lines of build output, tests that only pass from a subfolder, and a structured answer gathered from four files with two decoys. The same model ran every task through thunc and through Claude Code itself.

The answer was uncomfortable. On Claude Sonnet 5.5, thunc agents on the Claude Code backend passed 12 of 24 runs. Claude Code passed 24 of 24. The model was the same in every column; only the harness changed.

The report traced most of the gap to the text protocol. Current models are trained to call tools natively, and asked instead to write one JSON action as plain text, they slip back. They wrote a correct action and kept going, making up the tool's result and the next steps. One failed model call ended a whole run. search and list waded through what git ignores, and run had no working directory and kept only the end of long output, where the first error wasn't.

The fix was to stop fighting the model. On Claude Code, thunc now serves the agent's tools as an MCP server that one claude -p process per run calls natively; a small relay hands each call back to thunc, which carries it out with its own tools, permissions and records. Failed steps are retried. The tool gaps closed. And the Claude Code backend now loads none of your Claude Code settings, so files in an agent's working directory can't give it instructions.

Claude Sonnet 5.5, 24 runsPassedSeconds per task$ per task
thunc, before12/24990.084
thunc, after (native calls)24/24100.022
Claude Code24/24110.069

Same model, same tasks: thunc agents now pass every run, slightly faster than Claude Code and at about a third of its cost. That shipped as 0.2.2 on the morning of 5 October, together with thunc run --profile.

5–6 October · 0.2.3

Closing out 0.2

The benchmark left a list of open items, and 0.2.3 was scoped to finish all of them before anything new: eight workstreams, one pull request each, stacked on one another. A few worth telling:

In the rerun, every harness passed 24 of 24, the text protocol included (20 of 24 in 0.2.2) in about half its previous time, and through the Claude API both Sonnet 5.5 and Opus 5.5 passed 8 of 8. 0.2.3 shipped on 6 October as the last of the 0.2 line.

6–7 October · 0.3.0 and 0.3.1

Watching a program

A program that makes dozens of calls and runs agents is hard to follow from its logs. thunc watch app.py runs it with a dashboard in the terminal: the calls waiting on a model, retries and why each reply was rejected, timings for each function, each agent's steps as they happen, and the profile report when the program ends. --agents follows agent runs from any process, and --plain prints one line per event for CI.

The dashboard is written in Rust with ratatui, which raised a packaging question. Bundling a compiled binary would have turned thunc into a set of platform wheels and broken installs where no wheel exists. So the dashboard ships as its own package, thunc-watch, installed with pip install "thunc[watch]", and thunc stays pure Python. thunc only writes one JSON line per event to the file named by THUNC_EVENTS, and nothing at all when it isn't set. See Watching a program.

Functions that write themselves

The other 0.3 feature started from a simple question: if the model can answer a function's calls, why not have it write the function? With @thunc.function(write=True), the first call asks for three things side by side: this call's answer, a draft of the body from the docstring and signature, and five test calls from a separate request, each answered by the model. The draft is linted, run on all six inputs and must match every answer. A passing draft goes into your file in place of ..., the decorator is removed, the checked calls become doctest examples, and from then on the function is plain Python with no model calls.

stderr, first call
thunc: writing minutes() in durations.py (first call)
thunc: checked against 6 model answers: all agree
thunc: wrote durations.py lines 5-27 in 14s (answer 4s, draft 11s, test calls 9s; side by side). Removed @thunc.function. Review: git diff durations.py

It took two tries to build the right thing. The first attempt, a day earlier, built something else: hand-written function bodies that fall back to the model when they're unsure. It was a reasonable feature but not the one that had been asked for, so it was reverted the same day, unreleased. The second attempt started by writing down which idea it delivered. Its first version used a whole agent as the writer and took about 39 seconds on Codex; replacing it with one typed call, and asking for the test calls separately so they don't share the draft's blind spots, brought that to 20.

thunc write edits your source files, so it ships as an experiment. It's refused in CI, in installed code, in files outside the project and in files changed since they were imported, and every change is a diff to review. See thunc write.

The second release mistake. 0.3.0 was published on 7 October from a draft made by the release-notes workflow. Editing the draft to point at a newer commit didn't move it: the tag was made on the commit the draft was created on, and two changes merged just before the release missed the package. 0.3.1 shipped them 13 minutes later. Tags are now pushed by hand at the merge commit and checked before a release is published.

10 October · 0.3.2

A stability pass

After a few days of new features, a review of 0.3.1 looked for ways a run could fail that it shouldn't. It found two.

When the Claude API sends an error in the middle of a streamed reply, say an overloaded server, the SDK raises it with the stream's HTTP status, 200. thunc read that as permanent and the agent run failed, though the docs promise such errors are retried. Now the error's type decides. In a simulation of 10-step runs with real thunc agents and a scripted client:

Requests failing mid-replyRuns finished in 0.3.1In 0.3.2
1%90.8%100.0%
5%58.9%99.9%
10%33.4%98.8%
20%10.0%92.1%

The second was a trace file that couldn't be written, in a missing folder for example. The error replaced the call's outcome: a successful call lost its answer, already paid for, and a failed one hid its own error. Now thunc warns once per path and the call's result stands.

What we learned

What's next

thunc is in beta, and the 0.3 line is about using it on real programs. The open candidates are on the issue tracker: Enum and Pydantic return types (#6, #7), record and replay for tests (#8), a precise type for thunc.call(returns=...) (#10) and more examples (#11). Whether thunc write leaves its experimental label depends on how it holds up in your code. If you try it, tell us how it went.

pip install thunc
Get started

Edit this page on GitHub