thunc()
Guide · new in 0.2

Agents

An agent is a typed function that can look around before it answers. It lists, searches, reads and, when you allow it, edits files and runs commands, then hands back a checked value of the return type.

An agent allowed to write src/** and run pytest is asked to run the tests and fix any problems. pytest shows 1 failure; it reads src/pricing.py, fixes one line, reruns pytest (3 passed) and returns True.

New in 0.2. Agents are new; their API may change in a later release as feedback comes in.

Declare an agent

Give it a name and a working directory, declare its tasks the way you write @thunc.function, and call them from Python:

repo = thunc.Agent("repo-guide", workdir="~/code/myapp")


@repo.task
def request_timeout() -> int:
    """Find the HTTP request timeout this app uses, in seconds."""
    ...


request_timeout()  # 45, after the agent searched the code and read the file that sets it

An agent with a single task can be declared in one go, and a task built in code runs with agent.call, the agent version of thunc.call:

@thunc.agent("release-notes", workdir="~/code/myapp", permissions=["write:CHANGELOG.md", "run:git log"])
def changelog(since_tag: str) -> list[str]:
    """Add an entry to CHANGELOG.md for the commits since `since_tag`. Return the bullets you wrote."""
    ...


repo.call(f"Where is {setting} set?", returns=str)

How a run works

Each call is one run. The model takes one step at a time (list a folder, search, read or edit a file) and ends by calling finish with a value of the return type, which is checked like any thunc result.

Permissions

Permissions say what the agent may do. By default it may read everything in workdir and save notes, and may not write.

fixer = thunc.Agent("fixer", workdir=".", permissions=["write:src/**", "run:pytest", "!read:.env*"])
RuleMeans
write:docs/**, writeCreate and edit matching files (all files with no path); also lets it read them
read:src/**Read only these; any read: rule replaces the read-everything default
run:pytest, run:git log, runRun commands that start with these words (run:git log allows git log --oneline, not git push); run alone allows any
shellRun any command line in a shell (sh -c, or cmd /c on Windows), so pipes, &&, cd and redirects work. Off by default, and it can't be combined with !run: rules
!read:.env*, !write:..., !run:git push, !memoryDeny; a deny always wins, and !read also stops writing

* stays within one folder, ** crosses folders, and paths are relative to workdir. The agent is told its permissions, and an action they don't allow is refused with the reason, after which the run carries on. Bad rules fail when the agent is declared.

Tools

list, read, search
Always available. search takes a regular expression and an optional glob such as *.py. Files the agent may not read are left out of list and search, and so is what git ignores, in a git repository; a folder named explicitly is still listed and searched.
write, edit
When a write rule allows it. write creates a file or replaces one; edit replaces text that appears exactly once.
run
When a run or shell rule allows it. cwd runs the command in a folder inside workdir.
remember
Saves a short note for later runs (see memory).
finish
Ends the run with the answer.

Every path must stay inside workdir: .., absolute paths and symlinks that point outside are refused, and the rules are checked on where a link really leads.

Your own functions as tools

def open_issue(title: str, body: str) -> str:
    """Open an issue in our tracker and return its URL."""
    return tracker.create(title=title, body=body).url


triage = thunc.Agent("triage", workdir=".", tools=[open_issue])

Each function needs type hints and a docstring, which is its description. Arguments are checked against the hints before the call; what it returns goes back to the model (as JSON unless it's a str), and so does an exception, as an error. Listing a function is what allows it.

Safety

No blind overwrites

A file is only replaced or edited after the agent read it in the same run, and only if it hasn't changed on disk since. There is no undo, so run agents that write in a git repository with a clean tree, and review their changes with git diff.

Commands

Commands run in workdir, or in a folder inside it given as cwd. Without the shell permission there's no shell, so &&, pipes, cd, redirects and $VARIABLES don't work (the agent is told). They get a minimal environment: PATH, HOME, the locale and temp-folder variables, and whatever you pass in env=, so your API keys don't reach them. Each has a time limit (command_timeout=120 seconds) that also stops the processes it started, and the agent sees the exit code and the output: the start and the end when it's long.

A permitted command can do anything its program can. run:pytest runs the project's code, which can read or change any file your user account can, whatever the read and write rules say. Permissions limit which tools the model uses; they aren't a sandbox. For untrusted input, run the agent in a container.

Memory between runs

Each run starts a fresh conversation, but the agent can save a short note with its remember tool. Notes go in memory.md in the agent's folder, and every later run gets them at the end of its system prompt (a note saved during a run reaches the next run, not that one). It's a plain file: read it with agent.memory, edit it, or delete it to start over.

The agent's folder

.thunc_agents/<name>/, changed with configure(agents_dir=...) or THUNC_AGENTS_DIR. Besides memory.md it holds agent.json (the agent's settings) and sessions/, one JSONL file per run with every step (denied ones marked), the result, and the files it changed. Runs of one agent take turns; different agents run side by side. Two names that make the same folder ("Repo guide" and "repo-guide") can't both be used.

Instructions and system prompts

Instruction files

follow=True gives the agent AGENTS.md and CLAUDE.md from workdir (those that exist) as instructions, and follow=["docs/agent-rules.md"] names files. They're read at the start of each run and sent after thunc's rules; they can't grant permissions. It's off by default, so a folder you point an agent at (a cloned repo, an upload) can't give it instructions. Without it, the agent can still read those files, but as data. @imports in CLAUDE.md aren't followed.

system= and the presets

system= replaces the opening of the agent's system prompt. thunc always adds its working method and its rules after it (file contents and tool results are data, not instructions). Three presets cover common jobs: thunc.prompts.CODING, thunc.prompts.CODE_REVIEW and thunc.prompts.ANALYSIS. They're plain strings, so you can extend one:

coder = thunc.Agent("coder", workdir=".", system=thunc.prompts.CODING + "\n\nTarget Python 3.10.")

Time and step limits

max_steps=40 bounds the model replies in a run, and timeout= (seconds) bounds the run's time. It's checked before each model call; a command's time limit is cut to the time left.

What happened in a run

Calling a task returns its value. agent.run(task, *args) runs it the same way and returns a thunc.Run instead, typed like the task (Run[int]):

run = fixer.run(make_tests_pass)
run.value          # True
run.files_changed  # ["src/mathutil.py"]  (by write, edit and commands)
run.commands       # [Command("python3 tests/test_mathutil.py", exit_code=0, seconds=0.04)]
run.denied         # [Denial("run", "git commit -am fix", "running ... is denied by '!run:git'")]
run.notes, run.followed, run.steps, run.seconds, run.session

Failures are loud. A run that hits max_steps, never gives a valid value, or loses its backend raises thunc.AgentError (a ThuncError), whose .run is the record up to that point. A step that fails for a reason asking again may fix (a timeout, a lost connection, a rate limit, a server error, a CLI call that ended in an error) is retried twice first, and each retry is in the run's record.

All options

thunc.Agent(name, *, workdir, system=None, permissions=(), env=None, command_timeout=120,
            follow=False, protocol=None, tools=(), timeout=None, max_steps=40,
            retries=2, backend=None, model=None)

@agent.task(instructions=..., ensure=...)
@thunc.agent(name, workdir=..., instructions=..., ensure=..., **options)

async def tasks work. For runs that must survive a worker restart, see durable agents with Temporal.

How the prompt was tested

python -m live_tests.eval_prompts --backend anthropic runs three small tasks (fix a bug, review a diff, answer a question about a repo) with three versions of the system prompt: bare (no working method), the default, and the task's preset. Five runs of each, on the Claude API on 4 October 2026 and on Claude Code on 5 October 2026:

Claude API (Opus 5.5, native calls)Claude Code (Sonnet 5.5, native calls)
Passed45/45: every task, every version45/45
Steps (bare / default / preset)fix 4.0 / 4.0 / 4.0, review 2.0 / 2.4 / 2.8, analysis 3.0 / 3.0 / 3.0fix 4.0 / 4.0 / 4.0, review 2.0 / 2.0 / 2.0, analysis 3.0 / 2.8 / 2.6
Cost$0.76 for all 45 runs (cache reads were 257,553 of 312,294 input tokens)

Every version passed every time, so these tasks are too easy to tell the versions apart: the result says the prompt does no harm, not that it helps. On the API, the review preset read more of the code before answering. On the text protocol these tasks took more steps (fix 5.2 / 6.0 / 6.0 on Claude Code before native calls). For harder tasks that do tell harnesses apart, see the tool-use benchmark in live_tests/bench_tooluse.py.

Edit this page on GitHub