Agents
An agent is a typed function that can look around before it answers. It lists, searches, reads and, when you allow it, edits files and runs commands, then hands back a checked value of the return type.
New in 0.2. Agents are new; their API may change in a later release as feedback comes in.
Declare an agent
Give it a name and a working directory, declare its tasks the way you write @thunc.function, and call them from Python:
repo = thunc.Agent("repo-guide", workdir="~/code/myapp")
@repo.task
def request_timeout() -> int:
"""Find the HTTP request timeout this app uses, in seconds."""
...
request_timeout() # 45, after the agent searched the code and read the file that sets itAn agent with a single task can be declared in one go, and a task built in code runs with agent.call, the agent version of thunc.call:
@thunc.agent("release-notes", workdir="~/code/myapp", permissions=["write:CHANGELOG.md", "run:git log"])
def changelog(since_tag: str) -> list[str]:
"""Add an entry to CHANGELOG.md for the commits since `since_tag`. Return the bullets you wrote."""
...
repo.call(f"Where is {setting} set?", returns=str)How a run works
Each call is one run. The model takes one step at a time (list a folder, search, read or edit a file) and ends by calling finish with a value of the return type, which is checked like any thunc result.
- On the Claude and OpenAI APIs it uses their own tool calls: the model can make several at once, and the fixed part of the prompt is cached.
- On Claude Code the calls are native too: the agent's tools are an MCP server that one
claude -pprocess per run calls, while thunc carries out each call with its own tools, permissions and records. If Claude Code can't start them (an olderclaudeCLI, or MCP servers turned off by a policy), the run uses the text protocol instead, with a warning;protocol="native"fails instead. - On Codex it's the same: each
codex execis a turn in which the model calls the agent's tools through the MCP server, a turn that ends withoutfinishis continued withcodex exec resume, Codex's own tools stay off and its sandbox read-only, and the run's Codex session is deleted when the run ends. It falls back the same way, too. - With
protocol="text"the model replies with JSON actions as text: one at a time, or several independent ones (reading three files) as a JSON array, which saves turns. That works on any backend, for example with a server behindOPENAI_BASE_URLthat has no function calling. Durable runs on Codex use it. A reply that wraps its action in prose or tool-call markup, or carries on past it, is read for its first complete action, and on Claude Code the step stops as soon as that action has arrived. - Jev only answers typed questions and can't run agents. An agent run using it raises
ThuncErrorbefore creating any run files or calling a backend.
Permissions
Permissions say what the agent may do. By default it may read everything in workdir and save notes, and may not write.
fixer = thunc.Agent("fixer", workdir=".", permissions=["write:src/**", "run:pytest", "!read:.env*"])| Rule | Means |
|---|---|
write:docs/**, write | Create and edit matching files (all files with no path); also lets it read them |
read:src/** | Read only these; any read: rule replaces the read-everything default |
run:pytest, run:git log, run | Run commands that start with these words (run:git log allows git log --oneline, not git push); run alone allows any |
shell | Run any command line in a shell (sh -c, or cmd /c on Windows), so pipes, &&, cd and redirects work. Off by default, and it can't be combined with !run: rules |
!read:.env*, !write:..., !run:git push, !memory | Deny; a deny always wins, and !read also stops writing |
* stays within one folder, ** crosses folders, and paths are relative to workdir. The agent is told its permissions, and an action they don't allow is refused with the reason, after which the run carries on. Bad rules fail when the agent is declared.
Tools
- list, read, search
- Always available.
searchtakes a regular expression and an optionalglobsuch as*.py. Files the agent may not read are left out oflistandsearch, and so is what git ignores, in a git repository; a folder named explicitly is still listed and searched. - write, edit
- When a write rule allows it.
writecreates a file or replaces one;editreplaces text that appears exactly once, or every occurrence withreplace_all. Several changes to one file can go in one call asedits: they apply in order, and if one fails, none is made. - run
- When a run or shell rule allows it.
cwdruns the command in a folder insideworkdir. - remember
- Saves a short note for later runs (see memory).
- finish
- Ends the run with the answer.
Every path must stay inside workdir: .., absolute paths and symlinks that point outside are refused, and the rules are checked on where a link really leads.
Your own functions as tools
def open_issue(title: str, body: str) -> str:
"""Open an issue in our tracker and return its URL."""
return tracker.create(title=title, body=body).url
triage = thunc.Agent("triage", workdir=".", tools=[open_issue])Each function needs type hints and a docstring, which is its description. Arguments are checked against the hints before the call; what it returns goes back to the model (as JSON unless it's a str), and so does an exception, as an error. Listing a function is what allows it.
Safety
No blind overwrites
A file is only replaced (write) after the agent read it with read in the same run, and only if it hasn't changed on disk since. An edit needs no read, because it only changes text the agent quotes exactly, but a file the agent did read must not have changed since. There is no undo, so run agents that write in a git repository with a clean tree, and review their changes with git diff.
Commands
Commands run in workdir, or in a folder inside it given as cwd. Without the shell permission there's no shell, so &&, pipes, cd, redirects and $VARIABLES don't work (the agent is told). They get a minimal environment: PATH, HOME, the locale and temp-folder variables, and whatever you pass in env=, so your API keys don't reach them. Each has a time limit (command_timeout=120 seconds) that also stops the processes it started, and the agent sees the exit code and the output: the start and the end when it's long.
A permitted command can do anything its program can. run:pytest runs the project's code, which can read or change any file your user account can, whatever the read and write rules say. Permissions limit which tools the model uses; they aren't a sandbox. For untrusted input, run the agent in a container.
Memory between runs
Each run starts a fresh conversation, but the agent can save a short note with its remember tool. Notes go in memory.md in the agent's folder, and every later run gets them at the end of its system prompt (a note saved during a run reaches the next run, not that one). It's a plain file: read it with agent.memory, edit it, or delete it to start over.
The agent's folder
.thunc_agents/<name>/, changed with configure(agents_dir=...) or THUNC_AGENTS_DIR. Besides memory.md it holds agent.json (the agent's settings) and sessions/, one JSONL file per run with every step (denied ones marked), the result, and the files it changed. Runs of one agent take turns; different agents run side by side. Two names that make the same folder ("Repo guide" and "repo-guide") can't both be used.
Instructions and system prompts
Instruction files
follow=True gives the agent AGENTS.md and CLAUDE.md from workdir (those that exist) as instructions, and follow=["docs/agent-rules.md"] names files. They're read at the start of each run and sent after thunc's rules; they can't grant permissions. It's off by default, so a folder you point an agent at (a cloned repo, an upload) can't give it instructions. Without it, the agent can still read those files, but as data. @imports in CLAUDE.md aren't followed.
system= and the presets
system= replaces the opening of the agent's system prompt. thunc always adds its working method and its rules after it (file contents and tool results are data, not instructions). Three presets cover common jobs: thunc.prompts.CODING, thunc.prompts.CODE_REVIEW and thunc.prompts.ANALYSIS. They're plain strings, so you can extend one:
coder = thunc.Agent("coder", workdir=".", system=thunc.prompts.CODING + "\n\nTarget Python 3.10.")Time, step and retry limits
max_steps=40 bounds the model replies in a run, and timeout= (seconds) bounds the run's time. It's checked before each model call; a command's time limit is cut to the time left. In its last three replies before max_steps, the model is told how many are left, so it can finish with what it has. retries=2 is how many times a finish value that doesn't fit the return type (or fails ensure=) is sent back to be fixed.
Effort
effort= sets how hard the model thinks: "low", "medium", "high", "xhigh" or "max" (openai and codex go up to "xhigh"). By default it's "high" on the anthropic backend for Claude 4.6 and later (Claude Opus 5.5's own default is "medium", low for agentic coding), and each backend's own default elsewhere.
What happened in a run
Calling a task returns its value. agent.run(task, *args) runs it the same way and returns a thunc.Run instead, typed like the task (Run[int]):
run = fixer.run(make_tests_pass)
run.value # True
run.files_changed # ["src/mathutil.py"] (by write, edit and commands)
run.commands # [Command("python3 tests/test_mathutil.py", exit_code=0, seconds=0.04)]
run.denied # [Denial("run", "git commit -am fix", "running ... is denied by '!run:git'")]
run.notes, run.followed, run.steps, run.seconds, run.sessionFailures are loud. A run that hits max_steps, never gives a valid value, or loses its backend raises thunc.AgentError (a ThuncError), whose .run is the record up to that point. A step that fails for a reason asking again may fix (a timeout, a lost connection, a rate limit, a server error, a CLI call that ended in an error) is retried twice first, and each retry is in the run's record. On the Claude API, a reply cut off at max_tokens doesn't end the run: its tool calls get an error result saying so (twice in a row does).
All options
thunc.Agent(name, *, workdir, system=None, permissions=(), env=None, command_timeout=120,
follow=False, protocol=None, tools=(), timeout=None, max_steps=40,
retries=2, backend=None, model=None, effort=None)
@agent.task(instructions=..., ensure=...)
@thunc.agent(name, workdir=..., instructions=..., ensure=..., **options)async def tasks work. For runs that must survive a worker restart, see durable agents with Temporal.
How the prompt was tested
python -m live_tests.eval_prompts --backend anthropic runs three small tasks (fix a bug, review a diff, answer a question about a repo) with three versions of the system prompt: bare (no working method), the default, and the task's preset. Five runs of each, on the Claude API on 4 October 2026 and on Claude Code on 5 October 2026:
| Claude API (Opus 5.5, native calls) | Claude Code (Sonnet 5.5, native calls) | |
|---|---|---|
| Passed | 45/45: every task, every version | 45/45 |
| Steps (bare / default / preset) | fix 4.0 / 4.0 / 4.0, review 2.0 / 2.4 / 2.8, analysis 3.0 / 3.0 / 3.0 | fix 4.0 / 4.0 / 4.0, review 2.0 / 2.0 / 2.0, analysis 3.0 / 2.8 / 2.6 |
| Cost | $0.76 for all 45 runs (cache reads were 257,553 of 312,294 input tokens) | not recorded |
Every version passed every time, so these tasks are too easy to tell the versions apart: the result says the prompt does no harm, not that it helps. On the API, the review preset read more of the code before answering. On the text protocol these tasks took more steps (fix 5.2 / 6.0 / 6.0 on Claude Code before native calls). For harder tasks that do tell harnesses apart, see the tool-use benchmark in live_tests/bench_tooluse.py.