Metadata-Version: 2.4
Name: ratchet-harness
Version: 0.2.0
Summary: Ratchet: a small, careful agentic coding harness for Claude.
License: MIT
Keywords: claude,anthropic,agent,coding-assistant,cli
Classifier: Development Status :: 3 - Alpha
Classifier: Environment :: Console
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: MIT License
Classifier: Operating System :: POSIX :: Linux
Classifier: Operating System :: MacOS
Classifier: Programming Language :: Python :: 3
Classifier: Topic :: Software Development
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: anthropic<2,>=1.11
Requires-Dist: httpx2<3,>=2
Requires-Dist: rich>=13.7
Requires-Dist: prompt_toolkit>=3.0.48
Provides-Extra: dev
Requires-Dist: pytest>=8; extra == "dev"
Dynamic: license-file

# Ratchet

A small agentic coding harness for Claude that you can read in one sitting.

Ratchet runs in your terminal. You give it a task, and Claude reads code,
searches the project, edits files and runs commands in a loop until the task is
done. Edits and commands wait for your approval unless you say otherwise.

The name describes the design: **the conversation only moves forward.** History
is append-only. Ratchet never rewrites, prunes or reorders what it has sent.
Two things depend on that: the prompt cache, which only matches an unchanged
prefix, and Claude's thinking blocks, which stay valid only while the
conversation before them is unchanged.

## Install

Needs Python 3.10+ on Linux or macOS. `~/.local/bin` must be on your `PATH`
(it is on most distributions). Uninstall with `uv tool uninstall ratchet-harness`.

```bash
uv tool install ratchet-harness   # or: pipx install ratchet-harness; puts `ratchet` on your PATH
ratchet --set-key anthropic       # or: export ANTHROPIC_API_KEY=sk-ant-...
```

Working on Ratchet itself? From a clone, `uv tool install --editable .` makes
source edits apply immediately.

### Saving keys

Instead of exporting keys in every shell, save them once:

```bash
ratchet --set-key kimi         # prompts without echoing; also anthropic, openrouter, glm, custom
pass show zai | ratchet --set-key glm   # or pipe it in from a password manager
```

Inside a session, `/key [provider]` does the same and takes effect immediately.
Keys go to `~/.config/ratchet/keys.env` (or `$XDG_CONFIG_HOME/ratchet/keys.env`),
in a directory only you can read. It's a plain `NAME=value` file, so you can
also edit it by hand, and any setting Ratchet reads from the environment works
there too: `RATCHET_PROVIDER=glm`, `ZAI_BASE_URL=...`, `RATCHET_EFFORT=medium`.
Variables exported in your shell take precedence over the file.

### Using OpenRouter

```bash
export OPENROUTER_API_KEY=sk-or-...
ratchet --provider openrouter                       # anthropic/claude-opus-5.5 by default
ratchet --provider openrouter --model anthropic/claude-sonnet-5.5
export RATCHET_PROVIDER=openrouter                  # make it the default
```

Ratchet sends OpenRouter the same Anthropic Messages requests it sends Anthropic,
through OpenRouter's Anthropic-compatible endpoint (`https://openrouter.ai/api`,
or `$OPENROUTER_BASE_URL`), using the official SDK. Model names use OpenRouter's
slugs (`anthropic/claude-opus-5.5`). Differences from going direct:

- **Auth:** your key goes out as `Authorization: Bearer`. Even if
  `ANTHROPIC_API_KEY` is set, it is never sent to OpenRouter.
- **Sticky routing:** every request carries the session ID as OpenRouter's
  `session_id`, so a session stays with one upstream provider. Switching
  providers mid-session would lose prompt-cache hits and could invalidate
  thinking blocks.
- **Tool input streaming:** `eager_input_streaming` is off, because OpenRouter's
  schema doesn't include it. Tool inputs arrive whole instead of streaming in.
- **Refusal fallback:** Anthropic's `fallbacks: "default"` isn't sent, because
  OpenRouter's `fallbacks` field means something different.
- **Unchanged:** compaction, caching, thinking and effort work as they do direct.

Sessions remember their provider, so `ratchet --resume` picks it up without the
flag. Other Claude models work the same way.

### Bring your own key: Kimi, GLM, or any Anthropic-compatible API

Moonshot (Kimi) and Z.ai (GLM) both serve the Anthropic Messages API, so Ratchet
talks to them directly with your own key, without going through OpenRouter.

```bash
export MOONSHOT_API_KEY=sk-...                      # or KIMI_API_KEY
ratchet --provider kimi                             # kimi-k2.5 by default
ratchet --provider kimi --model kimi-k2-thinking

export ZAI_API_KEY=...                              # or ZHIPUAI_API_KEY / GLM_API_KEY
ratchet --provider glm                              # glm-4.7 by default
ratchet --provider glm --model glm-4.6
```

| Provider | Key | Default endpoint | Override with |
|---|---|---|---|
| `kimi` (alias `moonshot`) | `$MOONSHOT_API_KEY` | `https://api.moonshot.ai/anthropic` | `$MOONSHOT_BASE_URL` (China: `https://api.moonshot.cn/anthropic`) |
| `glm` (aliases `zai`, `zhipu`) | `$ZAI_API_KEY` | `https://api.z.ai/api/anthropic` | `$ZAI_BASE_URL` (China: `https://open.bigmodel.cn/api/anthropic`) |
| `custom` | `$RATCHET_CUSTOM_API_KEY` | none: set `$RATCHET_CUSTOM_BASE_URL` | model from `$RATCHET_CUSTOM_MODEL` or `--model` |

`custom` takes any other Anthropic-compatible endpoint (DeepSeek, MiniMax, a
self-hosted gateway, and so on). Give the API root: the SDK appends
`/v1/messages`. The key goes out as `Authorization: Bearer`, and
`ANTHROPIC_API_KEY` is never sent to these providers. Use `/provider` to switch
in a session (it starts a new conversation), and `/model` to pick from that
provider's models or type any ID it serves.

Ratchet only sends Claude-only features to Claude models:

- **Thinking:** non-Claude models get classic
  `thinking: {"type": "enabled", "budget_tokens": N}`. `--effort` picks the
  budget: 2k (low), 8k (medium), 16k (high), 32k (xhigh), or everything below
  `--max-tokens` (max). Adaptive thinking and `output_config.effort` go to
  Claude only.
- **Not sent:** automatic conversation caching (these providers cache prefixes
  on their own), refusal fallback, `eager_input_streaming`, and compaction.
  Long sessions on these models therefore aren't compacted. Watch `/context`.
- **Cost:** Ratchet only knows Claude prices, so the status line shows
  `cost n/a`. Check your provider's console.

### Using Portail

[Portail](https://portail.cc) users can run Ratchet on their Portail account
instead of an API key, and send build tasks to it from a Portail chat.

```bash
ratchet --link                 # sign in to Portail in the browser; links this computer
ratchet --provider portail     # run on your Portail credits (~anthropic-claude-opus-latest by default)
ratchet --watch -y             # wait for hand-offs from Portail chats and build each one
ratchet --unlink               # forget the link on this computer
```

`--link` opens `portail.cc/connect`, where you approve the sign-in. Portail sends
the browser back to a one-time port on `127.0.0.1`, and Ratchet trades what it
receives for a **device token** bound to your account. Ratchet keeps only that
token, as `PORTAIL_DEVICE_TOKEN` in `~/.config/ratchet/keys.env`. Unlinking
removes it here. Also remove the device from your linked devices in Portail,
or the token keeps working.

**Running on Portail credits.** `--provider portail` (or `/provider portail`)
sends each request to Portail's chat API, which bills your plan's credits like a
chat message. Model IDs are Portail's own: `~anthropic-claude-opus-latest`,
`~anthropic-claude-sonnet-latest`, `openai-gpt-5-3-codex` and others. `/model`
lists them and takes any other ID Portail offers. Differences from going direct:

- **Thinking:** `--effort` maps to Portail's `low`, `medium` and `high`;
  `xhigh` and `max` send `high`. Thinking streams in as usual but isn't kept in
  the conversation.
- **Not sent:** prompt caching, compaction and the refusal fallback. Portail
  takes requests up to 1 MB, so start a new conversation (`/clear`) when a long
  one is refused for its size.
- **Limits:** Portail allows one request at a time per account and 30 a minute.
  Ratchet waits and retries when it hits either.
- **Cost:** credits, not dollars, so the status line shows `cost n/a`. Token
  counts are estimates, since Portail doesn't report them.

**Hand-offs.** On Pro and Ultra plans, when a Portail chat turns into something
buildable, Portail can offer to send it to your linked computer. It turns the
conversation into a build brief and queues it. While Ratchet runs and is linked,
it checks in every 5 seconds, which is what keeps the computer "online" for
those offers. Waiting hand-offs show in the status line.

- **Interactive:** `/tasks` lists them. Pick one to build it in the current
  workspace in a new conversation. Your permission mode applies as usual.
- **Unattended:** `ratchet --watch` builds each hand-off as it arrives, in a new
  folder under `--root` named after it. Without `-y`, it still asks before edits
  and commands.

Ratchet reports `building`, then `completed` or `failed`, back to Portail with
the project's path. On Portail, a build uses the model picked in the chat.
Elsewhere it uses your current model.

## Use

```bash
ratchet                                  # interactive session in the current directory
ratchet "fix the failing test in tests/test_api.py"   # start with a task, stay interactive
ratchet -p "add type hints to utils.py"  # run one task to completion, then exit
ratchet -p -y "run the tests and fix what fails"      # same, without approval prompts
ratchet --mode plan "explain how auth works"          # can look, can't touch
ratchet --resume                         # pick up the latest session for this directory
```

### The interface

In a terminal:

- **Streaming output.** Claude's answers stream in as rendered markdown, each
  reply marked with `◆`, above a status line that says what's happening
  (`◜ Writing  12s · 1.2k tokens out · esc to interrupt`).
- **Thinking.** Summaries show dimmed behind a `┆` gutter; `/thinking` hides them.
- **Tool calls.** Each one shows as `▸ Read src/app.py`, with its result in a
  `│` gutter underneath (`✗` marks a failed call). Edits show a red/green diff
  with the file's line numbers, and shell commands show the first lines of
  their output.
- **Approvals.** A `? Edit file` or `? Run command` header shows the diff or
  command, then a menu: `y` allow once / `a` allow that kind for the session /
  `n` deny and say what to do instead. Your note goes to Claude with the denial.
- **Status line.** Under the input it shows the mode, model, effort, how full
  the context window is, and the session cost.

| Key | |
|---|---|
| `/` | commands, with completion as you type |
| `!command` | run a shell command yourself; its output goes along with your next message |
| `@path` | mention a file (Tab completes paths) |
| `?` | list shortcuts |
| Shift+Tab | cycle mode: default → accept edits → plan → auto |
| Esc | interrupt Claude (Ctrl-C works too) |
| Ctrl-C | clear the input; twice on an empty line to exit |
| `\` + Enter, Esc then Enter, Ctrl-J | new line |

| Command | |
|---|---|
| `/model [name]` | switch model: a picker, or a name like `sonnet`, `opus 5`, `fable`, or a full ID |
| `/effort [level]` | `low` · `medium` · `high` · `xhigh` · `max` |
| `/mode [mode]` | `default` · `accept-edits` · `plan` · `auto` |
| `/provider [name]` | `anthropic`, `openrouter`, `kimi`, `glm`, `portail` or `custom` (starts a new conversation) |
| `/key [provider]` | save a provider's API key to `~/.config/ratchet/keys.env` |
| `/thinking [on\|off]` | show or hide thinking summaries |
| `/status` | provider, model, effort, mode, session, and what's always allowed |
| `/usage`, `/cost` | tokens and estimated cost |
| `/context` | how full the context window is |
| `/clear`, `/new` | clear the screen and start a new conversation |
| `/resume [id]` | pick up an earlier conversation in this workspace |
| `/init` | have Claude write an `AGENTS.md` for the project |
| `/link`, `/unlink` | link this computer to your Portail account, or forget the link |
| `/tasks` | build a hand-off sent from a Portail chat |
| `/help`, `/exit` | |

Model and effort changes apply to the next request and keep the conversation.
Switching models means the prompt cache starts fresh.

**Modes.**

- `default` asks before every edit and command.
- `accept-edits` edits files without asking, but still asks before commands.
- `plan` is read-only. Claude explores and proposes a plan, and nothing
  changes until you leave plan mode. Entering and leaving plan mode is
  announced to Claude in your next message, so the conversation stays
  append-only.
- `auto` runs everything without asking. Use it only where mistakes are cheap.

When output isn't a terminal (pipes, `-p`), Ratchet falls back to plain text.

| Option | Default | |
|---|---|---|
| `--provider` | `anthropic` (or `$RATCHET_PROVIDER`) | `anthropic`, `openrouter`, `kimi`, `glm`, `portail` or `custom` |
| `--model` | `claude-opus-5-5` (or `$RATCHET_MODEL`); `anthropic/claude-opus-5.5` on OpenRouter, `kimi-k2.5` on Kimi, `glm-4.7` on GLM | Claude: any model with adaptive thinking (Opus 4.6+, Sonnet 4.6+, Fable). Other providers: any model ID they serve |
| `--effort` | `high` (or `$RATCHET_EFFORT`) | `low` · `medium` · `high` · `xhigh` · `max` |
| `--mode` | `default` | `-y` = `--mode auto`, `--read-only` = `--mode plan` |
| `--max-tokens` | 64000 | Output cap per response, thinking included |
| `--max-steps` | 200 | Model calls allowed per message you send |
| `--no-fallback` | on | Server-side refusal fallback (see below) |
| `--no-compaction` | on | Server-side context compaction (see below) |
| `--hide-thinking` | shown | Hide Claude's thinking summaries |
| `--link`, `--unlink` | | Link this computer to Portail, or forget the link, then exit |
| `--watch` | | Build Portail hand-offs as they arrive, each in a new folder under `--root` |
| `--root` | `.` | Workspace root |

Ratchet adds `RATCHET.md`, `AGENTS.md` or `CLAUDE.md` from the workspace root
to the system prompt, if any exist.

## Tools

| Tool | Needs approval | Notes |
|---|---|---|
| `read_file` | no | Line-numbered, pages through large files |
| `glob` | no | Newest first, skips `node_modules`, `.venv`, `.git` and similar |
| `grep` | no | Python regex, respects `.gitignore` |
| `edit_file` | yes | Exact, unique string replacement; keeps CRLF files CRLF |
| `write_file` | yes | Creates or overwrites whole files |
| `bash` | yes | Runs in the workspace root, with a timeout that kills the whole process group |
| `skill` | no | Loads one of Ratchet's built-in skills (see below) |

File tools can't reach outside the workspace root (symlinks included).
`edit_file` and `write_file` refuse to change a file Claude hasn't read in this
session, or one that changed on disk since, so Claude can't overwrite work it
hasn't seen. `bash` is not confined, so it asks first unless you pass `-y`.
Only use `-y` where a bad command can't do real damage.

When Claude asks for several read-only tools at once they run in parallel.
Anything that changes state runs one at a time, in order.

### Skills

Skills are guides for particular kinds of work. Each one is a
`src/ratchet/skills/<name>/SKILL.md` file, with `name` and `description` in
frontmatter. The system prompt lists only names and descriptions. When a task
matches, Claude loads the full text with the `skill` tool, so it arrives as a
tool result and the frozen prompt stays small. To add a skill, add a directory
with its own `SKILL.md`.

| Skill | |
|---|---|
| `frontend-design` | UI work: layout, type, color, states, motion, accessibility and a review checklist, drawn from Apple's HIG, Material Design and WCAG |

## How the loop works

All of it is in `src/ratchet/agent.py`.

- **Streaming and thinking.** Requests stream, with adaptive thinking and
  `display: "summarized"`, so you see Claude's reasoning summaries and progress
  notes as they arrive rather than a long silence.
- **Caching.** The system prompt and the tool list are fixed when the session
  starts and carry a cache breakpoint. Top-level automatic caching covers the
  growing conversation. `/usage` shows the cache reads.
- **Stop reasons.** `tool_use` runs the tools and continues. `end_turn` waits
  for you. `max_tokens` never runs a tool whose input may have been cut off;
  Claude gets an error result and can retry in smaller pieces. `refusal` runs
  nothing and discards the partial turn. `pause_turn` resumes.
- **Tool input validation.** Tool inputs stream as Claude writes them
  (`eager_input_streaming`), so the API doesn't validate them. Ratchet checks
  every input against the tool's schema before running it, and re-sends the
  request if the SDK couldn't parse the JSON at all.
- **No dangling calls.** Every `tool_use` gets a `tool_result` in the same
  append, even if you interrupt or deny it. If Claude never answered a message,
  Ratchet removes it instead, so a retry doesn't send it twice.
- **Refusal fallback.** On models that support it, requests carry
  `fallbacks: "default"`. If a safety classifier declines a request, the API
  retries it on a fallback model chosen for that refusal category. If a model
  switches mid-answer, Ratchet drops the declined part's tool calls and
  thinking before sending the turn back, and never runs those calls.
- **Compaction.** Long sessions use server-side compaction (`compact_20260112`).
  The API summarizes older context, and Ratchet keeps the compaction blocks in
  history as the API requires.
- **Sessions** are saved after every step to
  `~/.local/share/ratchet/sessions/`, outside your repo, together with the
  frozen system prompt, so a resumed session sends exactly the same prefix.

## Development

```bash
uv venv && uv pip install -e ".[dev]"
.venv/bin/python -m pytest
```

The tests run the real `anthropic` client against a scripted Messages API over
a mock HTTP transport that serves real SSE. They cover the SDK's streaming and
accumulation code as well as Ratchet's, and they check the exact request
bodies. The checks include byte-identical prefixes across turns, thinking
signatures echoed unchanged, and a valid `tool_use`/`tool_result` pairing
after every interrupt, denial, truncation, refusal and fallback. A fake Portail
(chat SSE, bridge and sign-in) covers the `portail` provider, linking and
hand-off builds the same way.

## Layout

```
src/ratchet/
  agent.py    the loop: streaming, stop reasons, tool execution, permissions
  tools.py    tool schemas, input validation, implementations
  prompt.py   system prompt (built once per session)
  skills.py   skill discovery; skills/<name>/SKILL.md holds the guides
  session.py  save / resume
  keys.py     saved API keys and settings (~/.config/ratchet/keys.env)
  portail.py  Portail: account linking, the hand-off bridge, the portail provider
  models.py   providers, model catalogs, provider model IDs, permission modes
  tui.py      interactive output: streaming markdown, tool lines, diffs, approval menus
  repl.py     interactive input: slash commands, completion, status line, shortcuts
  ui.py       plain-text output for -p and pipes
  cli.py      arguments, provider clients, one-shot mode, --watch
```
