Metadata-Version: 2.5
Name: rindo-runner
Version: 1.3.0
Summary: The Rindo project-agent runner: polls your Rindo server for agent jobs and runs your agent command on a machine you control
Project-URL: Homepage, https://rindo.io
Project-URL: Repository, https://github.com/rindohq/rindo
Author: Rindo
License-Expression: Apache-2.0
License-File: LICENSE
Keywords: agent,mcp,rindo,runner
Classifier: Development Status :: 5 - Production/Stable
Classifier: Intended Audience :: Developers
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Topic :: Software Development
Requires-Python: >=3.11
Requires-Dist: httpx>=0.28
Description-Content-Type: text/markdown

# rindo-runner

The Rindo **project-agent runner** (PRD AGT-7, architecture §5). A thin CLI that
gives a registered project agent hands: it long-polls the Rindo server for jobs
(triggered by an `@mention` or a task assignment), runs your configured command
against the job's context, keeps the job's lease alive, and reports the result.

Since 1.1 (decision #90, RP2) it is also the **shared runner**: one `rndw_`
credential — registered into a *pool* on **Account → Runners & pools** — serves
many agents of many projects (one pool, one runtime), and every job runs in a
NEW container of an image you name — or, since 1.2 (RP3b), the runner is itself
a one-job **ephemeral** runner that a provisioner starts inside a sandbox
(`executor = "process"`). See [Shared runner](#shared-runner-rndw_-11) and
[Ephemeral runner](#ephemeral-runner-executor--process-12) below; everything
else on this page describes the legacy single-agent runner unless it says
otherwise.

The server stays dumb — it only queues jobs. **This runner never calls an LLM and
never spawns work on the server.** All execution happens here, on a machine you
control, behind your own NAT if you like (the protocol is outbound-only HTTPS).

## Install

`rindo-runner` is Apache-2.0 (`LICENSE` beside this file) and is published on PyPI:

```sh
uv tool install rindo-runner          # or: pip install rindo-runner
uvx rindo-runner --help               # a throwaway environment on demand
```

To install an exact source commit instead (the repository is private — read access
needed): `uv tool install "rindo-runner @ git+ssh://git@github.com/rindohq/rindo.git@<sha>#subdirectory=runner"`,
where the ref is a commit sha or a product tag (`v*`). A git ref pins what is installed;
`rindo-runner --version` names the PACKAGE — the runner versions on its own line (1.1.0 takes a `rndw_`; 1.0.0 refuses it
at load), while `v*` tags are the PRODUCT's (owner ruling RP2-1).

Requires Python 3.11+ (provisioned automatically by `uv`/`uvx`). The legacy
single-agent runner runs on Linux, macOS, and Windows; the shared runner (a
container per job) is Linux-first and needs a container engine on the host; an
ephemeral runner (1.2) needs no engine — it runs its one job in its own process
inside the sandbox it was started in.

### Or: the agent-host image

`Dockerfile` here builds a ready-to-run **agent host** — the runner plus the
coding agents it exists to drive and the tools they reach for:

| Layer          | Contents                                                                     |
| -------------- | ---------------------------------------------------------------------------- |
| Agents         | **Claude Code** (`claude`), **Codex** (`codex`)                               |
| Forge CLIs     | **`gh`** (GitHub) and **`glab`** (GitLab) — Rindo integrates both providers  |
| Tooling        | `git` + `git-lfs`, `rg`, `fd`, `jq`, `make`, `curl`, `ssh`, `unzip`, `gnupg`  |
| Languages      | Node 24 (base), `python3`, `uv`/`uvx` for throwaway Python envs               |
| Runtime        | `rindo-runner` in its own venv, `tini` as PID 1, non-root user `rindo` (uid 1000), `/workspace` |

```sh
cd runner
docker build -t rindo-agent-runner:local .             # everything above (~1.5 GB)
docker build -t rindo-runner:local --target runner .  # no agent CLIs (~675 MB)

docker run --rm \
  -e RINDO_RUNNER_SERVER_URL=https://rindo.example.com \
  -e RINDO_RUNNER_TOKEN=rndr_… \
  -e RINDO_RUNNER_COMMAND='claude -p --dangerously-skip-permissions {payload_file}' \
  -e ANTHROPIC_API_KEY=sk-ant-… \
  rindo-agent-runner:local
```

`compose.example.yml` is the same thing as a long-lived service (persistent
`~/.claude` / `~/.codex` volumes, a `stop_grace_period` that fits the runner's
graceful shutdown, an opt-in healthcheck). One container = **one agent** on the
legacy path — the runner is serial, so a second agent means a second service with
its own token. The shared runner turns this around: it runs on the HOST and starts
a container of this image per JOB (the `agents` target, or a converter-capable image
built FROM the `runner` target with rindo-convert and LibreOffice — the slim `runner`
target alone carries neither), with the job's own credential inside. An ephemeral
runner (1.2) is the third shape: the SAME image is the sandbox template a
provisioner creates sandboxes from, and the runner inside it starts no container.

#### Rindo is wired as an MCP server automatically

Both agents can call [the MCP catalog](../docs/04-mcp-catalog.md) as the job's
own agent, with no setup — the server named `rindo` is already registered:

| | How it resolves | When |
| --- | --- | --- |
| Claude Code | `${RINDO_MCP_URL}` + `Authorization: Bearer ${RINDO_RUNNER_TOKEN}`, expanded by Claude Code itself | baked at build; resolved per connection |
| Codex | `url` derived from `RINDO_RUNNER_SERVER_URL` (+ `/mcp`, subpaths preserved); token via `bearer_token_env_var` | written at container start |

**No token is ever written to disk** — both sides hold only the *name* of the
env var, and the runner injects the real value per job, so rotating the runner
token takes effect on the next job with no rebuild. On the shared path the value
is the JOB's `rndj_` token, minted by the server at claim and dead when the job
ends — the runner's own `rndw_` never enters a job. The split exists because
Codex performs no env expansion in `url`; its token half needs no help.

The container-start half also installs the Codex API key (see Credentials
below). Set `RINDO_AGENT_AUTOCONFIG=0` to skip all of it and manage
`~/.codex/config.toml` yourself; `docker-entrypoint.sh` is the whole of it.

Notes:

- **Versions are pinned** (`CLAUDE_CODE_VERSION`, `CODEX_VERSION`, `GH_VERSION`,
  `GLAB_VERSION`, `NODE_BASE`) — bump with `--build-arg`, and note the image sets
  `DISABLE_AUTOUPDATER=1` so Claude Code cannot silently defeat the pin. Both
  forge CLIs are sha256-verified against their publishers' checksum lists.
- **Forge auth passes through** like any other env: `GH_TOKEN` / `GITHUB_TOKEN`
  for `gh`, `GITLAB_TOKEN` (plus `GITLAB_HOST` for a self-hosted instance) for
  `glab`. Neither is baked in.
- **Arch-native.** Both agent CLIs ship per-platform binaries, so build on the
  arch you will run on (or `docker buildx build --platform …`).
- **Credentials.** The LEGACY runner hands its whole environment to the spawned
  command, so Claude Code picks up `ANTHROPIC_API_KEY` with no wiring; a SHARED
  runner passes only the names you list in `pass_env` (nothing by default); an
  EPHEMERAL runner's child inherits the sandbox's environment — what enters the
  sandbox is the provisioner's `sandbox_env` allowlist, nothing by default. **Codex ignores
  `OPENAI_API_KEY` for auth** — `codex doctor` reports the variable as present
  while `auth mode` stays `none`, and every request 401s with "Missing bearer or
  basic authentication in header". The entrypoint therefore runs
  `codex login --with-api-key` for you, but only when Codex is logged out, so a
  ChatGPT session in a mounted `~/.codex` volume is never overwritten. To log in
  interactively instead, exec into a running container
  (`docker compose exec agent claude /login`) with the home volumes mounted —
  legacy path only: a shared runner's job starts from a fresh HOME every time, so
  agent auth there is an env key in `pass_env` (for an ephemeral runner, in the
  provisioner's `sandbox_env`), never an interactive login.
- **Codex's own sandbox needs a privilege Docker withholds.** `bubblewrap` is
  installed (its absence printed a warning into every job log), but it still
  wants unprivileged user namespaces, which Docker's default seccomp/AppArmor
  profiles block — so every shell command Codex runs dies with `bwrap: No
  permissions to create a new namespace`. Either run the container with
  `--security-opt seccomp=unconfined` (Codex keeps sandboxing itself), or turn
  Codex's sandbox off and let the container be the boundary. Claude Code is
  unaffected — it has no inner sandbox.
- **Every env var reaches model-generated shell commands.** Codex's default
  `shell_environment_policy` passes the environment through unfiltered (verified
  identical to `inherit=all`), so an injected instruction can read
  `RINDO_RUNNER_TOKEN`, `OPENAI_API_KEY` and `GH_TOKEN`. Narrow it with
  `shell_environment_policy` if your jobs handle untrusted content — and see the
  prompt-injection warning below.
- **Extra OS packages** without forking the Dockerfile:
  `--build-arg EXTRA_APT_PACKAGES="build-essential docker-cli"`.
- **`npm i -g` fails on purpose.** npm's prefix stays at `/usr/local`, which the
  runtime user cannot write, so nothing installed at runtime can shadow a pinned
  CLI — `codex update` would otherwise land an unpinned copy earlier on `PATH`.
  Use `npx` for one-off Node tools, and `uvx` / `uv tool install` (which target
  `~/.local/bin`) for Python ones.
- **Logs are JSON by default** in the image (`RINDO_RUNNER_LOG_JSON=1`) so a log
  collector can read them; set it to `0` for the human-readable form.
- **`tini` is PID 1** — it reaps grandchildren an ended job leaves behind and
  forwards `SIGTERM`, so `docker stop` runs the runner's graceful shutdown (the
  in-flight job's whole process tree is killed and reported failed, exit 0);
  `docker kill -s QUIT` drains instead (1.1 — see Shutdown).
- **Redistribution.** The image bakes in third-party CLIs under *their* licenses
  — Claude Code ships under Anthropic's commercial terms, not an OSS license.
  Build locally or push to a **private** registry; only `rindo-runner` itself is
  Apache-2.0. (This is why the image is absent from the repo's image-publish
  workflow.)
- **Containment is not a sandbox.** The container isolates the filesystem, not
  the agent's judgement — the prompt-injection warning below applies unchanged,
  and `--dangerously-skip-permissions` inside a container is still an agent
  acting with your token.

## Configure

Provide three things — **server URL, runner token, command template** — via any of
(highest precedence first) a CLI flag, an environment variable, or a config file:

| Setting       | Flag            | Env var                      | Config key    |
| ------------- | --------------- | ---------------------------- | ------------- |
| Server URL    | `--server-url`  | `RINDO_RUNNER_SERVER_URL`   | `server_url`  |
| Runner token  | `--token`       | `RINDO_RUNNER_TOKEN`        | `token`       |
| Command       | `--command`     | `RINDO_RUNNER_COMMAND`      | `command`     |
| Poll timeout  | `--poll-timeout`| `RINDO_RUNNER_POLL_TIMEOUT` | `poll_timeout`|
| Log level     | `--log-level`   | `RINDO_RUNNER_LOG_LEVEL`    | `log_level`   |
| JSON logs     | `--log-json`    | `RINDO_RUNNER_LOG_JSON`     | `log_json`    |
| Skills sync   | *(no flag)*     | `RINDO_RUNNER_SKILLS`       | `skills`      |
| Job image¹    | `--image`       | `RINDO_RUNNER_IMAGE`        | `image`       |
| Env allowlist¹| *(no flag)*     | `RINDO_RUNNER_PASS_ENV`     | `pass_env`    |
| Job server URL¹| *(no flag)*    | `RINDO_RUNNER_JOB_SERVER_URL` | `job_server_url` |
| Engine¹       | *(no flag)*     | `RINDO_RUNNER_ENGINE`       | `engine`      |
| Executor¹     | `--executor`    | `RINDO_RUNNER_EXECUTOR`     | `executor`    |
| Harness step² | *(no flag)*     | `RINDO_RUNNER_HARNESS`      | `harness`     |
| Fix passes²   | *(no flag)*     | `RINDO_RUNNER_HARNESS_FIX_PASSES` | `harness_fix_passes` |

¹ shared runner only (a `rndw_` credential) — refused beside a legacy `rndr_`
token, exit 2; see [Shared runner](#shared-runner-rndw_-11). `executor` is
`container` (the default) or `process` — the latter serves an EPHEMERAL
registration only and refuses the four container keys beside it; see
[Ephemeral runner](#ephemeral-runner-executor--process-12).
² the harness step (1.3): `harness` is ON by default for a shared credential and
OFF for a legacy token (opt in with `RINDO_RUNNER_HARNESS=1` / `harness = true` —
an explicit on under Windows on the process path exits 2: the step is POSIX-only
there); `harness_fix_passes` is 0..3 (default 1); see
[Harness step](#harness-step-harness-13).

Skills sync is **on by default** (set `RINDO_RUNNER_SKILLS=0` or `skills = false`
to disable — an exported-but-**empty** `RINDO_RUNNER_SKILLS=` is *not a value*:
per the runner's uniform env semantics it falls through to the file/default,
i.e. stays ON; the `0` is required): before each job spawns, the runner downloads the project's published
skills (delivered as short-lived signed URLs on the job) into `~/.claude/skills/`
so the agent CLI discovers them natively. It only ever touches directories carrying
its own `.rindo-skill.json` marker — a skill directory you placed by hand always
wins — and an unchanged skill costs zero downloads. A sync problem never fails or
delays a job; it is a warning in the runner's stderr.

Get a LEGACY runner token from **Settings → Agents** in Rindo (create/rotate/
revoke a token for your agent) — it starts with `rndr_`. A SHARED runner's
credential starts with `rndw_` and comes from **Account → Runners & pools**
(register a runner into a pool; the reveal shows it once, together with this
config file and the two commands below). A `rndw_` is accepted from the
environment or the config file ONLY — `--token` refuses it with exit 2, because
the command line is readable by every user on the host.

The config file is TOML. Default location:

- Linux/macOS: `$XDG_CONFIG_HOME/rindo-runner/config.toml` (i.e. `~/.config/rindo-runner/config.toml`)
- Windows: `%APPDATA%\rindo-runner\config.toml`

Override with `--config PATH`. The file holds your token — **`chmod 600` it.**

```toml
# ~/.config/rindo-runner/config.toml
server_url = "https://rindo.example.com"
token = "rndr_xxxxxxxxxxxxxxxxxxxxxxxx"
command = "claude -p --dangerously-skip-permissions {payload_file}"

# a SHARED runner (rindo-runner 1.1): a rndw_ instead, plus the job image —
# token = "rndw_xxxxxxxxxxxxxxxxxxxxxxxx"
# image = "rindo-agent-runner:local"
# pass_env = ["ANTHROPIC_API_KEY"]
# an EPHEMERAL runner (rindo-runner 1.2, started by a provisioner inside a
# sandbox): the same rndw_ with executor = "process" and NO image / pass_env
```

Then verify auth without claiming a job:

```sh
rindo-runner --check      # exit 0 = authenticates; 3 = rejected; 1 = unreachable; 2 = config
```

`--check` reads `GET /api/v1/runner/self` and prints the registration — agent and
project for a legacy token; pool, runtime, kind and credential expiry for a shared
runner — so you confirm the registration BEFORE the runner serves. Any
authenticated call bumps a last-seen stamp (a legacy agent reads "online"; a
shared runner's pool reads *ready* for a few minutes) and a shared runner's call
spends its pool's shared request window. Against a Rindo older than runner pools
(no `/self`) a legacy token falls back to the old job-0 probe — only a 404
`job_not_found` counts as authenticated — and a shared credential exits 2.

Run it:

```sh
rindo-runner                 # poll forever
rindo-runner --once          # handle one job (or one empty poll) then exit
rindo-runner -v --log-json   # debug logging as JSON lines
```

## The command template

The template is split into arguments **once** (shell-style quoting), and your
command is always run **without a shell** — so job content can never be interpreted
as a shell command. Exactly three placeholders are substituted, and only these:

| Token            | Becomes                                             |
| ---------------- | --------------------------------------------------- |
| `{payload_file}` | absolute path to a temp file holding the job JSON   |
| `{server_url}`   | your configured server URL                          |
| `{job_id}`       | the delivered job's id                              |

If your template contains **no** `{payload_file}`, the job JSON is fed to your
command on **stdin** instead.

Your command also receives these environment variables — this is how the spawned
agent connects back to Rindo over MCP **as the same agent**:

| Env var                | Value                                   |
| ---------------------- | --------------------------------------- |
| `RINDO_RUNNER_TOKEN`  | the runner token (for the MCP session)  |
| `RINDO_MCP_URL`       | `<server_url>/mcp`                       |
| `RINDO_SERVER_URL`    | the server URL                          |
| `RINDO_JOB_ID`        | the job id                              |
| `RINDO_PAYLOAD_FILE`  | the payload temp-file path              |
| `RINDO_SKILLS_DIR`    | where skills materialized (only when a sync ran) |

Inside a job container (shared runner) `RINDO_PAYLOAD_FILE` and `RINDO_SKILLS_DIR`
are the CONTAINER paths, and `RINDO_RUNNER_TOKEN` is the job's own `rndj_` token —
never the runner's `rndw_`. An ephemeral runner's child (1.2) gets the same `rndj_`
under the same name, with the sandbox's own paths.

The credential is **deliberately never** a template placeholder — putting a
secret in a command line would leak it to the process list. It travels in the
environment only; on the shared path the payload file carries no `job_token`
either (the child already holds it in env).

The job JSON looks like:

```json
{
  "id": 123,
  "trigger": "assignment",
  "entity_type": "task",
  "entity_id": 45,
  "project_id": 7,
  "payload": { "project_id": 7, "task": { "id": 45, "key": "PLX-3", "title": "…" } },
  "started_at": "…", "heartbeat_seconds": 60, "deadline_at": "…",
  "skills": {
    "bundles": [{ "name": "deploy-runbook", "url": "/api/v1/assets/…?exp=…&sig=…",
                  "sha256": "…", "size": 4096 }],
    "expires_at": "…"
  },
  "harness": [{ "binding_id": 3, "library_key": "SDLC", "version": 4,
                "schema_version": 1, "hash": "…", "harness": { "sensors": [] },
                "omitted": false }]
}
```

`skills` is the SK2 decoration (may be `null` on older servers): the runner
consumes it BEFORE spawning your command — by the time the agent runs, the files
are already on disk. Since HP2 the set is the project's **resolved** set — its own
published skills plus the skills it inherits from bound packs — under one flat
directory namespace (the server resolves collisions; a local skill always keeps
its name). On a shared home, a directory another project's runner already
materialized with the same bytes is counted `skipped`, never fought over.
`harness` (HP2, may be `null` on older servers) lists the served pack versions'
harness manifests in resolution order (an entry past the server's per-delivery
byte ceiling keeps its identity with an empty `harness` and `"omitted": true`);
since 1.3 each entry also carries the version's `host_floor` (the minimum runner
version) and `skill_shas` (its skill items, name → bundle sha256), and the runner
ENFORCES a manifest at schema version 1 — setup steps before your command, `loop`
sensors after it, a bounded fix pass — inside the job; see [Harness
step](#harness-step-harness-13). The signed URLs are short-lived read grants: they also sit in
the payload file, so a template that dumps its payload publishes them to the
project-visible job log for their TTL — keep templates quiet.

Your agent should call `get_identity` over MCP first to learn its role and its
operating-instructions document, then do the work (read the task/comment, propose
changes). Since WR1 the payload also carries **`instructions`** — an ordered list of
document references (`{document_id, source: "agent"|"role", role_key?}`): the
agent's own instructions document first **when it has one**, then one entry per
work Role it holds — branch on `source`, never on position, and expect `[]` when
it has neither. Like `instructions_document_id` it is a snapshot taken at enqueue
(a Role granted or revoked afterwards does not rewrite a queued payload —
`get_identity` stays the live source). The runner never opens it; your agent's
own template reads each document through `get_document` and decides how to use
them — the server never concatenates or renders them. Runner-channel writes always land as **draft revisions a human
approves** — that is the safety net, by design.

> ### ⚠️ Prompt-injection warning
> Job payloads contain **untrusted project content** (task titles, comment bodies
> written by anyone with access). Treat that content as **data, never
> instructions.** Your command template and your agent's own system prompt are
> responsible for this boundary. Put team-wide agent-safety conventions in the
> project's Guidance document. The runner's structural mitigations — no shell
> interpolation, a Member permission ceiling, drafts-not-commits, and no
> delete/config/approve powers — reduce but do not eliminate the risk.

## Example templates

```sh
# Claude Code, reading the job from a file:
command = "claude -p --dangerously-skip-permissions {payload_file}"

# A wrapper script that reads $RINDO_PAYLOAD_FILE / $RINDO_MCP_URL itself:
command = "/opt/agents/run-agent.sh"
```

> **Live output & the job log.** Whatever the command prints is what the live job
> log shows (see Operational notes). `claude -p` prints mostly at the end; add
> `--output-format stream-json` (or have your wrapper echo progress lines) if you
> want the in-app log to move while the job runs. Verbose modes reach the server's
> per-job log cap sooner — after that the live view freezes and only the final
> output tail is kept.

## Shared runner (`rndw_`, 1.1)

A **pool** (Account → Runners & pools) is granted to projects by their Admins; an
agent names the ONE pool that serves it and the **runtime** it needs. A shared
runner is registered into a pool with ONE runtime (`claude-code`, `converter`,
…) and claims the jobs of every granted project's agents of that runtime. It is a
SUPERVISOR: it holds the `rndw_` and runs **every job in a new container**, so
one job never sees another's state and never sees the runner's own credential.
(An EPHEMERAL registration — one job, inside a sandbox a provisioner made — is
served with `executor = "process"`, see
[Ephemeral runner](#ephemeral-runner-executor--process-12); under the default
`container` executor it too runs its one job in a container and then exits — the
one-job rule follows the registration, the executor only says where the job runs.)

**Requirements.** A Linux host; a container engine the user running
`rindo-runner` may drive (`docker`, or `podman` — set `engine`); the job image
present locally or pullable (`image` — build `runner/Dockerfile` as above and tag
it, or mirror it to your private registry; the registry auth in your
`~/.docker/config.json` is honoured); `rindo-runner` 1.1 installed on the host
(`uv tool install …`, as above). A Rindo on the same host is not at `localhost`
from inside a container — set `job_server_url` (e.g. `http://host.docker.internal:8000`
or the host's LAN IP) when `server_url` is loopback.

**Config.** `rindo-runner.toml`, `chmod 600`, then `rindo-runner --config
rindo-runner.toml --check` and `rindo-runner --config rindo-runner.toml` — the
reveal on Account → Runners & pools prints the same shape — the TOML in one well, the
`chmod` and the two commands in another, each with its own Copy:

```toml
server_url = "https://rindo.example.com"
token = "rndw_xxxxxxxxxxxxxxxxxxxxxxxx"          # env RINDO_RUNNER_TOKEN works too; never --token
image = "rindo-agent-runner:local"              # the JOB image — required to serve
command = "claude -p --dangerously-skip-permissions {payload_file}"
pass_env = ["ANTHROPIC_API_KEY"]                 # host env NAMES that cross into a job (none by default)
# job_server_url = "http://host.docker.internal:8000"   # the server as seen from INSIDE a container
# engine = "docker"                                       # or podman
```

**What a job container gets, and nothing else.** The image's entrypoint and your
`command` (`{payload_file}` is the container path); a CONSTRUCTED environment —
`RINDO_SERVER_URL`, `RINDO_MCP_URL`, `RINDO_RUNNER_TOKEN` (the job's `rndj_`),
`RINDO_JOB_ID`, `RINDO_PAYLOAD_FILE`, `RINDO_SKILLS_DIR` when this job synced,
`RINDO_RUNNER_SERVER_URL` for the image's Codex wiring, plus the names in
`pass_env` — never your shell's (a `pass_env` entry naming the runner's own
`RINDO_RUNNER_*`, a handle it constructs, or the engine client's `DOCKER_*` is
refused at load); at most two read-only mounts — the payload
(minus `job_token`; for a harness job the per-job directory holding it, the plan
and the driver) and the project's skills cache at `~/.claude/skills`; a fresh
HOME and `/workspace` in the container's own layer. The engine is driven under a
config derived from your `~/.docker/config.json` with its `proxies` REMOVED (the
CLI would otherwise inject `HTTP_PROXY` & co. into every container — pass them
through `pass_env` if a job needs them) and its context and registry auths KEPT,
so `pull` and `run` hit the same daemon; an env-var TLS setup (`DOCKER_HOST` +
`DOCKER_TLS_VERIFY`) keeps its certificates, since `DOCKER_CERT_PATH` is pointed at
your docker directory when unset and a `ca.pem`, `cert.pem` or `key.pem` sits there. On podman — decided
from the engine's own `--version`, so the `podman-docker` shim counts — the runner
also passes `--http-proxy=false`: podman would otherwise copy the host's proxy
variables into every container, around the allowlist (a `containers.conf` with
`env_host = true` or an `env` list would still add host variables; leave them
unset on a runner host). Before the first poll the runner reads
`GET /self` (401 → exit 3; a server without `/self` → exit 2; a transient answer →
backoff), then ensures the image is present (pulling on a miss) and pins it to its
ID, and removes job containers a crashed earlier run of this runner left behind;
an engine that does not answer, a client config it cannot derive, an image it
cannot get or a container listing it cannot read exits 2 (`executor unavailable`)
without claiming — and so does a stranded container of its own it cannot remove
while it may still be running: the next claim must never share the host with it
(remove it by hand, then restart; one that has exited — or, on podman, stopped — only warns). A runner
restarted after a crash sweeps its previous run's containers even when its
credential has since been retired (the runner id is remembered on disk, per
credential).

**Skills.** The supervisor syncs each project's published skills into a HOST
cache — `$XDG_CACHE_HOME/rindo-runner/skills/<server>-r<runner>/p<project>/`
(default `~/.cache/…`) — one directory per project, written by the runner alone
and mounted read-only into that project's jobs. Unchanged skills cost nothing on
the next job; a project's agents never see another project's skills.

**Failures.** A delivery whose runtime is not this runner's, or without a job
token, is reported failed as malformed and never executed. A container the engine
refuses to create or start fails THAT job (`runner error: executor …`), never
falls back to running in the supervisor's process, and the runner re-checks the
engine before its next poll — whether the engine refused the start or died
mid-job (exit 2 if it is gone). A job's exit status is
the container's own — `exit 125` inside is the command's 125, not an engine error.
Exit 3 means the RUNNER credential died (retired, expired, or its owner locked);
an agent's deactivation reaches a shared runner as the job's heartbeat 409, which
kills the job's container without a report.

### Ephemeral runner (`executor = "process"`, 1.2)

A **provisioner** (the reference one is `rindo-provisioner`, under `agents/provisioner/`
in the Rindo repository — decision #90, RP3) watches a pool's demand and, per job it
can serve, mints a ONE-JOB registration (`POST /api/v1/pool/runners` with the pool's
`rndp_` token — a `rndw_` whose `/self` says `kind = "ephemeral"` and carries a
claim-by `expires_at`), creates a sandbox from the agent-host image, and starts
`rindo-runner` inside it with the credential, the server URL and
`RINDO_RUNNER_EXECUTOR=process` in the runner's environment. Inside the sandbox the
runner:

- reads `GET /self` — it MUST say `ephemeral`; a persistent registration exits 2
  (`executor unavailable`) with zero polls — the process executor never serves a
  persistent `rndw_`, and there is no fallback to the container executor;
- polls as any shared runner (a 204 keeps polling: the TTL is the server's claim-by
  deadline, and an unused registration is refused at its expiry — the next poll's
  401 exits 3);
- runs its ONE job in its own process — the child's `RINDO_RUNNER_TOKEN` is the job's
  `rndj_`, the runner's `rndw_` never reaches it by name; skills sync into
  `~/.claude/skills` of the sandbox HOME; the payload file drops `job_token` as on
  the container path — then reports and **exits 0 without another poll**: the
  credential dies at the job's terminal (the server retires the registration), and
  the provisioner kills the sandbox at the runner's exit. A server-side end of the
  job — a cancel, a reaper end, a dropped grant — reaches the runner as a 401 at
  its next verb (the credential dies with the job) and exits 3 with the ephemeral
  copy; a 401 on the report after a natural end means the registration is spent
  (the report landed, or the job ended server-side) and the runner still exits 0.

The four container keys (`image`, `pass_env`, `job_server_url`, `engine`) are
refused beside `executor = "process"` (exit 2): the job runs in this process, so an
image or an allowlist would be a claim that never holds — what the job may see is
the SANDBOX's environment, which the provisioner's `sandbox_env` allowlist governs.
The isolation claim is the memo's P13, not P1: the runner and its one job share the
sandbox and its uid, so the `rndw_` is readable there — it can claim no second job
and reach no other project, and it dies with this one. Never serve a persistent
registration this way.

**Draining and retiring.** `SIGQUIT` to the SUPERVISOR process — `kill -QUIT <pid>`,
or under systemd `systemctl kill --kill-whom=main -s QUIT <unit>` (`--kill-who=` on
older systemd) — stops claiming, lets the current job finish untouched, and exits 0;
retire the runner on Account → Runners & pools AFTER that (retiring first ends the
credential at once and the running job as "lost contact"). The signal must reach
the supervisor ONLY: a cgroup-wide `systemctl kill` (its default, `--kill-whom=all`)
also hits the engine client the supervisor is waiting on — that client dies, the
job is reported failed (`runner error: executor lost: …`) and its container
removed. For the same reason a unit should set `KillMode=mixed`, so `systemctl
stop`'s `SIGTERM` reaches the supervisor alone and the supervisor stops the
container itself (10 s grace; with `uv tool install`, the unit's main process is
the runner itself). `SIGTERM` stays the hard stop. One supervisor per CREDENTIAL —
each runner credential has exactly ONE runtime, while a pool may mix runtimes
across its runners; scale by running more, each with its own `rndw_`.

## Harness step (`harness`, 1.3)

A pack bound to the job's project may carry a **harness manifest** (published by
the Library, checked by the server, version 1): **setup steps** to run before your
command and **sensors** to run after it — each a script shipped inside one of the
pack's own skills (`scripts/<name>.sh|.py|.js`), each sensor carrying the **fix
message** its author wrote. With `harness` on, the runner enforces it on every job
whose delivery carries one, the same way on all three shapes:

- A one-file **driver** becomes the job's command (your template rides it
  unchanged and is run by it; the driver is the supervisor's own file, so a job
  image needs only `python3` ≥ 3.11 — the documented image has it; an image
  without it runs its jobs as before, every entry that would otherwise run
  recorded `driver_unavailable`; no image rebuild). Your command sees nothing
  different but the payload file's path (it sits in the per-job directory, and
  each fix pass gets its own): the same arguments, environment, stdin, working
  directory and process group — the driver leads the group in its place. A
  command the shell cannot find exits 127 under the driver (1.2's process path
  reported it as a runner error).
- A step runs ONLY bytes the pack version pinned: the skill must be in this job's
  delivered set with the version's sha and its materialized directory must carry
  it (a same-named local skill, a hidden skill or a directory you placed by hand
  never runs; a same-sha directory another project's runner materialized on a
  shared home is the version's bytes and runs), the script a regular,
  non-symlink, non-hard-linked file under that skill's
  `scripts/`; it runs by its interpreter (`bash` / `python3` / `node`), no shell,
  with a CONSTRUCTED credential-free environment (`PATH HOME LANG LC_ALL TMPDIR`
  plus `RINDO_JOB_ID`, `RINDO_SKILLS_DIR`, `RINDO_HARNESS_PASS` — never the job
  token, the MCP URL or a `pass_env` value), its own process group and a timeout.
- **The fix loop** (`loop` sensors): when your command exited 0 and a sensor ran to
  completion and FAILED, the driver re-runs your command once more with a NEW
  payload file carrying `harness_feedback` (the failed sensors, each with its fix
  message), then re-runs the sensors — at most `harness_fix_passes` times per job
  (and never past the largest `max_fix_passes` a bound manifest declares; a pack
  declaring 0 never triggers a pass), only while the deadline allows, and only
  after a fresh heartbeat proved the job still runs. Exit 1 is FAIL whatever
  produced it (a Python or Node sensor's uncaught exception exits 1 — write
  sensors that exit 2 or more on their own errors); a sensor that exits 2 or
  more, a timeout, a cancelled job or an expired deadline never trigger a pass. **The
  same job id is no guarantee against repeated side effects:** the fix pass is a
  second run of your command over a result that already exists — its earlier
  comments, task updates and files already happened; the feedback says so, and
  your agent must treat `harness_feedback` as "fix, do not redo".
- **What you see:** the job's log ends with a `[rindo harness]` block (one line per
  pack and step; a failed sensor's fix message and the tail of its output), and
  the terminal report's `result.harness` carries the structured evidence (every
  entry and step with a closed reason when skipped, every run's status, exit code
  and output tail; `outcome` pass · fail · error · skipped; a large file is trimmed
  by the driver first — output tails, then fix messages — and marked `truncated`). **A sensor never
  changes the job's status** — succeeded or failed is your command's last run, as
  before; the sensors' verdict is recorded beside it.
- Time: a step starts only if its timeout fits before the job's deadline (minus a
  reserve for the report); the agent run itself is never cut by the driver;
  heartbeats keep flowing throughout.

## Operational notes

- **One job at a time.** The runner is single-threaded and serial: it finishes (and
  reports) one job before polling for the next. Run multiple processes with
  different tokens if you want more agents — or, on the shared path, more
  supervisors (one `rindo-runner` process per credential, each with its one
  runtime).
- **Skills land in `~/.claude/skills/` and persist across jobs** (legacy path — on
  the agent-host image that is the `claude-home` volume; a shared runner keeps the
  cache on the host, see above). Nothing is deleted per job — the sync
  converges: unpublished/renamed skills are pruned, unchanged ones are sha-skipped,
  and a very large first set may log `skills sync budget exhausted … owed=N` and
  start the job anyway (the next job finishes the remainder — not an error).
  Runners MAY share a `claude-home`: each managed directory's marker records the
  owning project, so one project's sync never prunes another project's skills,
  and the stray sweep leaves a live sibling process's in-flight state alone.
  ⚠ Sharing a home shares skill **visibility**: the CLI discovers every
  directory in the one flat `~/.claude/skills`, so every project's agent sees
  (and can follow) every project's skills — share a `claude-home` only among
  projects that trust each other's published content. Same-NAMED skills across
  projects contend for the one directory slot (first project in wins; the
  other reads `shadowed` — or `skipped` when the bytes are identical, the
  shared-Library-skill case).
- **Turning skills sync off freezes what is on disk.** The sync is the only
  thing that ever deletes a managed skill dir, so `RINDO_RUNNER_SKILLS=0`
  leaves previously-materialized (possibly since-unpublished) skills in place —
  remove them manually if the agent must stop seeing them.
- **Poll timeout.** The server holds an idle poll open for up to ~30s. The client
  read timeout (`--poll-timeout`, default 60s) must stay above that. If you raise
  the server's hold, raise this too.
- **Heartbeats & deadlines** are driven by values the server sends with each job
  (a beat cadence and a hard max-duration). If a job exceeds its max duration, the
  runner kills the command (and its whole process tree) and reports failure.
- **No automatic retries.** A failed job stays failed — re-mention or re-assign to
  try again. This matches the server contract; the runner never re-runs a JOB.
  (The harness step's fix pass re-runs your COMMAND inside the one job — one claim,
  one terminal report — see [Harness step](#harness-step-harness-13).)
- **Shutdown.** `SIGINT`/`SIGTERM` finishes gracefully: an in-flight job is
  terminated and reported as failed before exit, so job history stays truthful. An
  idle runner may take until the current poll returns (up to the poll timeout) to
  exit. **Drain (1.1, POSIX).** `SIGQUIT` stops claiming instead: a poll already in
  flight completes and a job it delivers is RUN (it is already claimed); the
  current job finishes with nothing shortened; then the runner exits 0 — an idle
  one within one poll hold. `SIGTERM`/`SIGINT` during a drain escalate to the hard
  stop. Windows has no drain signal.
- **Job output streams to the server (AGT-11).** While a job runs, the runner
  ships the command's merged stdout+stderr to the server in small chunks (~3s
  cadence) — project members can watch it live and read it later, up to a
  server-side size cap and retention window. Runner tokens are redacted before
  anything leaves the machine, but everything ELSE the command prints is visible
  to every project member: don't `env`, don't `set -x` with secrets in scope.
  Against an older server without the log route, the runner detects it once and
  streaming stays off for the session. The runner's OWN diagnostics still go to
  stderr only (human-readable, or `--log-json`).
- **Temp files.** Each job's payload and output go to files in the system temp dir
  (removed when the job ends). A hard crash can strand `rindo-job-*` files; the
  runner sweeps its own stale ones on the next start — and a shared runner likewise
  removes its own job containers whose supervisor died (by label).

## Exit codes

| Code | Meaning                                             |
| ---- | --------------------------------------------------- |
| 0    | clean shutdown, or a successful `--check`           |
| 1    | server unreachable during `--check`                 |
| 2    | configuration error — incl. a `rndw_` on `--token`, a missing `image`, an engine that does not answer (`executor unavailable`), a `rndw_` against a Rindo without `/self`, `executor = "process"` whose `/self` says the registration is persistent, an explicit `harness = on` under Windows on the process path |
| 3    | auth rejected — legacy: token revoked or agent deactivated; shared: the runner credential retired, expired or its owner locked |

## Develop

```sh
cd runner
uv sync
uv run ruff check . && uv run ruff format --check .
uv run pytest -q
```
