Metadata-Version: 2.4
Name: hyper-review
Version: 0.38.0
Summary: Multi-model read-only code review + AI panels.
Requires-Python: >=3.9
Description-Content-Type: text/markdown
Requires-Dist: pyyaml
Provides-Extra: test
Requires-Dist: pillow; extra == "test"
Requires-Dist: pytest<9,>=8; extra == "test"
Provides-Extra: live
Requires-Dist: telethon<2,>=1.34; extra == "live"
Provides-Extra: dashboard
Provides-Extra: rig-delegate

# review-cli

**multi-model code review from one command: diff review, cited quorum, brainstorm, visual review, and interactive spec-review tooling. CLI-first, harness-agnostic.**

> The review/quorum/brainstorm/just-ask/visual modes are **read-only** (the agents are
> caged — they cannot edit, run shell, or hit the network), with two explicit, narrow
> exceptions. The **`qa`** mode runs an **un-caged write/exec tester** that drives a
> System-Under-Test (see [QA — agent-as-tester](#qa--agent-as-tester-review-qa) for the
> safety model). And **`review diff --staged --task CODE --commit`** creates a checkpoint commit of
> the staged diff it just reviewed (opt-in via `--commit`; see
> [Diff review](#diff-review-review-diff) below). Don't assume every subcommand is
> read-only.

Runs your git diff through multiple AI backends **in parallel**, collects their findings,
and prints them side by side. Core review modes let you go from a quick pre-commit
sanity check all the way to a structured expert panel that builds consensus or explores
a design space. Built for use from any shell or AI agent harness (Claude Code, Codex,
opencode).

Beyond the core modes it also does **visual review** — attach a rendered screenshot to any
review with the composable `--visual` flag for a keep / rollback / repair verdict — and ships
**interactive spec-review tooling** so a markdown spec can be reviewed like a PR.

---

## Install

Pick one — each ends with `review` on your PATH:

```bash
# 1. uv — recommended; macOS, Linux and Windows (incl. Cygwin / Git Bash)
uv tool install hyper-review

# 2. pipx
pipx install hyper-review

# 3. one-liner — clones to ~/.local/share/review-cli, links review into ~/.local/bin, registers the agent skill
curl -fsSL https://git.hyperide.ai/ultrabricks/review-cli/raw/branch/main/install.sh | bash

# 4. from a clone — to hack on it; `git pull` updates the installed tool
git clone https://git.hyperide.ai/ultrabricks/review-cli && cd review-cli && ./install.sh
```

- After **1** or **2**, run `review install-skill` once so coding agents discover `review` (3 and 4 do it for you).
- **Unreleased `main`:** `uv tool install git+https://git.hyperide.ai/ultrabricks/review-cli` (or the same with `pipx install`).
- **Update:** `uv tool upgrade hyper-review` · `pipx upgrade hyper-review` · re-run the one-liner · `git pull`.
- **Installed before the rename** (as `review-cli`, from git)? Remove that first — `uv tool uninstall review-cli` or `pipx uninstall review-cli` — or two installs both claim `review`.
- **Via rig:** clone into `~/xp/review-cli` and list `review` under `tools.items` in `~/.config/rig/config.yaml` — `rig apply commit` runs its `install.sh` and keeps the clone fresh.
- **Windows:** use uv (option 1) — review runs natively under Cygwin / Git Bash (file locks via `msvcrt`, UTF-8 output on mintty). `./install.sh` and the one-liner do exactly that there. More: [rig-cli → Windows](https://git.hyperide.ai/ultrabricks/rig-cli#windows-cygwin-git-bash-powershell).

> **On PyPI as `hyper-review`** (the command is still `review`). The canonical repo is
> **git.hyperide.ai/ultrabricks/review-cli**; `github.com/alex-mextner/review-cli` is a frozen, archived mirror.

The `install-*` commands (`install-skill` / `install-commit-hook` / `install-hook tg` /
`register-module`) are idempotent — a re-run reports `already configured` (nothing
changed), `+ wrote/updated`, or `! conflict` (a foreign hook/symlink left untouched,
non-zero exit) so you know exactly what happened and can resolve a conflict and re-run.

---

## Quick start

Modes are **subcommands**: `review <mode> …`. The first verb selects the mode. A bare
`review` (no subcommand) prints the **help** — the diff review is `review diff`.
Every review iteration that is recorded in stats requires a task/issue code: pass
`--task CODE`, or set `REVIEW_TASK_CODE=CODE` once in the environment for automation.
Standalone `review visual IMAGE` is the exception because it is a single-image verifier,
not a multi-model review iteration.

```bash
# Review unstaged diff with your default backends
review diff --task HYP-742

# Review staged changes (pre-commit)
review diff --staged --task HYP-742

# Add backends to the defaults
review diff --task HYP-742 -m codex -m fable5 -m gemini

# Ask all backends a quick question (no diff needed)
review just-ask "Is a single-file Python CLI the right idiom for this tool?" --task HYP-742

# Settle a contested decision with cited evidence
review quorum "Should we cap brainstorm at 8 rounds?" --task HYP-742

# Open-ended design exploration
review brainstorm "How should we design the plugin system?" --task HYP-742

# Brainstorm ABOUT the current change (grounded in the working-tree / staged diff)
review brainstorm "Alternatives before I commit?" --task HYP-742 --diff
review brainstorm "Risks in this design?" --task HYP-742 --staged

# Save the result to a file — use -o, NOT `> file` (zsh noclobber-safe)
review diff --task HYP-742 -o review.md

# Later, inspect task-scoped iterations and transcripts
review task HYP-742
review task HYP-742 --detail 2
```

> **The diff review is `review diff` now.** A bare `review` (no subcommand) prints the
> help — it does **not** run a diff review (the old "bare review == a diff review" default
> was a mistake). The diff review moved from the stuttering `review review` to
> **`review diff`**. The removed `review review` verb and `review -C <repo>` (flags with
> no verb) print a one-line `review diff` pointer and exit non-zero. The meta flags
> (`--list-defaults` / `--show-board` / `--help`) still work with no subcommand.

> **Modes moved from flags to subcommands.** The old `--brainstorm` / `--quorum` /
> `--just-ask` flags are gone — use the `brainstorm` / `quorum` / `just-ask`
> subcommands. The flags now print a one-line pointer and exit non-zero. Visual review is
> **`review visual IMAGE`**; `--visual` remains a composable attachment for text modes
> such as `review brainstorm "is this good?" --task CODE --visual IMAGE`.

> **Write to a file with `-o file.md`, not `review … > file.md`.** Under zsh
> `noclobber` (a common default), `> file.md` refuses to overwrite an existing file
> and the command dies silently — no review, no error. `-o` writes the result with
> Python (`open(...,"w")`), bypassing the shell redirect entirely: it creates parent
> dirs, always overwrites, and still prints to stdout. See [Subcommands & flags](#subcommands--flags).

> **Deep help: `review help <topic>`.** Beyond `review --help` / `review <mode> --help`,
> `review help config` (alias `review --help config`) prints the configuration reference —
> the config file + cascade, the model/board selection, and keys/auth. The main `--help`
> lists the available topics.

> **Task history: `--task CODE` is required for review modes.** `review` stores the code
> in run stats and per-call logs, so `review task CODE` and the dashboard can show how many
> iterations were run for that task, which models participated, and the detailed
> conversations when logs are still present. Use `$REVIEW_TASK_CODE` in hooks/agents when
> repeating `--task` on every command would be noisy. `CODE` is one non-whitespace token,
> max 120 characters, with no control characters; examples: `HYP-742`, `review-cli-108`.

---

## Modes

### Diff review (`review diff`)

![review mode](docs/mode-review.svg)

N backends review your diff in parallel — one pass, no moderator. Best for pre-commit
checks where you want fast, independent perspectives without ceremony. The diff review is
the **`diff`** subcommand (`review diff`); **a diff is required**.

```bash
review diff --task HYP-742
review diff --staged --task HYP-742
review diff --staged --task HYP-742 --commit
git show --format= --no-ext-diff HEAD | review diff --task HYP-742 -m gemini,codex
```

> **Never use `git reset --hard` to discard a bad attempt mid-review** — a review→fix→
> re-review loop that resets hard can destroy unrelated uncommitted work from other
> sessions/agents sharing the same checkout (it has happened). Use `git checkout --
> <file>` to discard specific files, or `review diff --staged --task CODE --commit` (below)
> to checkpoint progress instead — undo a bad checkpoint safely with `git reset --soft
> HEAD~1`, which does not touch untracked/foreign files.

**Diff size cap.** The auto-detected diff is capped at 300,000 bytes by default
(`$REVIEW_DIFF_MAX_BYTES`, `<= 0` disables it) at DISPATCH time, before it is sent to
every backend — for `review diff`'s own two dispatch paths, and centrally for
`brainstorm`/`quorum`/`just-ask` too. `brainstorm` in particular auto-probes the
working-tree diff *by default* (no `--diff` needed) and sends it to every persona every
round plus the moderator — an oversized diff there was the worst token-burn multiplier
the 2026-08 investigation found, roughly an order of magnitude worse than the single
`review diff` panel. A real 6.5MB / 583-file diff (a debug harness's screenshot/video
capture scripts touching hundreds of files) was found being sent whole to every seat —
git already collapses each binary file to a one-line stub, so the cost is oversized
*text*. An over-cap diff is truncated with a visible marker naming the real total, not
silently dropped — scope the review (`git diff -- <path>`) or raise the cap to see the
full change. A diff piped on stdin (`git diff | review diff ...` / `... | review
just-ask ...`) is NEVER capped, in any mode — the user already explicitly, and
deliberately, scoped what they piped in. The cap never touches the canonical diff
`review diff --staged --commit`'s checkpoint re-derives to verify its integrity — but a
STAGED diff big enough to actually get truncated for dispatch REFUSES the checkpoint
(exit `EXIT_COMMIT_DIFF_TRUNCATED`) and skips the plain `--staged` commit-gate stamp too,
rather than certifying a partial review as "the full diff was reviewed": a stderr warning
alone doesn't protect non-interactive automation from silently checkpointing a diff no
seat fully saw. Scope the review or raise the cap, then re-run.

**Checkpointing a multi-round fix loop (`--commit`).** An agent iterating review → fix
findings → re-review may need several attempts, and a bad attempt needs a SAFE way back —
not `git reset --hard`. `--commit` (requires `--staged`) creates a real `git commit` of the
staged diff right after the review completes, so a bad next attempt can be undone with `git
reset --soft HEAD~1` instead. It checkpoints the *reviewed* diff, not a *clean* one: a
review that reports open findings still gets checkpointed (the pool producing usable
verdicts is what gates it, not "zero findings") — that's intentional, the same rule the
existing `--staged` commit-hook stamp already follows. The checkpoint is a real commit, so
it runs the repo's own commit-msg/pre-commit hooks; if a hook rejects it, `--commit` fails
loudly with a distinct exit code rather than silently skipping the checkpoint. `--commit`
without `--staged` is a usage error (there is no unstaged/piped diff to checkpoint against).
A review is multi-minute, so `--commit` also re-checks the staged index right before
committing and refuses (same distinct exit code) if it drifted from what was actually
reviewed — it never commits changes another process/session staged in the meantime. This
is the recommended default for any review loop that might need multiple rounds.

**Structured findings (`--json`).** `review diff --json` emits a machine-parseable
findings envelope on stdout instead of the prose panel — a flat `findings` array
(`file`, `line`, `severity`, `confidence`, `verdict`, `source`) plus a
`summary.expected`/`reported`/`degraded` seat breakdown — additive and opt-in: the
default prose output above is completely unaffected when the flag is absent. Each
backend is additionally asked for a structured block on top of its normal prose
review; a seat whose response can't be parsed is reported as `degraded` rather than
failing the whole run. Exit `0` = no findings, `1` = findings present, `2` = no
seat's structured response was usable at all. See `review diff --help` for the full
field/exit-code contract.

**Live board (interactive terminals).** Running `review diff` by hand in a real terminal
auto-shows a live per-seat status table (role, model, status -- queued/running/retrying/
promoted-from-reserve/done -- and elapsed time) plus a scrolling event log for retry/
failover/reserve-promotion notices, instead of a silent wait followed by one block of
prose findings. It activates ONLY when both stdout and stderr are a real interactive
terminal AND curses can actually start a real screen -- a piped/redirected/agent-driven
invocation (every CI/automation caller) is completely unaffected either way. Disable it
explicitly with `--no-tui` or `$REVIEW_NO_TUI=1`; a broken/dumb terminal falls back to the
existing per-call log files + `[review-cli] ... still waiting ...` heartbeat lines on its
own, within a bounded timeout.

---

### Just Ask

![just-ask mode](docs/mode-just-ask.svg)

Send a plain question to all selected backends in parallel. Diff is optional — pipe
one in or add `--staged` to attach it as context. One pass, no moderator, results
printed side by side.

```bash
review just-ask "Does this change need a migration?" --task HYP-742
git diff | review just-ask "Is this safe to merge?" --task HYP-742
```

---

### Quorum

![quorum mode](docs/mode-quorum.svg)

Two-phase structured panel. **Phase 1:** every expert answers in parallel and must cite
concrete evidence (file/line/fact); if they lack an evidence base they must say
`INSUFFICIENT EVIDENCE` rather than guess. Each seat also reasons from an assigned
role/lens (the same persona pool `brainstorm` rotates through — pragmatic staff
engineer, security-paranoid reviewer, skeptical SRE, etc.), shown in the transcript as
`glm [Security-paranoid reviewer]`; when a distinct-model pool is scarce or some
models are near their usage limit, one model can fill several seats (`fable#1`,
`fable#2`, ...) and each of ITS seats gets a different lens, up to the size of the
persona pool (currently 6) — beyond that a model's lens can repeat. A model covering
multiple roles in parallel is a fully valid panel shape; each of its seats is a genuine,
undiscounted opinion in its own right. **Phase 2:** a moderator runs sequentially, reads
all expert answers, and emits a structured summary with three sections — QUORUM (points
of majority agreement with evidence), DISAGREEMENT / NO QUORUM, and ABSTAINED.

Use when a question has real stakes and you want cited consensus, not vibes.

**`--adversarial-check` — an opt-in refutation pass for ship-gate-critical
questions.** After Phase 2 reaches a CLEAN verdict (no blocking disagreement), this
spawns one more pass whose only job is to try to REFUTE "no issues found" — not
another vote, an explicit attempt to find what the panel missed. If it finds a
genuine problem, that's surfaced as a new finding (never silently discarded) and the
run exits non-zero. Skipped automatically when the panel already disagreed (nothing
clean to refute) or when the flag is absent, so it costs nothing on routine
questions — reach for it specifically before a merge/ship decision, not every
`review quorum` call:

```bash
review quorum "should we ship this?" --adversarial-check --task HYP-742
```

**Live logs & partial output.** Each backend call streams its output in real time
to a per-call log in the OS-standard per-user log dir — **macOS** `~/Library/Logs/review-cli/`,
**Linux** `$XDG_STATE_HOME/review-cli/logs/` (default `~/.local/state/review-cli/logs/`);
override with `$REVIEW_LOG_DIR`; files are private, mode 0600. Panel modes print the
log path to stderr at the start of each call, so you can `tail -f` it to watch a long
run progress instead of staring at a frozen terminal. Review/panel agent CLI calls use
`--timeout` as a **silence** timeout: if the backend writes stdout/stderr, the timer
resets; if it stays quiet for the idle window, the partial output captured so far is
still returned (with a `[review-cli] TIMEOUT after Ns without output]` marker and exit
124) rather than being thrown away. Normal review runs allow at least 20 minutes of quiet
thinking time for subprocess backends. REST calls still use their HTTP request timeout;
QA and vision calls keep wall-clock timeout caps. Advanced override:
`REVIEW_IDLE_TIMEOUT_SECONDS=N` sets the review/panel subprocess idle window; `0` disables
idle reap and uses wall-clock `--timeout`. Values under 60s stay exact for tests/probes;
otherwise the normal 20m floor applies when the env var is unset. Idle mode treats any
stdout/stderr as progress, including output inherited from child processes; the internal
review backstop is the hard guard for a chatty but otherwise wedged process tree.

**Heartbeats during a silent wait.** A live backend call can legitimately sit quiet
for a while (see the idle timeout above), and in-seat retry (`--retry`, default on)
can double that once for a timeout. To make sure a long silent wait always reads as
"alive, still waiting" rather than "hung", every STREAMED backend call — regardless of
mode, including the plain `review diff` path that does not print a `tail -f` log
pointer — prints a `[review-cli] <backend> round <n>: still waiting, Ns elapsed (...)`
line to stderr every 60 seconds once the call is actually RUNNING (the trailing note is
`no output yet`, `producing output`, or `produced output, now idle`). Override with
`REVIEW_HEARTBEAT_SECONDS=N`; any value `<= 0` (not just a literal `0`) disables
heartbeats entirely; a positive value below the poll granularity (0.5s) clamps up to
it. (Scope: this covers the review/panel seat dispatch path, `_run_streamed`, from the
moment its subprocess spawns — the one implicated in the incident that motivated it.
Two things it does NOT cover: a seat still queued on `REVIEW_MAX_CONCURRENCY`'s
concurrency cap, waiting for a slot before it can even spawn; and a handful of short,
bounded, non-seat subprocess calls elsewhere — e.g. `git worktree add` during QA
isolation setup — that use the plain, non-streamed runner.)

**Run stats & a startup ETA — never short-timeout `review`.** `review` is
multi-model and (for the panel modes) multi-round, so a run takes **minutes**, and a
short shell `timeout` around it kills the run before its synthesis. Every dispatch
appends a run-stats record and prints a one-line ETA to stderr keyed on
`(mode, pool_size)`, so you know up front roughly how long to wait instead of
guessing. See `review help runtime` for the run-stats JSONL schema, the exact ETA
line format, and the full read-that-line guidance.

**No external timeout — `review` carries its own internal ≤4h backstop.** Do not
put *any* external `timeout` on `review`: it is designed to run unbounded from the
outside. The only time bound is an **internal**, last-resort backstop of **≤4h** that
the binary arms itself (`reviewlib.backstop`) — a watchdog that force-terminates a
genuinely wedged run (exit `124`). So a healthy run never needs an external cap (it
finishes in minutes, far under the ceiling) and a stuck run can't run forever either.
`$REVIEW_BACKSTOP_SECONDS` can only **lower** that ceiling, never raise it past 4h.

```bash
review quorum "Should we cap brainstorm at 8 rounds?" --task HYP-742
git diff | review quorum "Is this diff safe to merge?" --task HYP-742 -m codex,gemini,fable5
review quorum "Should we switch to a plugin architecture?" --task HYP-742 --moderator gemini
```

---

### Brainstorm

![brainstorm mode](docs/mode-brainstorm.svg)

Iterative ideation loop. Each round assigns at least three distinct **rotating personas**
(Pragmatic Staff Engineer, Security-Paranoid Reviewer, Developer-Experience Designer,
Skeptical SRE, Product-Minded Architect, Cost-Conscious Perf Engineer) to your panel
backends in parallel. After each round a moderator summarizes and decides STOP/CONTINUE — but
**cannot stop before `--rounds`** (minimum and default: 3). `--max-rounds` (default 8)
is a hard cap. Ends with a full moderator synthesis: best ideas, tradeoffs, and a
concrete recommendation.

Use for genuinely open design questions where you want the discussion to build across
rounds rather than converge in one shot.

**Brainstorm about a specific change (`brainstorm` + a diff).** brainstorm is composable
with the diff. When there IS a diff — an uncommitted working-tree diff in `-C` (pass
`--diff`), a `--staged` diff, or a piped diff — the moderator sees it as constant
**grounding context** on every round, and personas see it on the first round of the
invocation (Alex, 2026-08-28: later rounds rely on the shared transcript instead of
re-reading the diff — cheaper, since personas rotate and the transcript already carries
forward what earlier rounds found), so you can brainstorm concretely ABOUT a change
instead of in the abstract. With **no** diff present it stays pure ideation, exactly as
before. The diff is optional: an absent diff or a non-repo `-C` degrades silently to
ideation.

```bash
# brainstorm grounded in the current uncommitted working-tree diff
# (the subcommand leads; -C and the other shared options follow it)
review brainstorm "Is this caching approach sound? What are the risks?" --task HYP-742 --diff -C <repo>
review brainstorm "Alternatives to this design before I commit?" --task HYP-742 --staged -C <repo>
git diff main... | review brainstorm "How else could we structure this?" --task HYP-742 -C <repo>
```

The whole conversation is also written **incrementally** to a single discussion log
(`<logdir>/<stamp>-brainstorm.md`, path printed to stderr at the start) — each round
and moderator decision is flushed as it lands, so a timeout or interruption leaves the
discussion-so-far on disk instead of losing everything that was only being held in
memory for the final print. That log is also what makes a crashed brainstorm
**resumable** — see [`review sessions`](#review-sessions--list--resume-brainstorm-sessions).

The growing transcript is fed to the **claude and codex** backends over **stdin**
(not a `-p`/argv argument), which removes review-cli's own argv overhead. Note the
ceiling isn't fully gone: `claude-p`'s inner `claude` exec re-argv's the prompt, and
the **opencode** backend's CLI only takes the message as argv — so a very large
transcript (~1 MB+) can still hit `ARG_MAX` on those paths. `_payload` prints a size
WARNING as it approaches the limit; keep `--max-rounds` and diffs reasonable.

```bash
review brainstorm "How should we design the plugin system?" --task HYP-742
review brainstorm "API shape for the cache layer" \
  --task HYP-742 \
  --rounds 5 --max-rounds 10 \
  -m codex,gemini --moderator gemini
```

---

### QA — agent-as-tester (`review qa`)

The first mode that needs a **write/exec** agent, not a read-only reviewer. `review qa`
brings up a System-Under-Test (SUT), drives it against **human-authored prose test
suites**, and reports bugs **with proof** (logs / exit codes / expected-vs-actual). It is
**report-only**: a found bug never fails the build (it prints findings and exits 0); only
"couldn't run the tester" / `BLOCKED` is non-zero, and `--strict` flips any finding to 10.

Suites live at `docs/tests/suites/*.md` (relative to the SUT). Each `*.md` is a suite;
each `## Case:` block is one case the tester must exercise and verdict PASS / FAIL /
BLOCKED. With **no** authored suite, qa fails the no-suites gate (exit 6) and teaches you
how to author one — a green qa run with zero cases is a lie.

```bash
review qa <sut> --task HYP-742 --suites docs/tests/suites/*.md       # default: claude tester, isolated worktree, 1 case
review qa <sut> --task HYP-742 --kind backend --max-cases 5          # cap the run (cost control); 0 = full suite
review qa <sut> --task HYP-742 --in-place                            # run in the SUT tree (riskier; opt-in)
REVIEW_QA_TESTER=codex review qa <sut> --task HYP-742                # use the codex write/exec seat instead of claude
```

**Safety — read this.** qa is the first review-cli mode that runs an **un-caged** agent
(bash + write, no permission gate — claude runs `--permission-mode bypassPermissions`, codex
`--full-auto`). It runs WITH its working directory set to a throwaway `git worktree` of the
SUT by default, so an agent that stays in its cwd writes only into a disposable tree. **But
the worktree is NOT an OS sandbox.** An un-caged shell with absolute paths can **read AND
write anywhere on the filesystem** (other repos, `~/.ssh`, system files) and reach the
**network** — the worktree only bounds the *default* working directory, not what the agent
*can* touch. The only real write/exec boundary would be a container/VM, which qa does not yet
provide. **Run qa only against SUTs and suites you fully trust** (a malicious suite file or
SUT README could prompt-inject the un-caged agent), prefer the (default) worktree over
`--in-place`, and treat it like handing a shell to an LLM. `--in-place` is refused over a
tree with uncommitted/unknown git state (for BOTH the claude and codex seats). **Single-seat**
(one tester driving one SUT; the panel/`--pool` are ignored for qa), with a **long timeout**
default (not the short chat-panel cap) and token/wall accounting in the report.

**Deterministic Tier-1 harnesses (no un-caged agent).** Two SUT shapes have a **deterministic**
path that needs NO write/exec agent — "send input → assert output" is mechanical once the SUT is
up, so it runs as plain Python, off the agent-cage blast radius and reproducible in CI with zero
model spend:

- **`--kind bot`** (with a `sut.bot` mock config): a hermetic fake Telegram Bot-API server. The
  bot polls the fake via `TG_API_BASE`; the driver injects synthetic `getUpdates` and asserts the
  captured `sendMessage` calls against each `## Case:`'s `Send:` / `Expect:` / `Expect-no:` /
  `Expect-silent` grammar.
- **`--kind web`** (with a `sut.web` config): a real headless browser. The harness boots the app's
  dev server (`sut.web.command`), health-gates it reachable at `base_url`, then drives it in
  Playwright/Chromium against each `## Case:`'s `Goto:` / `Click:` / `Fill:` / `Expect-text:` /
  `Expect-no:` / `Expect-url:` grammar, classifying PASS/FAIL with a screenshot on failure. The
  browser is heavy, so it is **gated behind `REVIEW_QA_PLAYWRIGHT=1`** — off (or with Chromium not
  installed) a web run is a clear `BLOCKED` with the install command, not a crash. (Install once:
  `pip install playwright && python -m playwright install chromium`.)

```yaml
# docs/tests/qa.yaml — web Tier-1
sut:
  kind: web
  web:
    driver: playwright
    base_url: http://127.0.0.1:8080
    command: [npm, run, dev]      # or any dev server; omit for an already-running base_url
    ready_path: /
```

```bash
REVIEW_QA_PLAYWRIGHT=1 review qa <web-sut> --task HYP-742 --kind web   # deterministic headless-browser run
```

Both emit the SAME `## QA RESULTS` contract the un-caged tester does, so the verdict→exit mapping
is identical. Both guarantee teardown (the bot/fake, the dev server, the browser). A bot/web SUT
WITHOUT the matching config falls back to the un-caged tester, whose prose runbook tells the agent
to stand a mock / drive the site by hand.

---

### When to use which

| Subcommand | Reach for it when... |
|------------|----------------------|
| `diff` | Pre-commit diff check — fast, parallel, no overhead |
| `just-ask` | Quick multi-model second opinion on any question |
| `quorum` | A contested decision that needs cited evidence to settle |
| `brainstorm` | An open design space you want to explore across multiple rounds (optionally grounded in a diff — pass `--diff` / `--staged` or have an uncommitted diff to brainstorm about a specific change) |
| `qa` | Acting as a tester: bring up a running system and drive it against authored `## Case:` suites, reporting bugs with proof (report-only) |

---

## `review sessions` — list / resume brainstorm sessions

Every `review brainstorm` run is persisted as a round-by-round discussion log
(`<logdir>/<stamp>-brainstorm.md`, written incrementally as each round lands). `review
sessions` reads those logs so you can **list** past brainstorms — including ones that
**crashed, were killed, or timed out** before the final synthesis — and **resume** an
interrupted one instead of starting over.

```bash
review sessions          # recent COMPLETED sessions (those that reached a synthesis)
review sessions -a       # ALL sessions, including dead/interrupted ones (no synthesis)
review sessions -s <id>  # RESUME: reload the transcript, continue the round loop, synthesize
```

**Listing.** Each row shows a short **session id** (derived from the log's UTC
timestamp, e.g. `20260616T013310`), the **status** (`completed` = has a Final synthesis /
`interrupted` = crashed before one), the number of **rounds** captured (`r3`), the
**timestamp**, and the **topic**:

```
Brainstorm all sessions (incl. interrupted) — newest first; resume with `review sessions -s <id>`:

  20260616T020000  [interrupted]  r2  2026-06-16 02:00 UTC  resilient retry policy
  20260616T013310  [completed  ]  r5  2026-06-16 01:33 UTC  how to cache the widget
```

By default (no `-a`) the list shows only **completed** sessions, newest first, capped at
the 20 most recent — the "recent finished work" subset. `-a`/`--all` adds the
dead/interrupted sessions and lifts the cap, so you can find a crashed run to resume.
Listing is **read-only**, so it is safe to run against a brainstorm that is still
in progress (it just parses as a shorter transcript).

**Resuming.** `review sessions -s <id>` does **not** start from scratch: it reloads the
prior transcript, **reuses the saved topic, panel, and moderator**, and continues the
round loop from `completed_round + 1` to the original `--max-rounds` (respecting the
min-rounds / moderator-STOP rules), then produces the final synthesis. The continued
rounds and synthesis are **appended to the same log**, so the resumed run is one
continuous session, not a new file. Override the saved panel/moderator with `-m` /
`--moderator`, point at a repo with `-C`, or cap per-call time with `--timeout`. If the
session crashed *after* the moderator already decided to STOP (but before the synthesis
was written), resume skips straight to the synthesis — it does not run extra rounds.

The original `--diff`/`--staged` grounding is **not** persisted in the log, so a resumed
grounded brainstorm would otherwise continue ungrounded. Pass `--diff` (working-tree) or
`--staged` to **re-attach** the current diff as grounding for the resumed rounds and
synthesis.

The `<id>` can be the short displayed id or any **unambiguous prefix**. Edge cases are
handled explicitly:

| Situation | Behaviour |
|-----------|-----------|
| Unknown id | Error (exit 2), suggests `review sessions -a` to list ids |
| Ambiguous prefix (two runs in the same second) | Error (exit 2), lists the full ids to disambiguate |
| Already-completed session | Refused (exit 2) with a message — pass `--force` to re-synthesize from the saved transcript |
| Zero usable rounds | Degrades to a fresh run over the saved topic (nothing to continue) |

---

## `review task` — task-scoped review history

Every recorded review mode requires `--task CODE` (or `$REVIEW_TASK_CODE`) so iterations can
be grouped by the external task/issue that caused them. The CLI history view reads the
append-only run-stats store for the authoritative iteration count and model list, then joins
against dashboard logs for the detailed conversations when those logs are still present.

```bash
review task                 # list all task codes seen in run-stats
review task HYP-742         # iterations, models, ok/fail counts, log session ids
review task HYP-742 --json  # machine-readable history
review task HYP-742 --detail 2
review task HYP-742 --detail sess-20260703T101500_123456
```

`--detail` accepts either an iteration number or a dashboard session id. It prints the
brainstorm discussion body when present, then each backend call transcript and stderr block.
If the stat record exists but the old per-call logs have been deleted, the iteration still
appears in the summary and the detail command reports that transcript logs are unavailable.
The `--json` shapes are public CLI output contracts: task listings contain `task_code`,
iteration counts, model/mode lists, timestamps, duration, and ok/fail counts; detail output
uses the dashboard session-detail schema (`session_id`, calls, errors, brainstorm, roles).

JSON top-level shapes:

| Command | Shape |
|---------|-------|
| `review task --json` | `{"tasks": [{"task_code": str, "iterations": int, "models": [str], "modes": [str], "first_ts": str, "last_ts": str, "duration_seconds": number, "ok_count": int, "fail_count": int}]}` |
| `review task CODE --json` | `{"task_code": str, "iterations": [run_stats_record], "sessions": [dashboard_session_summary]}` |
| `review task CODE --detail N --json` | `dashboard_session_detail` with `session_id`, `task_code`, `calls`, `errors`, `brainstorm`, and `roles` |
| `review task CODE --check --json` | `{"task_code": str, "passed_iterations": int, "total_iterations": int, "distinct_models_passed": int, "models": [str], "min_iter": int, "min_models"?: int, "min_models_source"?: "explicit" \| "fallback", "passed": bool, "error"?: str, "identity_verification": "ran" \| "disabled" \| "skipped_unresolvable", "verified_iterations"?: int, "unverifiable_iterations"?: int, "excluded_mismatched_iterations"?: int, "searched_key"?: str, "near_miss_keys"?: [str], "near_miss_counted"?: bool, "near_miss_ever_passed"?: bool, "mismatch_details"?: [{"iteration": int, "reason": str, "recorded_repo_id": str, "recorded_diff_files": [str] | null, "ts": str}], "mismatch_details_truncated"?: true, "stalled_models"?: [{"model": str, "reason": str, "remaining_seconds": int, "consecutive_failures": int}], "min_roles_suggestion"?: str, "roles"?: [str], "distinct_roles_passed"?: int, "min_roles"?: int, "min_models_advisory"?: str, "quorum_mode_fallback"?: str, "role_tracking_gap"?: str}` — self-merge-authority gate; only iterations whose run came back clean count toward `passed_iterations`/`distinct_models_passed` (see `--check`'s own help). `mismatch_details` is capped at 50 entries — `excluded_mismatched_iterations` is always the uncapped true count among PASSED iterations specifically (its scope predates and is unchanged by the bare-check zero-role fallback below — a non-passed iteration excluded as mismatched during the fallback's OWN, separate check is not reflected in this counter). `searched_key`/`near_miss_keys` (review-cli#389) appear only on a failed/empty check whose task code has a same-digits, different-`#` twin already recorded in the store (`105` vs `#105`) -- `near_miss_counted` (review-cli#431) is `true` only when at least one of the twin's iterations actually landed in the gated population above (so it already counts toward `passed_iterations`); `near_miss_ever_passed` is `true` when the twin passed at all, even if that pass was excluded above as a repo/diff mismatch (`near_miss_counted=false` with `near_miss_ever_passed=true` means exactly that) -- both are `false` together when the twin's only record(s) never passed. `stalled_models` (review-cli#221) is present only when the bar ISN'T met and names any ATTEMPTED model that's currently cooling down (an unavailable-sentinel response or a session-limit/usage-credits notice — the two chronic signals `seat_cooldown` records; a plain timeout does NOT currently start a cooldown, see the "Deliberately narrow" note above) — the same signal `stalled: <model> (...)` lines print in text mode. `min_models` (review-cli#246: now optional) appears only when `--min-models` was explicitly given, OR the bare-check zero-role fallback fired (see `quorum_mode_fallback` below) — `min_models_source` (review-cli#246 follow-up) accompanies it EVERY time, naming which of the two put it there, since the two cases now produce an otherwise-identical shape. `roles`/`distinct_roles_passed`/`min_roles` (review-cli#221, `--min-roles`) appear whenever a role-based floor is ACTIVE — either `--min-roles` was explicitly given, or NEITHER flag was given and the task is NOT eligible for the case-(a) legacy model-counting fallback (review-cli#246's new default, see below). Note (review-cli#252 correction): this is NOT the same as "has at least one role-tagged iteration" — it also covers case (b), a task with genuinely ZERO role-tagged iterations that STILL isn't eligible for the fallback (at least one iteration isn't recorded as predating role-tracking; see `role_tracking_gap` below), which produces this same `roles: []`/`min_roles` shape and simply fails on it. `min_roles_suggestion` appears only when `--min-roles` was NOT passed and the gate failed because `distinct_models_passed` fell short of an explicit `--min-models`. `min_models_advisory` (review-cli#246) appears whenever `--min-models` was explicitly given AND at least one iteration is recorded for the task AND the store was readable (met, not-met, or a diff-identity-mismatch denial all qualify) — a non-blocking nudge, never affecting `passed`; it is NOT added to any of the fail-closed shapes that carry NO real history to evaluate (invalid task code, unreadable store, zero recorded iterations, a floor-validation failure), and it never appears alongside `quorum_mode_fallback` (the two are mutually exclusive — the advisory is explicit-`--min-models`-only). `quorum_mode_fallback` (review-cli#246 follow-up, NARROWED in review-cli#252) appears only on the bare-check default path (neither flag given), when at least one of the task's recorded iterations (of ANY verdict — passed, failed, or unverdicted) is NOT diff-identity-mismatched against the current repo/diff, that non-mismatched set carries genuinely ZERO role-tagged iterations, AND every one of those iterations provably predates role-tracking (PR #246's merge) — it names why the gate fell back to distinct-model counting instead of the new role-based default. `role_tracking_gap` (review-cli#252) is the counterpart for the OTHER zero-role-data case: at least one non-mismatched iteration is NOT confirmed to predate role-tracking — either it genuinely ran AFTER role-tracking became possible, or its recorded timestamp is missing/malformed/unverifiable (fails closed the same direction, since an unverifiable date proves nothing about age) — and yet carries no role data (reviewed via quorum/just-ask/brainstorm/qa, or a role-less `review diff` run) — this does NOT get the model-counting fallback (it would let a task dodge the role-based gate forever by choice of review mode), so the gate stays role-based and fails on its own merits, with this key naming the specific gap and pointing at `review diff --task CODE` through a role-tracking board (or, if a role-less diff run is part of the gap, at configuring a reviewer board) as the fix; `quorum_mode_fallback` and `role_tracking_gap` never both appear (each covers a disjoint slice of the "zero role data" case). A task whose ENTIRE recorded history is mismatched (e.g. a task-code collision with an unrelated repo) does NOT trigger either — it keeps the pre-existing role-based fail-closed shape instead of a misleadingly-worded model-counting one. |

**`--min-roles N`** (review-cli#221, defaults changed in review-cli#246) switches the
SECOND half of the gate — normally "N distinct MODEL NAMES among the passed iterations"
(`--min-models`) — to "N distinct BOARD ROLES (architect/correctness/security/…) covered
by the passed iterations" instead.

**Explicit-vs-default semantics (review-cli#246):** both `--min-models` and `--min-roles`
are optional; the gate reacts differently depending on which were actually TYPED:

- **Neither flag given** (the true default): the bar is role-based, at the SAME numeric
  floor `--min-models` used to default to (3) — "at least 3 distinct board roles", not "at
  least 3 distinct models". This is the new default everywhere (Alex's direction: role-based
  counting by default, no default model-count floor). **Zero-role fallback (follow-up to
  review-cli#246, NARROWED in review-cli#252):** a task whose entire history is recorded as
  predating role-tracking (PR #246's merge, `_ROLE_TRACKING_CUTOFF`), or comes only from
  `quorum`/`just-ask`/`brainstorm`/`qa` (modes that never record roles at all) AND every one
  of those iterations is recorded as predating that same cutoff, has ZERO role-tagged
  iterations — treating that the same as "recorded some roles, just short of the floor" is a
  guaranteed hard fail even when the task genuinely has 3+ distinct MODELS that would have
  satisfied the old model-counting default. So whenever the store is readable, at least one
  recorded iteration is NOT diff-identity-mismatched against the current repo/diff, AND
  **every** iteration in that non-mismatched set is recorded as predating role-tracking, the
  bare default falls back to the OLD distinct-model count at the same floor instead — the
  `--json` payload gains a `quorum_mode_fallback` key (and text mode a matching `  note: ...`
  line) explaining that this was a fallback, PLUS a `min_models_source: "fallback"` key
  (`"explicit"` for a genuine `--min-models` request) as the direct, positive way to tell the
  two apart (`"min_models" in payload` alone no longer does). If instead at least ONE
  non-mismatched iteration is recorded AFTER the cutoff (or has a timestamp that can't be
  verified as older) yet still carries no role data — a task reviewed via `quorum`/
  `just-ask`/`brainstorm`/`qa` (or a role-less `review diff` run with no board config) AFTER
  role-tracking became available, just never through a role-tracking board — the fallback
  does NOT fire: silently granting it there would let a task dodge the role-based gate
  forever, permanently, just by never choosing a role-tracking review mode, not the one-time
  migration accommodation it's meant to be for genuinely old history. That task stays
  role-based (fails on its own merits, same "N/M distinct roles" ratio a role-count shortfall
  always prints) and gains a `role_tracking_gap` key (plus a text-mode `  hint: ...` line)
  naming exactly what's missing and what to run instead (`review diff --task CODE` through a
  role-tracking board) — `quorum_mode_fallback` and `role_tracking_gap` never both appear.
  Two more cases are deliberately EXEMPT from either outcome, keeping the pre-existing
  role-based fail-closed shape instead: a task with NO recorded history at all (never
  reviewed — a different problem with a different fix), and a task whose ENTIRE history is
  diff-identity-mismatched (e.g. a task-code collision with an unrelated repo — that task's
  role data is real, just for a different repo/diff, so a misleadingly-worded model-counting
  shape would be wrong). A task with even ONE role-tagged iteration of ANY verdict in its
  non-mismatched history is unaffected either way: it stays role-based and must pass or fail
  on that basis alone, never falling back just because the role floor isn't met. Because the
  fallback now depends on WHEN history was recorded, the gate is non-monotonic: a task
  currently passing via the model-counting fallback can flip to failing the moment one more
  role-less `quorum`/`just-ask` review is recorded against it — doing additional review work
  can move the gate from pass to fail, which is the intended incentive (stop taking the
  role-blind shortcut) but is worth knowing before an automated `gh ship` pipeline trips on
  it.
- **Only `--min-models` given**: reproduces the exact pre-#221 PASS/FAIL decision — a
  strict count of distinct model-name strings decides the gate. The `--json` payload
  itself gains one new key here (the `min_models_advisory` described below), so a
  strict-key-set consumer of the OLD shape still needs updating even though the gate's
  logic hasn't changed.
- **Only `--min-roles` given**: reproduces the exact pre-#246 `--min-roles` behavior — a
  count of distinct board roles decides the gate; `--min-models`/`distinct_models_passed`
  are still computed and reported for visibility, but never gate.
- **BOTH given explicitly**: an explicitly requested floor is now ALWAYS enforced — the gate
  requires `--min-iter` AND the model floor AND the role floor to all be met (an AND, not
  "whichever is passed governs"). This fixes a review-cli#221 round-3 review finding: a
  caller explicitly passing `--min-models 5 --min-roles 1` used to get `passed: true` from a
  single-model review, because `--min-roles` silently outvoted the explicit model floor —
  Alex's policy is "if a model limit is explicitly set, it must work". Whenever `--min-models`
  is explicit, the response also carries a non-blocking `min_models_advisory` note (both in
  `--json` and as a `  note: ...` stderr/stdout line in text mode) suggesting that role-based
  coverage is usually sufficient on its own — it never affects `passed`.

This exists because the board's shortage-resilience behavior (PR #207:
`select_pool_with_reuse`, review-cli#207) fills an otherwise-empty role by reusing an
already-picked model instead of shrinking the panel — each duplicated pass reviews under
its OWN distinct role, but `--min-models` still only counts it as the SAME model string it
already counted once. A task whose recent history is "2 real distinct models + 1 role-fill
pass that reused one of them under a third role" can never satisfy `--min-models 3`, even
though 3 genuinely distinct facets of the diff were reviewed — `--min-roles 3` counts that
correctly. Only a role that is an actual `REVIEW_ROLES` key counts (an unknown/typo'd role
string on a custom config `board:` entry is kept for its own diagnostics but never a
distinct FACET — the seat degrades to the generic prompt, so it earns no lens credit here).
When an EXPLICIT `--min-models` (with no `--min-roles`) fails, the text/`--json` denial
includes a `min_roles_suggestion` pointing at `--min-roles` — but ONLY when re-running with
`--min-roles` set to the SAME number, IN PLACE of `--min-models`, would actually pass: the
iteration floor is already met, the recorded history covers enough distinct roles, AND at
least 2 distinct models actually EARNED one of those roles (a model sharing a multi-seat
record with a valid-role seat, but whose own role was unknown/typo'd, doesn't count toward
that "2 distinct models" check — only a model whose own recorded role is a real
`REVIEW_ROLES` key does). A task with no recorded roles at all, one that's ALSO short on
`--min-iter`, or one whose qualifying role coverage traces back to a single model (a
monoculture — technically `--min-roles`-passable, but not what the hint should ever steer a
caller toward), never gets the hint. Role-mode text output also prints a secondary
`models: N distinct model (...)` line naming which models actually reviewed. On the MET
path this is skipped when `--min-models` was ALSO explicitly given, since that floor's own
count clause already lists the names inline; on the NOT-met path the count clause is always
bare numbers (no names, matching the pre-#246 convention), so the secondary line prints
whenever a role floor is active, REGARDLESS of whether `--min-models` was also explicit —
including on a diff-identity-mismatch denial, which still carries real model/role counts
alongside its error message — so a self-merge-authority audit trail never loses which
models actually reviewed just because roles (not models) governed the verdict. Only
iterations from a mode
that records per-seat roles (currently: the `review diff`/`review visual` board dispatch)
contribute a role — `quorum`/`just-ask`/`brainstorm`/`qa` iterations count toward
`--min-iter` as always, but never toward `--min-roles` — and a role only ever counts from
a genuinely PASSED iteration that diff-identity verification did NOT exclude (the same
verified-or-unverifiable gate `--min-models` already respects; see below).

`--check` also runs **diff-identity verification** by default: it resolves `-C`/cwd
to a repo id (the normalized `origin` remote, or a local path when there is no
remote) and the current diff's touched files, then EXCLUDES any recorded PASSED
iteration whose own repo differs, or whose touched files share nothing with the
current diff — instead of trusting a task-code match alone. This closes a real
incident class: task-code reuse (a typo, a shared parent-ticket convention, or an
accidental/naive substitution) that let one diff's real reviews silently count
toward a completely different diff's quorum. **Threat-model boundary:** the store
is a local, self-reported JSONL an agent can append to directly (as the test
suite does, deliberately, to simulate history) — this closes the "wrong string
still matches history that's actually unrelated" class of bug, not a
cryptographic guarantee against a FULLY malicious agent fabricating a fresh
record with spoofed `repo_id`/`diff_files` that happen to match. Nor is file-set
overlap a strong signal in a repo where every PR conventionally touches the same
file (a CHANGELOG, a lockfile) — a shared "always touched" file alone can make
an unrelated same-repo iteration look verified; this gate substantially narrows
the incident classes it targets, it does not eliminate every same-repo
false-positive (tracked: issue #214).
`identity_verification` is the MACHINE-READABLE form of whether verification
actually ran — `"ran"`, `"disabled"` (`--no-verify-identity`), or
`"skipped_unresolvable"` (`-C` didn't resolve to a real directory) — so a
`--json`-only caller (which never sees the stderr warning printed in the same
two skip cases) can assert on it directly instead of inferring "did
verification run" from whether the `verified_iterations`/etc. keys are present.
`unverifiable_iterations` covers history recorded before this field existed,
which still counts (fail-open only for "no data to check", never for a confirmed
mismatch). `--no-verify-identity` restores the old task-code-only behavior.

---

## `review stat` — per-harness/per-model usage + health report

Detailed breakdown parsed from the real per-call logs under `log_dir()`: calls/ok/fail
per harness (backend), a byte-size proxy for usage (the only cross-harness signal that
exists — see below), **real** token counts where a backend actually emits them, the
SKILL.md/MEMORY.md context-pollution rate for agentic seats, the Fable (board seat)
dispatch/failure pattern, and the largest individual calls recorded.

```bash
review stat                        # last 7 days (default), text report
review stat --days 30              # last 30 days
review stat --days 0               # all recorded history (can be slow on a long-lived install)
review stat --since 2026-08-01T00:00:00+00:00
review stat --harness codex        # narrow the per-harness TABLE to codex (other sections stay whole-window)
review stat --top 20               # list the 20 largest calls instead of 10
review stat --json                 # machine-readable
```

**Token/cost honesty, same posture as `review dashboard`'s Metrics panel:** review-cli has
never recorded real token/cost numbers in its stats store (`reviewlib.stats`) — only a
handful of REST backends (`z.ai`, `commandcode`, `gemini`, `openrouter`, and `claude` in
API mode) ever had a token count to begin with, because they alone get a structured JSON
response with a `usage` field; the four agentic CLI harnesses this command breaks out by
name — `oc` (opencode), `omp` (Oh My Pi), `codex`, and `claude` in its default CLI mode —
carry **no** token number in any call log review-cli writes today, because review-cli
invokes each of them in its plain human-readable-stdout mode. This is not because the
underlying tools have nothing to offer: verified live, `codex exec --json`, `opencode
run --format json`, `omp --mode json`, and `claude -p --output-format json` each emit an
exact usage object (the last two also emit a real cost in USD) — but that structured
mode *replaces* the tool's readable stdout wholesale, and review-cli's live-log tailing
plus its paywall/auth/context-pollution detection all currently scan that readable
stdout as prose. Wiring it in is tracked separately (review-cli#186), not silently left
unfixed. `cc` aliases to the
single `commandcode` backend (`reviewlib.config.MODEL_ALIASES`), not a distinct harness.
`--harness` accepts this and other common short aliases directly (`glm`/`zai`/`oc`/`cc`
all normalize to the underlying backend name — see `--harness`'s own help text), so you
don't need to know the exact internal spelling. `bytes` (the call log's file size) is
reported for every harness as the best available proxy; `tokens_real` is `true` only for
a harness/call where an exact count was actually parsed.

JSON top-level shape:

| Command | Shape |
|---------|-------|
| `review stat --json` | `{"log_dir": str, "since": str\|null, "call_count": int, "retry_event_count": int, "tokens_recorded_backends": [str], "harnesses": {name: {calls, ok, fail, running, bytes_total, bytes_avg, bytes_p50, bytes_p90, bytes_max, tokens_real: bool, tokens_prompt, tokens_output, calls_with_real_tokens, skill_md_calls, memory_md_calls}}, "models": {model_id: {calls, bytes_total, classes}}, "fable": {dispatch_attempts, cached_skips, paywall_sentinel_calls, failure_rate, retry_events, retry_event_reasons, cached_skip_retry_events_excluded}, "retry_events_by_kind": {kind: int}, "top_oversized_calls": [{backend, model, task_code, started, size_bytes, diff_git_files, binary_stub_files, skill_md, memory_md, path}]}` |

---

## `review entities` — harnesses/providers/models/accounts/effort-lenses overview

Read-only structural report of review-cli's own configured entities: which harnesses
(codex/claude/opencode/omp/commandcode/z.ai/gemini/openrouter) the active board actually
routes seats through, each seat's provider/role/effort/availability, the registered
Claude multi-account roster (`reviewlib.claude_accounts`), the effort/role fallback
lenses (`reviewlib.role_models.ROLE_FALLBACK_MODELS`), and a compact per-harness
call-health summary for the last 7 days (the same data `review stat` reads). It makes no
new external/network calls and prints no keys — every field is reused from data
review-cli already resolves for a normal run.

```bash
review entities                    # text report
review entities --json             # machine-readable
```

This is the Phase 0 (read-only list view) slice of the fuller interactive
entity/relationship browser with drill-down and live refresh tracked in issue 487.

---

## `--detach` / `review jobs` / `review status` / `review wait` — background jobs

`review` advertises no EXTERNAL timeout — a healthy run finishes in minutes, so agents
must not wrap it in a short shell `timeout`. But some CALLERS genuinely cannot block for
the whole run (a subagent whose own shell tool has a short foreground cap, a pre-commit
hook with a tight budget). `--detach` is the supported way to stop blocking WITHOUT
resorting to an external timeout or a bypass: it spawns the identical review as a
session-detached background process and returns almost immediately with a job-id.

```bash
review diff --staged --task HYP-1234 --detach
# [review-cli] detached job 20260718T211853-0db637ed started (pid 86989)
# [review-cli]   log:    ~/Library/Caches/review-cli/jobs/20260718T211853-0db637ed.log
# [review-cli]   result: ~/Library/Caches/review-cli/jobs/20260718T211853-0db637ed.result.txt
# [review-cli]   check:  review status 20260718T211853-0db637ed

review jobs                                   # list every recorded job, newest first
review jobs --json
review status 20260718T211853-0db637ed        # status + paths + a log tail
review status 20260718T211853-0db637ed --json
review wait 20260718T211853-0db637ed          # block until it finishes (for a human /
                                               # unbounded caller — NOT for a capped subagent;
                                               # poll `review status` instead)
```

The detached run is the SAME review — same backstop, same `-o`/quorum-stamp output, same
child-reaping on kill (`reviewlib.process.install_signal_reaper`, review-cli#160) — just
running in the background; a caller's own `-o FILE` is honored as the job's result file,
otherwise a default path under the job's cache dir is used. Job status is one of
`running`, `done`, `failed`, or `unknown-terminated` (the process died — crash, SIGKILL, a
reboot — before it could record its own terminal status). `--detach` is rejected for
`dashboard`/`spec-web`, which already have their own `start`/`stop` service lifecycle.

---

## `review dashboard` — local web dashboard (managed service)

A **managed service** (`run` / `start` / `status` / `stop` / `enable` / `disable`, via the
shared `agenttools_service` lib) that serves a single-page web app over every review-cli run —
chat logs, per-task/mode/model stats, and an overseer's free-text feedback/PR-ticket
annotations, all read from the existing per-call logs (no new run-record format).

```bash
review dashboard start      # background daemon; review dashboard status / stop
review dashboard enable     # OS autostart at login (launchd / systemd --user)
```

It binds **127.0.0.1 only** by default — the logs persist prompts/diffs that may
carry secrets — so it is not exposed on the network unless you pass `--host 0.0.0.0` (e.g.
to reach it over Tailscale).

See **[docs/dashboard.md](docs/dashboard.md)** for the full command/flag reference, the
autostart matrix, data sources, panels, and the annotation store format.

---

## `review visual` — visual verification

**Give it a screenshot; it judges keep / rollback / repair.** `--visual` is image-only
visual verification: pixels in → verdict out — no DOM, no page, no capture, ever; every
check runs on the **image** (pixel heuristics plus an AI-vision model), never a stylesheet.
The pipeline (`cvGate → local pre-classifier → AI-vision → policy engine`) treats the vision
model as a **witness**, not the judge — a deterministic policy engine outside the model makes
the final call.

```bash
review visual after.png                                                   # standalone verdict
review brainstorm "is this layout good?" --task CODE --visual after.png   # composable context
```

`--visual` is also composable on `brainstorm` / `quorum` / `just-ask` / `diff`, so those
modes' models see the screenshot as multimodal context instead of running this standalone
pipeline.

See **[docs/visual-review.md](docs/visual-review.md)** for the full standalone/composable
matrix, vision-backend selection, the module system (built-ins, per-project modules, and the
untrusted-module quarantine), the `tg --photo` pre-send hook, and the exit-code reference.

---

## `review spec-web` — multi-spec web reviewer (a persistent daemon)

**Render any markdown spec server-side and review it like a GitHub PR** — a bidirectional
channel between a human reviewer and the agent that launched it. Select text in the rendered
spec to leave a **question** or **remark** anchored to it; **Submit review** finalizes the
batch and delivers the structured review straight back to the launching agent (no copy-paste
export step). One persistent daemon (`review spec-web start`) serves every registered spec by
name at `/spec/<name>`, live-reloading a page in place when the `.md` changes on disk.

```bash
review spec-web add docs/specs/my-spec.md --agent ext    # register + print its /spec/<name> URL
review spec-web serve docs/specs/my-spec.md --agent ext --exit-on-submit   # add + block for one review
```

**Security.** Reads (spec, assets, comments, **drafts**) are open to anyone who can reach the
port — there are no secrets in a spec review; only figures the markdown *references* are
served (an unrelated file in the assets dir 404s), SVGs are served with a `sandbox` CSP so a
directly-opened one can't run script, and symlinked assets that escape the assets dir are
refused. Writes (post comment / reply / submit / import / draft autosave) are origin-guarded
against both CSRF and DNS rebinding: a write requires (1) the request's **Host** to be in the
allowlist — **loopback + this machine's Tailscale identity (discovered at runtime via
`tailscale status`, never hardcoded) + `$REVIEW_SPECWEB_ALLOWED_HOSTS`** (comma-separated) —
and (2) the **Origin/Referer** to match that Host (classic CSRF check), plus
`Content-Type: application/json` and a body-size cap. The Host allowlist is what stops a
rebound attacker hostname from posting same-origin. So to accept writes from a phone over
Tailscale, the Tailscale name/IP must be discovered or listed in
`$REVIEW_SPECWEB_ALLOWED_HOSTS` (e.g. `REVIEW_SPECWEB_ALLOWED_HOSTS=ultras-mbp.tailbfe8ea.ts.net,100.123.113.82`).
Note the daemon binds `0.0.0.0` (not loopback-only) by default — the Host allowlist above is
what makes that safe; see `--host` in the full reference to restrict it to `127.0.0.1`.

See **[docs/spec-web.md](docs/spec-web.md)** for the full command/flag reference, the
submit → agent stdout handoff framing, the tmux `--agent` delivery mechanism, drafts, layout,
rendering, persistence paths, and the seed/import JSON shape.

---

## Agent workflows

`review` earns its keep on a hard call: `review brainstorm "<the decision>" --task
CODE` runs multi-model, multi-round ideation; the agent posts the top options to
Telegram via [`tg`](https://git.hyperide.ai/ultrabricks/tg-cli) for a phone decision, and
for the closest calls builds rival approaches in parallel **git worktrees** to compare
for real before committing.

Before every commit, `review diff --staged --task CODE` is a multi-model gate,
optionally *enforced* with `review install-commit-hook` (bypass with `REVIEW_SKIP=1
git commit` or `--no-verify`); small trailing follow-ups on an already-reviewed
baseline are tolerated (`$REVIEW_TRIVIAL_DELTA_LINES`, default 10 lines) without a
full re-review — see `review help config` for the exact tolerance rules.

---

## Model backends

Each backend runs as a **`cli`** subprocess, a **`api`** REST call, or both:

| Specifier | Transport | What runs under the hood |
|-----------|-----------|--------------------------|
| `codex` / `codex:<model>` | cli | `codex exec -s read-only --ephemeral` |
| `claude` / `claude:<model>` | api \| cli | `claude-p` CLI, or the Anthropic-compatible Messages API |
| `fable` / `fable5` | api \| cli | Alias for `claude:claude-fable-5` |
| `gemini` / `gemini:<model>` | api | Gemini REST API (`gemini-3.5-flash` by default) |
| `zai:<model>` / `glm` / `glm52` … | api | z.ai (GLM) OpenAI-compatible REST API — needs `ZAI_API_KEY` |
| `commandcode:<model>` / `cc` | api | Command Code OpenAI-compatible Provider API — needs `COMMANDCODE_API_KEY` |
| `openrouter:<model>` / `openrouter` | api | OpenRouter OpenAI-compatible aggregator (400+ models) — needs `OPENROUTER_API_KEY` (bare `openrouter` → `openrouter/auto`) |
| `oc:<model>` / `opencode:<model>` | cli | `opencode run --agent read-only-reviewer --dir <repo>` (reads the real repo, read-only) |
| `oc:xai/grok-4.6` / `grok` / `grok46` / `grok-4.6` | cli | Grok 4.6 via opencode's native `xai` provider (oauth-authenticated via `opencode providers login`); GROK_SEAT, the default board's priority-12 seat. `grok45` pins `oc:xai/grok-4.5`. |
| `omp:opencode-go/mimo-v2.5` / `mimo` / `mimo25` / `mimo-2.5` | cli | MiMo 2.5 via omp OpenCode Go (`MIMO_SEAT`, the default board's priority-13 `tests` reserve). The `oc:opencode/mimo-v2.5-free` catalog spelling is not DEFAULT-safe (`default_routes_live` False: provider `opencode` under `oc:` is unnamed and not in `_AGENTIC_ONLY_PROVIDERS`). Explicit `-m oc:opencode/mimo-v2.5-free` still works at runtime. |
| `omp:anthropic/claude-sonnet-5` | cli | Claude Sonnet 5 through the omp harness (`OMP_CLAUDE_SEAT`, the default board's priority-8 `quality` reserve) — the SAME model as `claude:claude-sonnet-5` but a DIFFERENT auth path (omp's own profile-scoped agent dir/credentials, not the Claude Code CLI's `CLAUDE_CONFIG_DIR`/keychain), so it survives an outage hitting every `claude:*` seat's shared account. Rotates across the comma-separated `$REVIEW_OMP_CLAUDE_PROFILES` when set. |
| `omp:<provider>/<model>` / `omp` | cli | `omp -p --no-session --tools read,grep,glob --add-dir <repo> @<payloadfile>` (reads the real repo, read-only) |
| anything else | cli | Treated as an opencode model id |

**Transport split.** Each backend declares which transports it supports — `cli`, `api`,
or both — shown in the *Transport* column above. `REVIEW_<NAME>_MODE` forces one; forcing
a mode a backend doesn't support is a hard error, never a silent fall-through. (Today:
codex/opencode/omp are cli-only, gemini/z.ai/commandcode/openrouter are api-only, claude does
both and auto-picks — CLI if the binary is present, API when it isn't and a key is set.)

The opencode backend is **agentic and read-only**: it runs in the **real `-C`
repository** (via `opencode run --dir <repo>`), exactly like the codex backend, so an
`oc:` seat can **read any project file** — not just the diff in the prompt. Safety is
enforced by the `read-only-reviewer` agent, which **denies** `edit`/`write`/`bash`/
`webfetch`: opencode may open files but can never mutate the worktree, run a command,
or hit the network.

It falls back to an isolated temp dir (diff-only) in two cases:
- `-C` is **not a git repo** (e.g. a `just-ask` from a scratch dir) — nothing to read;
- the repo **ships its own opencode config** (`.opencode/` or `opencode.json`/`.jsonc`).
  A repo-local agent definition can **override** the global `read-only-reviewer` and
  re-enable `write`/`bash` (verified: project config wins, and no opencode env flag
  suppresses it), so to keep the sandbox trustworthy on a potentially adversarial repo,
  review refuses to run agentically there and reviews the diff in a clean dir instead.

The **omp (Oh My Pi) backend** (`omp:<provider>/<model>`, e.g. `omp:kimi-code/k3`) is
likewise **agentic and read-only**: it reads the real `-C` repository with the tool
set restricted to `read,grep,glob` (`--tools`), extension/skill discovery disabled
(`--no-extensions --no-skills`), and no session persisted (`--no-session`). Two
hardened boundaries, both verified live against omp v17 (review of review-cli#174):

- omp **executes project-shipped code from its launch cwd** (a repo's `.mcp.json`
  spawns its MCP server command; `.omp/tools/*.js` is imported at startup) and mounts
  **user-scope MCP servers** (`~/.claude.json` et al.) whose tools run arbitrary code,
  so omp is launched from a **neutral empty temp dir** with **HOME pointed at an empty
  subdir** (`PI_CODING_AGENT_DIR` pins omp's real agent dir so auth still resolves) and
  the repo is mounted read-only as a workspace via `--add-dir` — every project file
  stays readable, no project or user-scope code is ever executed.
- omp's `read` tool accepts **https URLs** (an outbound exfiltration channel) and the
  `xd://` device transport carries write/edit/bash **around `--tools`**, so a per-run
  `--config` overlay disables `fetch`, `tools.xdev`, and project MCP config. All three
  boundaries are covered by permanent LIVE assertions
  (`REVIEW_OMP_CAGE_LIVE=1 python3 tests/test_omp_cage_live.py`).

The prompt+diff is handed over as an `@<tempfile>` message arg — omp does not read
prompts from stdin, and the `@file` transport dodges the ~1 MB ARG_MAX ceiling
argv-passing would hit. The selector after `omp:` goes to omp's `--model` fuzzy
matcher verbatim. Availability is probed offline: the `omp` binary on PATH plus a
non-disabled credential row for the seat's provider in omp's own auth db
(`~/.omp/agent/agent.db`, honoring `PI_CODING_AGENT_DIR` / `OMP_PROFILE`) —
authenticated via omp's own setup (`omp setup`), never a review-cli key.

> **Note on commandcode / z.ai (review-cli#24).** These were historically kept as raw
> diff-only `api` backends — opencode's `@ai-sdk/openai-compatible` adapter did not
> reliably drive the Command Code gateway, and z.ai/GLM was not an opencode-native
> provider. That has been **re-investigated and resolved**: with `commandcode` and `zai`
> registered as opencode **custom providers** (`~/.config/opencode/opencode.json`, auth via
> `opencode auth login`), the default board's Kimi/GLM/Qwen/DeepSeek seats now run
> agentically through opencode (`oc:commandcode/…`, `oc:zai/glm-5.2`) like the rest of the
> board. The raw keyed-HTTP `commandcode:`/`zai:` backends remain for explicit `-m cc` /
> `-m glm` and config-board seats on hosts without opencode.

---

## Subcommands & flags

The mode is a **subcommand** (`review <mode> …`); the flags below are shared options
available to the relevant subcommands. A bare `review` (no subcommand) prints this help —
the diff review is `review diff`.

```
SUBCOMMANDS
diff                Diff review across the reviewer board (requires a diff).
visual IMAGE        Visual verification for a screenshot; add --diff to include git diff.
brainstorm TOPIC    Multi-round persona ideation; composable with --diff/--staged grounding.
just-ask QUESTION   Single-shot multi-model answer to a question (diff optional).
quorum QUESTION     Experts cite evidence + a moderator finds quorum/disagreement.
qa SUT              Agent-as-tester mode for authored QA suites.
dashboard           Local web dashboard over review-cli runs.
sessions            List / resume brainstorm sessions (-a all, -s <id> resume).
task [CODE]         List task-coded review iterations, models, and transcript details.
jobs                List detached (--detach) review jobs.
status JOB-ID       Show one detached job's status, paths, and a log tail.
wait JOB-ID         Block until a detached job finishes; exits with its status.
spec-web            Multi-spec web reviewer daemon (start/status/stop/add SPEC; also `spec-web SPEC.md`).
install-skill | install-commit-hook | install-hook tg | register-module

TOP-LEVEL / SHARED FLAGS (shown by `review --help`; subcommand help shows what applies)
-m / --model        Backend to run; repeat or comma-separate. Default (no -m) is mode-aware:
                    `review diff` runs the active default preset board — "light" as of
                    2026-08-28, see --preset — (or your config `models:`);
                    brainstorm uses `brainstorm_models:`, just-ask/quorum the defaults.
                    Each subcommand's `--help` shows its own effective default.
-C / --cwd DIR      Run against a different repository directory.
--task CODE         Task/issue code for this review iteration. Required for recorded
                    review modes; standalone `review visual IMAGE` is the exception.
                    Can also be supplied by $REVIEW_TASK_CODE for automation.
-o / --output FILE  Write the result to FILE via Python (creates parent dirs, always
                    overwrites) while still printing to stdout. Use this instead of
                    `review … > FILE`, which fails silently under zsh noclobber.
--timeout N         Per-call timeout in seconds. Review/panel agent CLIs use an idle/silence
                    timeout with a 20m default floor on normal review runs; REST calls use it
                    as the HTTP request timeout, and QA/vision keep wall-clock caps. Values
                    under 60s stay exact for tests/probes; REVIEW_IDLE_TIMEOUT_SECONDS
                    overrides the review/panel idle window when set.
--detach            Spawn this review as a session-detached BACKGROUND process and return
                    almost immediately with a job-id, instead of blocking the caller for
                    the whole run. See "review --detach / jobs / status / wait" below.
                    Not supported on `dashboard`/`spec-web` (they have their own
                    start/stop lifecycle).
--list-defaults     Print effective default backends and exit.
--show-board        Print the active reviewer board (model -> role + availability) and exit. Also prints per-role fallback model ids (bare names; harness auto-resolved).
--preset NAME       Diff-review and --show-board preset: light for quick preflight,
                    default for routine review, heavy for release/risky changes with
                    Sol (Fable is excluded from every preset — a confirmed ~100%
                    dispatch failure rate; it sits last-resort in the raw board
                    instead). Other subcommands reject it.
--pool N            How many of the selected preset/board's seats to run (default
                    preset-dependent: 4 for default/heavy, 2 for light; a recorded
                    `--task` run floors that to 3 so gh ship can pass); the first N seats run,
                    the rest are kept in reserve. The board is never off — --pool only sizes
                    it. N<=0 (e.g. --pool 0) runs all seats in the selected preset/board.
                    An explicit `--pool N` below 3 with `--task` still wins but warns.
                    Ignored for exact multi-model -m; `--task -m` one-model pads that model
                    across 3 roles. Recorded `--task --pool 0` still pads a same-role
                    config roster to the ship floor.
--effort LEVEL|PROVIDER=LEVEL
                    Run-scoped reasoning effort, overriding each seat's config effort for
                    THIS run. A bare level (minimal/low/medium/high/xhigh/max) applies to
                    every seat; PROVIDER=LEVEL (e.g. codex=high, opencode=max) overrides one
                    backend route; per-provider wins over the global level. Repeat or
                    comma-separate. Reaches the codex/claude/opencode reasoning levers plus
                    the screenshot vision call; falls back to the seat's config effort where
                    the flag is silent.

SUBCOMMAND-SCOPED FLAGS (shown by `review <mode> --help`, not the global list)
--diff / --staged   Diff source: working-tree (--diff) or staged (--staged). On the diff
                    review the diff is required; optional grounding for brainstorm.
--prompt TEXT       (review diff) Override the diff-review prompt.
--moderator M       (quorum / brainstorm) Override the auto-picked moderator.
--rounds N          (brainstorm) Minimum rounds before STOP is allowed (default 3).
--max-rounds N      (brainstorm) Hard cap on rounds (default 8).
--visual IMAGE …    Composable visual-verification group for text modes: attach/verify a
                    render; rides subcommands such as
                    `review brainstorm "Q" --task CODE --visual shot.png`.
                    Companions: --before/--intent/--expect/--check/--json/--strict/--no-ai/
                    --no-local-model/--vision-timeout/--project. For standalone use
                    `review visual IMAGE` or `review visual IMAGE --task CODE --diff`.
                    See `review <mode> --help`.
```

> **Modes are subcommands, not flags.** `--brainstorm` / `--quorum` / `--just-ask` were
> removed; use `review brainstorm …` / `review quorum …` / `review just-ask …`. The old
> flags print a one-line pointer and exit non-zero. The diff review is `review diff` (a
> bare `review` prints help; the removed `review review` verb points at `review diff`).

---

## Configuration

Personal defaults live in `~/.config/review-cli/config.yaml`:

```yaml
# Priority roster for `review diff`: the first available seats fill the live pool and the
# rest are reserve. just-ask/quorum use this same list as their flat default panel.
models:
  - codex
  - fable5

# Brainstorm can use a wider panel (falls back to `models` if absent)
brainstorm_models:
  - codex
  - gemini
  - fable5

# Visual review uses a separate priority list from the text reviewer board.
# Opus is tried first; GROK_SEAT is next; unavailable/unusable vision calls fall through.
visual_models:
  - claude:claude-opus-4-8
  - oc:xai/grok-4.6
  - oc:zai/glm-4.5v
  - oc:commandcode/moonshotai/Kimi-K2.7-Code
  - gemini

# Optional: skip every seat under providers whose subscription/billing is currently
# unavailable. Same as REVIEW_UNPAID_PROVIDERS=commandcode,fireworks.
# unpaid_providers:
#   - commandcode
#   - fireworks
```

Run `review --list-defaults` to see the effective (normalized) models after config is
applied.

Code defaults (when no config file exists): `codex`, `gemini`,
`commandcode:moonshotai/Kimi-K2.7-Code`.

---

## Reviewer board

The plain `review diff` runs a **reviewer board**: a panel where each model is given its OWN
review role/lens, so the panel covers the diff broadly instead of every model doing the same
generic pass. The board is the default panel out of the box — no config file required.

The raw built-in board is a **priority-ordered** list of 15 models, strongest WORKING model
first, chosen by **priority + availability**. A bare `review diff` runs the `light` preset (a
pool of 2 at medium effort) for cheap routine checks; `--preset default` steps up to a pool of
4 at high effort; `--preset heavy` runs Sol/Opus/GLM-cc/Kimi at `xhigh`, with the rest of the
board as `max`-effort reserve. Fable is excluded from every preset — chronic dispatch failures
demoted it to the board's final, last-resort seat. Two layers of failover keep the pool full:
**startup failover** skips an unreachable higher-priority seat and pulls up the next one, and
**mid-run failover** promotes the next-priority reserve if an active seat fails mid-review.
`--pool N` sizes the pool (`--pool 0` runs every available seat in the selected preset/board);
the board itself is never disabled.

The built-in board, in **priority order** (the `Tier` column shows the `heavy` preset split on
a fully-keyed environment):

| # | Tier | Reviewer | Backend | Role | Lens focus |
|---|---|---|---|---|---|
| 1 | pool | Sol | `codex:gpt-5.6-sol` | `consistency` | cross-file consistency, dead refs, contract drift, whole-repo coherence |
| 2 | pool | Opus | `claude:claude-opus-4-8` | `correctness` | logic bugs, regressions, edge cases, null/async/race, off-by-one (also the moderator) |
| 3 | pool | GLM-cc | `commandcode:zai-org/GLM-5.2` | `performance` | complexity, hot paths, allocations, async/concurrency, N+1 (GLM 5.2 via the Command Code gateway; diff-only, read-only by construction) |
| 4 | pool | Kimi | `oc:commandcode/moonshotai/Kimi-K2.7-Code` | `quality` | readability, naming, duplication, code smells, idiom |
| 5 | reserve | Astra | `codex:gpt-6-astra` | `consistency` | cross-file consistency, dead refs, contract drift, whole-repo coherence (GPT-6-Astra, OpenAI's flagship codex model; explicitly pinned, unlike a bare `codex` seat which would silently track the CLI's own default; duplicates Sol's lens — a pre-existing, deliberately accepted trade-off, see review-cli#382) |
| 6 | reserve | Terra | `codex:gpt-5.6-terra` | `performance` | complexity, hot paths, allocations, async/concurrency, N+1 (a distinct, already-paid codex model — live `performance` fallback for GLM-cc, review-cli#382) |
| 7 | reserve | Sonnet | `claude:claude-sonnet-5` | `quality` | readability, naming, duplication, code smells, idiom (a distinct, already-paid Anthropic model — live `quality` fallback for Kimi, review-cli#382) |
| 8 | reserve | Sonnet-omp | `omp:anthropic/claude-sonnet-5` | `quality` | readability, naming, duplication, code smells, idiom (the SAME claude-sonnet-5 model through the omp harness — a different auth path than the direct `claude:` seat right above, so it survives a Claude-account-wide outage that takes out every `claude:*` seat sharing one account, review-cli#168) |
| 9 | reserve | Qwen | `oc:commandcode/Qwen/Qwen3.7-Max` | `security` | injection, authz, secrets, unsafe deserialization, path traversal, SSRF |
| 10 | reserve | DeepSeek | `oc:commandcode/deepseek/deepseek-v4-pro` | `tests` | missing tests, untested branches, boundary conditions, error-path coverage |
| 11 | reserve | Gemini | `gemini` | `contracts` | public API shape, contracts, types, backward-compat, interface design |
| 12 | reserve | Grok | `oc:xai/grok-4.6` | `security` | injection, authz, secrets, unsafe deserialization, SSRF, path traversal (opencode's native xai provider, oauth-authenticated; agentic, no diff-only REST fallback exists; placed directly before the slow GLM seat so failover reaches it first, review-cli#165/#415) |
| 13 | reserve | MiMo | `omp:opencode-go/mimo-v2.5` | `tests` | missing tests, untested branches, boundary conditions, error-path coverage (OpenCode Go via omp; live paid `tests` fallback because DeepSeek is commandcode and often unpaid; `oc:opencode/mimo-v2.5-free` is not DEFAULT-safe, review-cli#450) |
| 14 | reserve | GLM | `oc:zai/glm-5.2` | `security` | injection, authz, secrets, unsafe deserialization, path traversal, SSRF (z.ai subscription route; **deprioritized — pathologically slow under load, and the single point of failure in the 2026-09-05 quota incident that motivated review-cli#382** — live `security` fallback for Qwen, freed from `quality` now that Sonnet covers it) |
| 15 | reserve (last, no preset) | Fable | `claude:claude-fable-5` | `architect` | architecture, design coherence, API shape, abstraction boundaries (**deprioritized to last-resort — confirmed ~100% dispatch failure rate as of 2026-09-05**, review-cli#fable-seat-reliability) |

See **[docs/board.md](docs/board.md)** for the full mechanics: in-seat retry vs.
reserve-replace, the cross-invocation cooldown's escalating window, Claude account rotation
across multiple accounts, the memory-aware concurrency cap, and how `board:`/`models:` config
precedence resolves.

**Per-project lens overrides (review-cli#547):** drop `.review/lenses/<role>.md` at
your repo's root (`<role>` is one of the known roles above) to teach a board role
something specific to your codebase — a payments invariant, a house style rule —
without forking or patching review-cli. The file's text is **appended** to that
role's built-in lens (never replaces it), so the seat keeps the built-in coverage
plus your addition; an absent, unreadable, or blank/whitespace-only file is treated
as no override, and a symlink at the override path or an intermediate `.review`/
`.review/lenses` directory that escapes the repo root is never followed (also treated
as absent). An **oversized** file is not rejected the same way: it is silently
truncated to the first 8 KiB (`LENS_OVERRIDE_MAX_BYTES`) and the truncated text still
becomes the seat's active override — so a repo that never adds this directory behaves
exactly as before, but a project author who exceeds 8 KiB gets a quietly truncated
lens, not a rejected one.
`review --show-board` marks any seat whose role has an active override with a
`(lens overridden: .review/lenses/<role>.md)` suffix. (A bigger docs reorg is
tracked separately in review-cli#552 and will eventually move this into
`docs/board.md`.)

---

## Auth

Each backend authenticates independently; review-cli never writes a credential to disk,
only reads one. **Gemini** and the diff-only keyed-HTTP backends (`-m cc` / `-m glm` /
`-m openrouter:<model>`) read their own review-cli env var — `GEMINI_API_KEY`
(or `GOOGLE_API_KEY`), `COMMANDCODE_API_KEY`, `ZAI_API_KEY` (or `ZHIPU_API_KEY`),
`OPENROUTER_API_KEY` — from the environment or `~/.config/review-cli/.env`
(override the search path with `$GEMINI_ENV_FILE`). **Codex / opencode / omp** must be
on PATH and authenticated through their OWN CLI setup (`codex` login, `opencode auth
login` / `opencode providers login`, `omp setup`) — no review-cli key covers them.
**Claude** normally runs through its own agentic CLI too, but falls back to a
diff-only Anthropic API call on a host with no `claude` CLI but a set
`ANTHROPIC_API_KEY` / `ANTHROPIC_AUTH_TOKEN` (+ optional `ANTHROPIC_BASE_URL`).

**The most surprising precedence fact:** the default AGENTIC board seats that run
*through* opencode (Kimi, Qwen, DeepSeek, GLM, Grok) and through omp (MiMo) authenticate
via **opencode's / omp's own provider config**, NOT via `COMMANDCODE_API_KEY` /
`ZAI_API_KEY` / a review-cli env var — those keys only gate the separate diff-only
`-m cc` / `-m glm` keyed-HTTP path. A missing binary or missing provider auth is caught
by the startup probe and the seat is skipped gracefully (pool backfills from reserve);
`unpaid_providers:` / `REVIEW_UNPAID_PROVIDERS` skips a provider that's authenticated but
not currently billed/entitled.

Full per-backend reference — every env var, the shared `.env` file, z.ai's Coding-Plan
endpoint requirement for `glm-5.2`, OpenRouter model-slug selection, and the
opencode/omp startup-probe details — lives in **`review help config`**. Seat-level board
history and the per-seat cooldown/retry mechanics live in
[docs/board.md](docs/board.md).

---

## Architecture — `lib | cli | mcp` + the mode registry

review-cli is layered so the same panel engine is reusable beyond the CLI:

- **lib** — `reviewlib/` is the engine: `panel.py` (parallel fan-out, moderator,
  failover board), `backends.py` (the model transports), `config.py` (board/defaults).
  It has no argparse dependency and is callable directly.
- **cli** — `reviewlib/cli.py` is a **thin** argparse front-end. It resolves the diff,
  models, and `--visual` context, then dispatches to a mode handler. It owns no review
  logic of its own.
- **mcp** — *not built yet*, but the seam is kept clean: an MCP wrapper (or another CLI
  — a future research-cli / task-cli `just-ask`) can call the lib + a mode handler
  directly without dragging the argparse surface along. Each mode handler is thin over
  the lib for exactly this reason.

**Modes are plugin-directory modules** (generalized from the per-project
`features/visual` MODULE registry):

```
reviewlib/modes/
  contract.py     # ModeSpec descriptor + ModeContext (mirrors features/visual/contract.py + module_api.py)
  registry.py     # MODES list, get_mode / known_subcommands / default_mode (mirrors features/visual/registry.py)
  review.py       # MODE = ModeSpec(subcommand="diff", diff_policy="require", handler=…)
  brainstorm.py   # MODE = ModeSpec(subcommand="brainstorm", diff_policy="optional", …)
  just_ask.py     # MODE = ModeSpec(subcommand="just-ask", diff_policy="none", …)
  quorum.py       # MODE = ModeSpec(subcommand="quorum", diff_policy="none", …)
```

Each mode is a **self-describing module** that exposes a top-level `MODE = ModeSpec(…)`
declaring the subcommand it registers, its default diff policy, the CLI arguments it
adds, and its thin handler — exactly how a visual module exposes a top-level `MODULE`.
The CLI looks the subcommand up in the registry and dispatches; **adding a mode = drop a
`modes/<name>.py` and list it in `registry.MODES` — no `cli.py` surgery.** A bare
`review` with no recognized subcommand prints help rather than running any mode (see
[Quick start](#quick-start)); `--visual` is a composable flag orthogonal to the mode, so it
rides any subcommand.

---

## How review compares

AI code review tools cluster into four camps. **PR-bots** (Qodo PR-Agent, CodeRabbit,
GitHub Copilot code review) run a single model against a *pull request* — they live on
the platform, comment inline, and are great once a PR exists. **In-agent review** (Claude
Code `/review`, Codex review) runs one model on the local diff inside the harness you are
already in. **Autonomous-loop tools** ([ralphex](https://github.com/umputun/ralphex)) fold
review *into* a full code-gen loop — they drive a coding agent through a plan and review its
output as one fused, opinionated pipeline (see [review vs ralphex](#review-vs-ralphex--the-whole-loop-vs-the-review-primitive) below).
**Supervised multi-agent review CLIs** ([revmux](https://github.com/umputun/revmux)) spawn
`claude`/`codex` as subprocesses they own and watch in a TUI, returning one stable JSON
findings contract — a narrower, more auditable shape than a panel (see
[review vs revmux](#review-vs-revmux--panel--decision-tool-vs-supervised-multi-agent-process) below).

`review` sits outside all four camps above: it runs **several models in parallel on the local working-tree diff**
before you ever push, then goes further — a cited **quorum** (consensus with evidence) and
a multi-round **brainstorm** panel for open design questions. It adds **visual review**
(attach a render with `--visual` for a keep / rollback / repair verdict) and **interactive
spec-review tooling** (review a markdown spec like a PR), all from the same binary. It is
**read-only** (never edits your code), **CLI-first** (no PR, no hosted service — it shells
out to model CLIs you already have), and **harness-agnostic** (callable from Claude Code,
Codex, opencode, or a plain shell).

| Tool | Multi-model in parallel | Local pre-PR diff | Consensus / quorum | Design brainstorm | Read-only | No hosted service | Generates code |
|---|---|---|---|---|---|---|---|
| **review** | ✓ | ✓ | ✓ (cited) | ✓ (multi-round) | ✓ | ✓ (your own model CLIs) | — (review only, by design) |
| Qodo PR-Agent | — (1 call) | ~ (CLI, PR-oriented) | — | — | — (suggests edits) | ~ (self-host or hosted) | — |
| CodeRabbit CLI | — (1 service) | ✓ | — | — | — (one-click fixes) | — (hosted) | — |
| GitHub Copilot review | — (1 model) | — (PR / IDE) | — | — | — (suggests edits) | — (hosted) | — |
| Claude Code `/review` | — (1 model) | ✓ | — | — | ~ | — (in-harness) | — |
| Codex review | — (1 model) | ✓ | — | — | ~ | — (in-harness) | — |
| ralphex | ✓ (5 review agents + opt. codex) | ✓ | — | — | — (drives an agent that edits) | ✓ (local Go binary) | ✓ (drives the coding agent) |
| revmux | ✓ (claude + codex, supervised subprocesses) | ~ (caller-written round; any subject) | — | — (`triage` rates relevance, not open ideation) | ✓ | ✓ (your own claude/codex CLIs) | — (review only, by design) |

`~` = partial. PR-bots shine *after* a PR exists and can apply fixes; `review` is the
pre-commit, multi-perspective second opinion that runs from any shell and decides nothing
for you — it surfaces findings and consensus, you stay in control of the edit.

### review vs ralphex — the whole loop vs the review primitive

[**ralphex**](https://github.com/umputun/ralphex) is the *extended Ralph loop*: a single
local binary that takes a written plan and runs the **entire** autonomous loop — it drives
a coding agent (Claude Code / codex / Copilot CLI) to write code task-by-task in fresh
sessions, then runs its own multi-agent review pipeline (5 parallel agents → optional GPT-5
codex cross-review → a final pass), committing after each step. It genuinely owns the part
`review` does not touch at all: **code generation**. If you want "write a plan, walk away,
come back to reviewed-and-committed code" in one opinionated tool, that is exactly what
ralphex is for, and `review` is `—` on code-gen on purpose.

`review` is the other half of that picture: not a loop, but the **review component an agent
plugs into a loop it controls itself**. You (or your agent) drive the loop and call `review`
only for the critique step — which is the step agents do *worst* on their own. The two shapes:

**ralphex** — one opaque binary that encapsulates *both* code-generation and review in a single
autonomous loop:

![ralphex — all-in-one encapsulated loop: one binary drives a code-gen agent, runs its own built-in review, and commits, repeating until the plan is done](docs/compare-ralphex.svg)

**`review`** — does *only* the review part (the part agents do badly), transparently and
controllably, driven by the agent that keeps code-gen and every decision:

![review — a critique primitive the agent orchestrates: the agent writes code, calls review (multi-model, read-only, never edits) for the critique step, then decides fix / ship / re-loop](docs/compare-review.svg)

In one sentence: **ralphex's strength is that it is all-in-one** — code-gen and review
fused into one walk-away binary; **`review`'s strength is that it is focused and
controllable** — a read-only, multi-model review primitive the agent composes into its own
Ralph loop, so the agent keeps full control of code-gen and of every fix/ship/re-loop
decision instead of handing the loop to a black box.

### review vs revmux — panel + decision tool vs supervised multi-agent process

[**revmux**](https://github.com/umputun/revmux) spawns `claude --print` and `codex exec` as
subprocesses it owns, watches them in a terminal TUI, and returns one structured JSON report
with a stable schema — findings, severity, confidence, verdict, fixed exit codes. It is
absolutely read-only, reviews *any* subject (a branch, a PR, a design doc, a filed issue, not
only a diff) through a caller-written **task round**, and checks its review standard
(`.revmux/` profiles and lenses) into the target repo as versioned files.

`review` overlaps on intent — stop one model grading its own homework, stay out of code-gen,
leave a durable log behind — but is architecturally wider: a **15-seat priority board** across
many providers (not just claude+codex) with three layers of failover and per-seat reliability
tracking, plus modes revmux has no equivalent for — cited **quorum**, multi-round
**brainstorm**, **visual** screenshot review, and the one narrow un-caged **`qa`** write/exec
tester. Its core interface is a human-readable panel, not a stable findings-JSON contract; the
JSON side of `review` lives in separate management commands (`task`, `stat`, `entities`, `jobs`, `visual`).

Neither replaces the other — they don't share config and don't conflict. See
**[docs/compare-revmux.md](docs/compare-revmux.md)** for the full, detailed, feature-by-feature
comparison, and for where the rest of the field fits.

---

## Ecosystem

Part of the [HyperIDE.ai](https://hyperide.ai) agent toolchain:

- **[tg-cli](https://git.hyperide.ai/ultrabricks/tg-cli)** — simple Telegram CLI to send messages, photos & files, and a two-way agent bridge (reports, Q→buttons, voice/rich)
- **[rig-cli](https://git.hyperide.ai/ultrabricks/rig-cli)** — umbrella dev-env driver: sets up a repo from config — skills, hooks, CI, dep-bootstrap; reconciles drift
- **[agent-tools](https://git.hyperide.ai/ultrabricks/agent-tools)** — the shared catalog `rig` applies: portable agent skills, agent-hooks, the global git-hook dispatcher, CI gates, and MCP servers
- **[draw-cli](https://github.com/alex-mextner/draw-cli)** — text-to-image via Hugging Face
- **[3d-cli](https://github.com/alex-mextner/3d-cli)** — scriptable CLI for the full 3D FDM lifecycle: modeling, mesh repair, slicing, and print monitoring
- **[hyperide.ai](https://hyperide.ai)** — Figma replacement inside VS Code. Edit React components directly through AST/LSP without AI hallucinations, token waste, or context-window limits. Works for indie vibe-coding and for enterprise teams with split design/dev roles.

Each CLI registers a skill into your agent harnesses (`<tool> install-skill`) so agents know it exists — see Install.

---

## Comparisons

See [How review compares](#how-review-compares) for the field at a glance, and
**[docs/compare-revmux.md](docs/compare-revmux.md)** for the deepest one-on-one —
`review` vs [revmux](https://github.com/umputun/revmux), feature by feature, plus where the
rest of the field (2ndOpinion, Open Code Review, Calimero ai-code-reviewer, ralphex, and more)
fits.
