The not-logged-in check reads a worker's prose as an auth failure: match the CLI's own error, not the report text
Implemented and committed as `43b5228`.
Implemented and committed as `43b5228`.
Plate V · context-garden
What is moving, what comes next, and what it adds up to.
Implemented and committed as `43b5228`.
Other design's excerpt omitted to preserve independent design.
No assistant text recorded.
No assistant text recorded.
No assistant text recorded.
Each queue moves on its own.
Worker temp files live on disk and are pruned: TMPDIR under the work root, per-run cleanup at reap, and a sweep of finished tasks' worktree venvs and caches
The task page shows the decision a worker's no-change or question report needs: the same card and actions as the Inbox, right where the notification sends you
Inbox decision cards lay out at full width: the text column no longer collapses to one word per line and the action buttons no longer overlap the evidence list
Design Now 2: astra's take on the Now page, information architecture, visual system, motion, and a static mock of every state
UI changes are reviewed against rendered pages: a template or style change captures the affected pages as screenshots at two widths, the reviewer and the personas read them, and the walkthrough uses the same capture
A dispatch that fails before its process starts closes the run record at once, and the orphan sweep closes any running record with no live process
The task page renders a trial in progress: the trial panel treats winner, scores and PRs as optional, and a test renders a task mid-trial with a failed contender
The web app serves design documents, mocks and run captures: /design/<file> for the product's docs/design and a run page link to each capture
Onboarding skill and command: analyse an existing project and its environment to create a garden product, principles, setup config and a first phase
The retro captures its own walkthrough, and hand merges and tick duration are in garden metrics and the rail
CG-217 merged with a criterion marked 'no evidence given' and CG-158 with a plac
The built-in economy, balanced and fast stops hardcode Claude model ids and mode
A brief never ships with an empty or unresolved reading list, and a revise brief restates the criteria and the concrete blocker
Docs match the mechanism: the scaffolded operate skill, design.md and roadmap.md non-goals, and the architecture module map
The retro's Numbers section reads the operator ledger where the owner keeps it, and reports spend and share
Retro questions are deduplicated across reconcile runs, and a second judge can run with task filing off
Spike: OpenRouter harness shape and adapter CLI
parse_result accepts the GARDEN_RESULT marker wrapped in markdown emphasis or code and a JSON payload that spans lines
redispatch kills the superseded worker, and pin runs the canary, installs and restarts after a tick
A draft's acceptance criteria and reading list can be edited inline on the task page
Planning sequences dependent tasks and inlines retro evidence into the brief
One vocabulary and one place for each fact across rail, Config, CLI and Inbox
The walkthrough renderer skips hidden elements and attributes, and check runs retry once on a signal exit
Update docs: docs/architecture.md
Update docs: docs/worker-protocol.md
Update docs: docs/design.md
Persona runs are recorded per phase and the retro validates the run id it reads
Small follow-ups: retro_question kind, backlog noscript, CI flake, ambient config dir in tests, a second opinion for self products
No eligible merge candidates.
Hold untrusted config changes before live reload
Progress in context

Papaver argemone
10 / 45A garden in someone else’s environment.
phase-04
phase-06



Outcomes, not activity alone
8 merged in this window
$28.45 recorded window spend
$4.53 / accepted task, as exported.
0% first-pass approval
Scroll to compare models →
| Difficulty | claude-sonnet-5 | gpt-5.6-terra | claude-opus-4-8 | claude-fable-5-1 |
|---|---|---|---|---|
| easy | $3.00=n=2 † | —n=0 | —n=0 | —n=0 |
| medium | $5.22↓n=1 † | $3.08↑n=3 | —n=0 | —n=0 |
| hard | —n=0 | —n=0 | $7.27=n=2 † | —n=0 |
| Difficulty | claude-sonnet-5 | gpt-5.6-terra | claude-opus-4-8 | claude-fable-5-1 |
|---|---|---|---|---|
| easy | 0%=n=1 † | —n=0 | —n=0 | —n=0 |
| medium | —n=0 | —n=0 | —n=0 | —n=0 |
| hard | —n=0 | —n=0 | —n=0 | 0%=n=1 † |
| Difficulty | claude-sonnet-5 | gpt-5.6-terra | claude-opus-4-8 | claude-fable-5-1 |
|---|---|---|---|---|
| easy | —n=0 | —n=0 | —n=0 | —n=0 |
| medium | $0.62↓n=2 † | $0.53↑n=2 † | —n=0 | —n=0 |
| hard | —n=0 | —n=0 | —n=0 | $11.32=n=2 † |
| Difficulty | claude-sonnet-5 | gpt-5.6-terra | claude-opus-4-8 | claude-fable-5-1 |
|---|---|---|---|---|
| easy | 0.0=n=2 † | —n=0 | —n=0 | —n=0 |
| medium | 1.0↑n=1 † | 1.3↓n=3 | —n=0 | —n=0 |
| hard | —n=0 | —n=0 | 0.0=n=2 † | —n=0 |
| Difficulty | claude-sonnet-5 | gpt-5.6-terra | claude-opus-4-8 | claude-fable-5-1 |
|---|---|---|---|---|
| easy | 94 min=n=2 † | —n=0 | —n=0 | —n=0 |
| medium | 115 min↓n=1 † | 97 min↑n=3 | —n=0 | —n=0 |
| hard | —n=0 | —n=0 | 119 min=n=2 † | —n=0 |
↑ best · ↓ worst · = tied or only value · † n < 3, provisional
Grounds compare within a row. Colour is never the only signal.
10 merged in this window
$88.90 recorded window spend
$4.41 / accepted task, as exported.
42% first-pass approval
Scroll to compare models →
| Difficulty | claude-sonnet-5 | gpt-5.6-terra | claude-opus-4-8 | claude-fable-5-1 | gpt-6-astra |
|---|---|---|---|---|---|
| easy | $2.50=n=3 | —n=0 | —n=0 | —n=0 | —n=0 |
| medium | $5.80↓n=2 † | $3.08↑n=3 | —n=0 | —n=0 | —n=0 |
| hard | —n=0 | —n=0 | $7.27=n=2 † | —n=0 | —n=0 |
| Difficulty | claude-sonnet-5 | gpt-5.6-terra | claude-opus-4-8 | claude-fable-5-1 | gpt-6-astra |
|---|---|---|---|---|---|
| easy | 67%=n=3 | —n=0 | —n=0 | —n=0 | —n=0 |
| medium | 50%↑n=2 † | 0%↓n=4 | —n=0 | —n=0 | —n=0 |
| hard | —n=0 | —n=0 | 100%↑n=2 † | 0%↓n=1 † | —n=0 |
| Difficulty | claude-sonnet-5 | gpt-5.6-terra | claude-opus-4-8 | claude-fable-5-1 | gpt-6-astra |
|---|---|---|---|---|---|
| easy | $1.94=n=3 | —n=0 | —n=0 | —n=0 | —n=0 |
| medium | $3.32↓n=5 | $0.72↑n=10 | —n=0 | —n=0 | —n=0 |
| hard | —n=0 | —n=0 | $6.24n=2 † | $11.32↓n=2 † | $5.98↑n=1 † |
| Difficulty | claude-sonnet-5 | gpt-5.6-terra | claude-opus-4-8 | claude-fable-5-1 | gpt-6-astra |
|---|---|---|---|---|---|
| easy | 0.0=n=3 | —n=0 | —n=0 | —n=0 | —n=0 |
| medium | 1.0↑n=2 † | 1.3↓n=3 | —n=0 | —n=0 | —n=0 |
| hard | —n=0 | —n=0 | 0.0=n=2 † | —n=0 | —n=0 |
| Difficulty | claude-sonnet-5 | gpt-5.6-terra | claude-opus-4-8 | claude-fable-5-1 | gpt-6-astra |
|---|---|---|---|---|---|
| easy | 80 min=n=3 | —n=0 | —n=0 | —n=0 | —n=0 |
| medium | 84 min↑n=2 † | 97 min↓n=3 | —n=0 | —n=0 | —n=0 |
| hard | —n=0 | —n=0 | 119 min=n=2 † | —n=0 | —n=0 |
↑ best · ↓ worst · = tied or only value · † n < 3, provisional
Grounds compare within a row. Colour is never the only signal.
100 merged in this window
$896.16 recorded window spend
$7.15 / accepted task, as exported.
77% first-pass approval
Scroll to compare models →
| Difficulty | claude-sonnet-5 | claude-opus-4-8 | claude-fable-5-1 | gpt-5.6-terra | gpt-6-astra |
|---|---|---|---|---|---|
| easy | $4.62=n=50 | —n=0 | —n=0 | —n=0 | —n=0 |
| medium | $11.69n=15 | $8.98n=23 | $13.36↓n=1 † | $5.58↑n=4 | —n=0 |
| hard | —n=0 | $7.27↑n=2 † | $12.20↓n=4 | —n=0 | —n=0 |
| Difficulty | claude-sonnet-5 | claude-opus-4-8 | claude-fable-5-1 | gpt-5.6-terra | gpt-6-astra |
|---|---|---|---|---|---|
| easy | 86%=n=49 | —n=0 | —n=0 | —n=0 | —n=0 |
| medium | 57%n=14 | 85%↑n=26 | 0%↓n=1 † | 0%↓n=4 | —n=0 |
| hard | —n=0 | 100%↑n=2 † | 80%↓n=5 | —n=0 | —n=0 |
| Difficulty | claude-sonnet-5 | claude-opus-4-8 | claude-fable-5-1 | gpt-5.6-terra | gpt-6-astra |
|---|---|---|---|---|---|
| easy | $2.25=n=73 | —n=0 | —n=0 | —n=0 | —n=0 |
| medium | $2.33n=41 | $4.56↓n=49 | $3.83n=3 | $0.60↑n=12 | —n=0 |
| hard | —n=0 | $6.24n=2 † | $7.31↓n=9 | —n=0 | $5.98↑n=1 † |
| Difficulty | claude-sonnet-5 | claude-opus-4-8 | claude-fable-5-1 | gpt-5.6-terra | gpt-6-astra |
|---|---|---|---|---|---|
| easy | 0.3=n=50 | —n=0 | —n=0 | —n=0 | —n=0 |
| medium | 1.4↓n=15 | 0.2↑n=23 | 1.0n=1 † | 1.2n=4 | —n=0 |
| hard | —n=0 | 0.0↑n=2 † | 0.8↓n=4 | —n=0 | —n=0 |
| Difficulty | claude-sonnet-5 | claude-opus-4-8 | claude-fable-5-1 | gpt-5.6-terra | gpt-6-astra |
|---|---|---|---|---|---|
| easy | 47 min=n=50 | —n=0 | —n=0 | —n=0 | —n=0 |
| medium | 153 min↓n=15 | 46 min↑n=23 | 55 minn=1 † | 114 minn=4 | —n=0 |
| hard | —n=0 | 119 min↓n=2 † | 39 min↑n=4 | —n=0 | —n=0 |
↑ best · ↓ worst · = tied or only value · † n < 3, provisional
Grounds compare within a row. Colour is never the only signal.
10 merged in this window
$88.22 recorded window spend
$4.41 / accepted task, as exported.
42% first-pass approval
Scroll to compare models →
| Difficulty | claude-sonnet-5 | gpt-5.6-terra | claude-opus-4-8 | claude-fable-5-1 | gpt-6-astra |
|---|---|---|---|---|---|
| easy | $2.50=n=3 | —n=0 | —n=0 | —n=0 | —n=0 |
| medium | $5.80↓n=2 † | $3.08↑n=3 | —n=0 | —n=0 | —n=0 |
| hard | —n=0 | —n=0 | $7.27=n=2 † | —n=0 | —n=0 |
| Difficulty | claude-sonnet-5 | gpt-5.6-terra | claude-opus-4-8 | claude-fable-5-1 | gpt-6-astra |
|---|---|---|---|---|---|
| easy | 67%=n=3 | —n=0 | —n=0 | —n=0 | —n=0 |
| medium | 50%↑n=2 † | 0%↓n=4 | —n=0 | —n=0 | —n=0 |
| hard | —n=0 | —n=0 | 100%↑n=2 † | 0%↓n=1 † | —n=0 |
| Difficulty | claude-sonnet-5 | gpt-5.6-terra | claude-opus-4-8 | claude-fable-5-1 | gpt-6-astra |
|---|---|---|---|---|---|
| easy | $1.94=n=3 | —n=0 | —n=0 | —n=0 | —n=0 |
| medium | $3.32↓n=5 | $0.72↑n=10 | —n=0 | —n=0 | —n=0 |
| hard | —n=0 | —n=0 | $6.24n=2 † | $11.32↓n=2 † | $5.98↑n=1 † |
| Difficulty | claude-sonnet-5 | gpt-5.6-terra | claude-opus-4-8 | claude-fable-5-1 | gpt-6-astra |
|---|---|---|---|---|---|
| easy | 0.0=n=3 | —n=0 | —n=0 | —n=0 | —n=0 |
| medium | 1.0↑n=2 † | 1.3↓n=3 | —n=0 | —n=0 | —n=0 |
| hard | —n=0 | —n=0 | 0.0=n=2 † | —n=0 | —n=0 |
| Difficulty | claude-sonnet-5 | gpt-5.6-terra | claude-opus-4-8 | claude-fable-5-1 | gpt-6-astra |
|---|---|---|---|---|---|
| easy | 80 min=n=3 | —n=0 | —n=0 | —n=0 | —n=0 |
| medium | 84 min↑n=2 † | 97 min↓n=3 | —n=0 | —n=0 | —n=0 |
| hard | —n=0 | —n=0 | 119 min=n=2 † | —n=0 | —n=0 |
↑ best · ↓ worst · = tied or only value · † n < 3, provisional
Grounds compare within a row. Colour is never the only signal.
Inspection sheet · all examples are simulated
The live page uses these treatments in place. This sheet keeps them visible for design review.
Validating the generated context against the project.
Every PR reaches the review queue
Waiting for CI before merge. Keep for eight seconds, then settle into recent activity.
Base branch checks failed.
Other workers continue. This reason stays visible until the hold clears.
Account quota reached.
New codex work waits; in-flight clocks keep counting. No retry time reported.
The acceptance check returned an error.
Run stopped at 1:08. The failure stays until a retry or a decision resolves it.
Plan work from the phase goals or create the first task.
Review existing drafts before workers can start.
Waiting for an approved task.
Phase and past outcomes stay visible. There is no idle animation.
Last received at 02:44 UTC.
Retain the last snapshot. Clocks show elapsed since start; they cannot prove a worker is still alive.
Queue unavailable · typical not established · progress not mapped.
Costs and outcomes show —. Missing history is not an empty garden.