context-garden
Recorded snapshot · 2026-09-06T02:34:47+00:00. Main regions render the supplied garden export. State atlas below is simulated. Task, run and Inbox links show live-app destinations; they require a running app. Clocks replay from capture time; this is not a live connection.

Plate V · context-garden

Work in view.

What is moving, what comes next, and what it adds up to.

September 6, 2026
5 run records · 3 without a process

Now

Ⅱ Merge held

Hold untrusted config changes before live reload

Process missing · revise

`garden dispatch <id>` (CLI) allows dispatching a draft directly

No assistant text recorded.

No assistant text recorded.

Typical not established
Process missing · revise

Hold untrusted config changes before live reload

No assistant text recorded.

No assistant text recorded.

Typical not established
Process missing · revise

Isolate planner execution from operator state

No assistant text recorded.

No assistant text recorded.

Typical not established

Next

Each queue moves on its own.

Workers · 5 / 5 recorded occupied

  1. Worker temp files live on disk and are pruned: TMPDIR under the work root, per-run cleanup at reap, and a sweep of finished tasks' worktree venvs and caches

  2. The task page shows the decision a worker's no-change or question report needs: the same card and actions as the Inbox, right where the notification sends you

  3. Inbox decision cards lay out at full width: the text column no longer collapses to one word per line and the action buttons no longer overlap the evidence list

  1. Design Now 2: astra's take on the Now page, information architecture, visual system, motion, and a static mock of every state

  2. UI changes are reviewed against rendered pages: a template or style change captures the affected pages as screenshots at two widths, the reviewer and the personas read them, and the walkthrough uses the same capture

  3. A dispatch that fails before its process starts closes the run record at once, and the orphan sweep closes any running record with no live process

  4. The task page renders a trial in progress: the trial panel treats winner, scores and PRs as optional, and a test renders a task mid-trial with a failed contender

  5. The web app serves design documents, mocks and run captures: /design/<file> for the product's docs/design and a run page link to each capture

  6. Onboarding skill and command: analyse an existing project and its environment to create a garden product, principles, setup config and a first phase

  7. The retro captures its own walkthrough, and hand merges and tick duration are in garden metrics and the rail

  8. CG-217 merged with a criterion marked 'no evidence given' and CG-158 with a plac

  9. The built-in economy, balanced and fast stops hardcode Claude model ids and mode

  10. A brief never ships with an empty or unresolved reading list, and a revise brief restates the criteria and the concrete blocker

  11. Docs match the mechanism: the scaffolded operate skill, design.md and roadmap.md non-goals, and the architecture module map

  12. The retro's Numbers section reads the operator ledger where the owner keeps it, and reports spend and share

  13. Retro questions are deduplicated across reconcile runs, and a second judge can run with task filing off

  14. Spike: OpenRouter harness shape and adapter CLI

  15. parse_result accepts the GARDEN_RESULT marker wrapped in markdown emphasis or code and a JSON payload that spans lines

  16. redispatch kills the superseded worker, and pin runs the canary, installs and restarts after a tick

  17. A draft's acceptance criteria and reading list can be edited inline on the task page

  18. Planning sequences dependent tasks and inlines retro evidence into the brief

  19. One vocabulary and one place for each fact across rail, Config, CLI and Inbox

  20. The walkthrough renderer skips hidden elements and attributes, and check runs retry once on a signal exit

  21. Update docs: docs/architecture.md

  22. Update docs: docs/worker-protocol.md

  23. Update docs: docs/design.md

  24. Persona runs are recorded per phase and the retro validates the run id it reads

  25. Small follow-ups: retro_question kind, backlog noscript, CI flake, ambient config dir in tests, a second opinion for self products

Merges

No eligible merge candidates.

Hold untrusted config changes before live reload

Reviews · 0 / 3 slots occupied

Hold untrusted config changes before live reload

Progress in context

Where we are

prickly poppy, botanical plate by Thomé, 1885

Papaver argemone

10 / 45in leaf

A garden in someone else’s environment.

  • Onboardingnot started · 0 / 1 linked tasks
  • Any model, at the right pricenot started · 0 / 1 linked tasks
  • Shared quotasnot started · 0 / 1 linked tasks
  • Any machinenot started · 0 / 1 linked tasks
  • What the phase-04 retro addsin flight · 3 / 4 linked tasks
Dryopteris filix-mas

phase-04

Paeonia mascula

phase-06

Earlier specimens · closed phases

Digitalis purpurea
Plate III
30 done
Pisum sativum
Plate I
19 done
Rubus fruticosus agg.
Plate II
90 done

Outcomes, not activity alone

The last period

8 merged in this window

$28.45 recorded window spend

$4.53 / accepted task, as exported.

0% first-pass approval

Scroll to compare models →

Cost per accepted tasklower is better · n = accepted tasks · exported measurements
Difficultyclaude-sonnet-5gpt-5.6-terraclaude-opus-4-8claude-fable-5-1
easy$3.00=n=2 n=0n=0n=0
medium$5.22n=1 $3.08n=3n=0n=0
hardn=0n=0$7.27=n=2 n=0
First-pass approvalhigher is better · n = reviewed tasks (export cohort) · exported measurements
Difficultyclaude-sonnet-5gpt-5.6-terraclaude-opus-4-8claude-fable-5-1
easy0%=n=1 n=0n=0n=0
mediumn=0n=0n=0n=0
hardn=0n=0n=00%=n=1
Work-run costlower is better · n = work runs · exported measurements
Difficultyclaude-sonnet-5gpt-5.6-terraclaude-opus-4-8claude-fable-5-1
easyn=0n=0n=0n=0
medium$0.62n=2 $0.53n=2 n=0n=0
hardn=0n=0n=0$11.32=n=2
Revise roundslower is better · n = accepted tasks · exported measurements
Difficultyclaude-sonnet-5gpt-5.6-terraclaude-opus-4-8claude-fable-5-1
easy0.0=n=2 n=0n=0n=0
medium1.0n=1 1.3n=3n=0n=0
hardn=0n=00.0=n=2 n=0
Median lead timelower is better · n = accepted tasks with lead times · exported measurements
Difficultyclaude-sonnet-5gpt-5.6-terraclaude-opus-4-8claude-fable-5-1
easy94 min=n=2 n=0n=0n=0
medium115 minn=1 97 minn=3n=0n=0
hardn=0n=0119 min=n=2 n=0

↑ best · ↓ worst · = tied or only value · † n < 3, provisional
Grounds compare within a row. Colour is never the only signal.

Recorded spend by activity

check
$0.00
other
$22.65
rebase
$0.49
review
$3.01
revise
$2.30

Runs and hand steps

claude:claude-fable-5-1
2 runs
garden:review
5 runs
claude:claude-sonnet-5
3 runs
codex:gpt-5.6-terra
2 runs
codex:gpt-5.6-luna
1 runs
garden:rebase
2 runs
garden:check
9 runs
Hand steps
3 recorded
Manual merges
not supplied
dispatch_paused
1
dispatch_resumed
2

10 merged in this window

$88.90 recorded window spend

$4.41 / accepted task, as exported.

42% first-pass approval

Scroll to compare models →

Cost per accepted tasklower is better · n = accepted tasks · exported measurements
Difficultyclaude-sonnet-5gpt-5.6-terraclaude-opus-4-8claude-fable-5-1gpt-6-astra
easy$2.50=n=3n=0n=0n=0n=0
medium$5.80n=2 $3.08n=3n=0n=0n=0
hardn=0n=0$7.27=n=2 n=0n=0
First-pass approvalhigher is better · n = reviewed tasks (export cohort) · exported measurements
Difficultyclaude-sonnet-5gpt-5.6-terraclaude-opus-4-8claude-fable-5-1gpt-6-astra
easy67%=n=3n=0n=0n=0n=0
medium50%n=2 0%n=4n=0n=0n=0
hardn=0n=0100%n=2 0%n=1 n=0
Work-run costlower is better · n = work runs · exported measurements
Difficultyclaude-sonnet-5gpt-5.6-terraclaude-opus-4-8claude-fable-5-1gpt-6-astra
easy$1.94=n=3n=0n=0n=0n=0
medium$3.32n=5$0.72n=10n=0n=0n=0
hardn=0n=0$6.24n=2 $11.32n=2 $5.98n=1
Revise roundslower is better · n = accepted tasks · exported measurements
Difficultyclaude-sonnet-5gpt-5.6-terraclaude-opus-4-8claude-fable-5-1gpt-6-astra
easy0.0=n=3n=0n=0n=0n=0
medium1.0n=2 1.3n=3n=0n=0n=0
hardn=0n=00.0=n=2 n=0n=0
Median lead timelower is better · n = accepted tasks with lead times · exported measurements
Difficultyclaude-sonnet-5gpt-5.6-terraclaude-opus-4-8claude-fable-5-1gpt-6-astra
easy80 min=n=3n=0n=0n=0n=0
medium84 minn=2 97 minn=3n=0n=0n=0
hardn=0n=0119 min=n=2 n=0n=0

↑ best · ↓ worst · = tied or only value · † n < 3, provisional
Grounds compare within a row. Colour is never the only signal.

Recorded spend by activity

check
$0.00
other
$32.74
rebase
$0.74
review
$13.29
revise
$4.38
work
$37.75

Runs and hand steps

claude:claude-sonnet-5
10 runs
claude:claude-fable-5-1
2 runs
garden:review
19 runs
claude:claude-opus-4-8
2 runs
codex:gpt-5.6-terra
10 runs
codex:gpt-6-astra
1 runs
garden:edit
21 runs
garden:kickoff
1 runs
codex:gpt-5.6-luna
1 runs
garden:rebase
5 runs
garden:check
29 runs
Hand steps
31 recorded
Manual merges
not supplied
decision_resolved
4
dispatch_paused
1
dispatch_resumed
3
moved
2
suggestion
21

100 merged in this window

$896.16 recorded window spend

$7.15 / accepted task, as exported.

77% first-pass approval

Scroll to compare models →

Cost per accepted tasklower is better · n = accepted tasks · exported measurements
Difficultyclaude-sonnet-5claude-opus-4-8claude-fable-5-1gpt-5.6-terragpt-6-astra
easy$4.62=n=50n=0n=0n=0n=0
medium$11.69n=15$8.98n=23$13.36n=1 $5.58n=4n=0
hardn=0$7.27n=2 $12.20n=4n=0n=0
First-pass approvalhigher is better · n = reviewed tasks (export cohort) · exported measurements
Difficultyclaude-sonnet-5claude-opus-4-8claude-fable-5-1gpt-5.6-terragpt-6-astra
easy86%=n=49n=0n=0n=0n=0
medium57%n=1485%n=260%n=1 0%n=4n=0
hardn=0100%n=2 80%n=5n=0n=0
Work-run costlower is better · n = work runs · exported measurements
Difficultyclaude-sonnet-5claude-opus-4-8claude-fable-5-1gpt-5.6-terragpt-6-astra
easy$2.25=n=73n=0n=0n=0n=0
medium$2.33n=41$4.56n=49$3.83n=3$0.60n=12n=0
hardn=0$6.24n=2 $7.31n=9n=0$5.98n=1
Revise roundslower is better · n = accepted tasks · exported measurements
Difficultyclaude-sonnet-5claude-opus-4-8claude-fable-5-1gpt-5.6-terragpt-6-astra
easy0.3=n=50n=0n=0n=0n=0
medium1.4n=150.2n=231.0n=1 1.2n=4n=0
hardn=00.0n=2 0.8n=4n=0n=0
Median lead timelower is better · n = accepted tasks with lead times · exported measurements
Difficultyclaude-sonnet-5claude-opus-4-8claude-fable-5-1gpt-5.6-terragpt-6-astra
easy47 min=n=50n=0n=0n=0n=0
medium153 minn=1546 minn=2355 minn=1 114 minn=4n=0
hardn=0119 minn=2 39 minn=4n=0n=0

↑ best · ↓ worst · = tied or only value · † n < 3, provisional
Grounds compare within a row. Colour is never the only signal.

Recorded spend by activity

check
$0.00
other
$45.66
persona
$119.36
rebase
$20.63
retro
$11.72
review
$153.77
revise
$66.80
work
$478.23

Runs and hand steps

claude:claude-sonnet-5
156 runs
claude:claude-opus-4-8
51 runs
garden:review
200 runs
garden:persona
24 runs
claude:claude-fable-5-1
12 runs
garden:retro
6 runs
codex:gpt-5.6-terra
12 runs
codex:gpt-6-astra
1 runs
garden:edit
21 runs
garden:kickoff
1 runs
garden:compare
2 runs
codex:gpt-5.6-luna
1 runs
garden:rebase
78 runs
garden:check
202 runs
codex:work
2 runs
garden:trial
1 runs
codex:trial
6 runs
garden:work
1 runs
Hand steps
91 recorded
Manual merges
not supplied
budget_set
3
config_override
2
decision_accepted
5
decision_resolved
21
dispatch_paused
5
dispatch_resumed
8
moved
10
resumed
12
suggestion
21
triaged
4

10 merged in this window

$88.22 recorded window spend

$4.41 / accepted task, as exported.

42% first-pass approval

Scroll to compare models →

Cost per accepted tasklower is better · n = accepted tasks · exported measurements
Difficultyclaude-sonnet-5gpt-5.6-terraclaude-opus-4-8claude-fable-5-1gpt-6-astra
easy$2.50=n=3n=0n=0n=0n=0
medium$5.80n=2 $3.08n=3n=0n=0n=0
hardn=0n=0$7.27=n=2 n=0n=0
First-pass approvalhigher is better · n = reviewed tasks (export cohort) · exported measurements
Difficultyclaude-sonnet-5gpt-5.6-terraclaude-opus-4-8claude-fable-5-1gpt-6-astra
easy67%=n=3n=0n=0n=0n=0
medium50%n=2 0%n=4n=0n=0n=0
hardn=0n=0100%n=2 0%n=1 n=0
Work-run costlower is better · n = work runs · exported measurements
Difficultyclaude-sonnet-5gpt-5.6-terraclaude-opus-4-8claude-fable-5-1gpt-6-astra
easy$1.94=n=3n=0n=0n=0n=0
medium$3.32n=5$0.72n=10n=0n=0n=0
hardn=0n=0$6.24n=2 $11.32n=2 $5.98n=1
Revise roundslower is better · n = accepted tasks · exported measurements
Difficultyclaude-sonnet-5gpt-5.6-terraclaude-opus-4-8claude-fable-5-1gpt-6-astra
easy0.0=n=3n=0n=0n=0n=0
medium1.0n=2 1.3n=3n=0n=0n=0
hardn=0n=00.0=n=2 n=0n=0
Median lead timelower is better · n = accepted tasks with lead times · exported measurements
Difficultyclaude-sonnet-5gpt-5.6-terraclaude-opus-4-8claude-fable-5-1gpt-6-astra
easy80 min=n=3n=0n=0n=0n=0
medium84 minn=2 97 minn=3n=0n=0n=0
hardn=0n=0119 min=n=2 n=0n=0

↑ best · ↓ worst · = tied or only value · † n < 3, provisional
Grounds compare within a row. Colour is never the only signal.

Recorded spend by activity

check
$0.00
other
$32.74
rebase
$0.74
review
$12.62
revise
$4.38
work
$37.75

Runs and hand steps

claude:claude-sonnet-5
10 runs
claude:claude-fable-5-1
2 runs
garden:review
18 runs
claude:claude-opus-4-8
2 runs
codex:gpt-5.6-terra
10 runs
codex:gpt-6-astra
1 runs
garden:edit
21 runs
garden:kickoff
1 runs
codex:gpt-5.6-luna
1 runs
garden:check
29 runs
garden:rebase
4 runs
Hand steps
31 recorded
Manual merges
not supplied
decision_resolved
4
dispatch_paused
1
dispatch_resumed
3
moved
2
suggestion
21

Inspection sheet · all examples are simulated

Every state has something to say.

The live page uses these treatments in place. This sheet keeps them visible for design review.

Running · beyond typical
in leafwork

Workers on independent hosts

Validating the generated context against the project.

Validating the generated context against the project.

typical duration from export · sample n not supplied
Finishing

Review approved.

Every PR reaches the review queue

Waiting for CI before merge. Keep for eight seconds, then settle into recent activity.

Ⅱ Held

Merge held.

Base branch checks failed.

Other workers continue. This reason stays visible until the hold clears.

Ⅱ Paused

codex is paused.

Account quota reached.

New codex work waits; in-flight clocks keep counting. No retry time reported.

× Failedwilted

Checks failed.

The acceptance check returned an error.

Run stopped at 1:08. The failure stays until a retry or a decision resolves it.

Emptyseed

No tasks yet.

Plan work from the phase goals or create the first task.

Drafts awaiting approval

Work is ready to consider.

Review existing drafts before workers can start.

Quiet

Nothing running.

Waiting for an approved task.

Phase and past outcomes stay visible. There is no idle animation.

Disconnected

Updates disconnected.

Last received at 02:44 UTC.

Retain the last snapshot. Clocks show elapsed since start; they cannot prove a worker is still alive.

Unavailable

State not supplied.

Queue unavailable · typical not established · progress not mapped.

Costs and outcomes show —. Missing history is not an empty garden.

Now 2 · design target for /now2 · recorded snapshot + simulated atlas · full specification in docs/design/now-2.md