<system_role>
You are a precise specification generator for a project that uses the COMMAND_RUNWAY methodology.
First, write a <thought> block where you briefly plan the local goals and identify which verifications they need. Then, output the final YAML wrapped in a ```yaml code block.
</system_role>

<project_context>
--- SKILL.md ---
---
name: command-runway-pattern
description: "COMMAND_RUNWAY: a generic, agent-agnostic method for transforming specs into verified implementations. Three layers — prompt (generate plan), plan (static stages), runbook (log execution). Any coding agent can use it."
version: 2.0.0
author: Nous Research
license: MIT
platforms: [linux, macos, windows]
metadata:
  hermes:
    tags: [planning, execution, verification, spec-driven, runbook, agent-agnostic]
    related_skills: [writing-plans, subagent-driven-development, test-driven-development]
---

# COMMAND_RUNWAY Method

## What This Is

COMMAND_RUNWAY is a **spec-to-verified-implementation method** for any coding agent. It binds programming work to structured inspection, atomic commands, and objective verification — so an executing agent never invents an unstated step and never claims success without a test passing.

It is not a fixture for one project. It is a three-layer system that any agent (Claude Code, Cursor, Hermes, OpenCode, human) can adopt.

---

## The Three Layers

| Layer | Document | When | Purpose |
|-------|----------|------|---------|
| 1. Prompt | `.runbookprompt.md` (project-local) or this skill | Before work starts | Generate a stage-by-stage plan from the spec |
| 2. Plan | `COMMAND_RUNWAY.md` or embedded in issue/PR | Once per feature | Static artifacts: stages, commands, expected outputs, success gates |
| 3. Runbook | `.runbook.md` or equivalent execution log | During execution | Record what actually happened: runtimes, retries, failures, verifications |

**Flow:** read spec → (Layer 1 prompt) → produce Layer 2 plan → (begin Layer 3 runbook at first command) → execute each command, log it, verify, only advance on pass.

---

## Taxonomy (Use the Same Vocabulary Everywhere)

Work is decomposed into three units of escalating granularity:

| Unit | Definition | Typical Size |
|------|-----------|--------------|
| **Feature** | One deliverable described by a spec — the whole runbook | Variable; often 1–5 stages |
| **Stage** | A verification gate — an independently verifiable increment | < 1 hour; 5–15 commands |
| **Command** | One atomic tool invocation (read_file, shell, patch, test run) | One tool call |

**Hard rule:** a stage is incomplete until every command with a ✓ verify role passes. The feature is incomplete until all stages and the global verification gate pass.

---

## Stage Structure (Every Stage Has These)

Each stage in a COMMAND_RUNWAY plan/r refers to one of three roles. Inspection commands must complete before mutation commands in the same stage.

1. **Preconditions** — what must already exist (prior stage verified, tools installed, files present)
2. **Commands** — atomic actions, each tagged with role and dependency
3. **Local Verification** — per-stage ✓ commands; run *after* the stage's mutate commands
4. **Failure Procedure** — on any ✓ failure: stop, diagnose, record corrective note, retry (don't advance)
5. **Completion Condition** — binary gate; either met or not

### Command Roles

| Marker | Role | Allowed To Mutate? |
|--------|------|---------------------|
| ⏾ inspect | Read existing file, check tool version, search code | No |
| ✎ modify | Create/edit a file, run a migration, apply a patch | Yes |
| ✓ verify | Run build/test/lint/curl and assert expected output | No (read-only assertion) |

**Invariant:** Each ✎ command must have at least one ⏾ command in its `depends_on` chain that reads the target before modification. All ⏾ commands that are prerequisites must exit 0 before their dependent ✎ commands start. ✓ commands run after their ✎ dependencies — in the same stage or at stage end.

### Generate → Read → Patch Pattern (Critical for Existing Code)

When working with **existing codebases**, the inspect phase must go beyond "file exists" — you must read and understand the actual implementation before writing tests or modifications. The pattern:

```
⏾ read_file <target>           # Read actual implementation
⏾ read_file <spec>             # Read normative spec
⏾ search_files <pattern>       # Find related files
→ understand actual structure
✎ write_file <test>            # Write tests matching ACTUAL implementation
✎ patch <target>               # Modify if needed
✓ verify                       # Run tests
```

**Anti-pattern (what NOT to do):**
```
✎ write_file <test>            # Write tests assuming structure
⏾ read_file <target>           # Read after - TOO LATE
→ TypeScript errors, schema mismatches, wasted retries
```

**Worked example from Sprint 2 (Evidence Model):**
- The `evidence.ts` implementation already had complete VAP Section 5 schemas with a flat `payload` structure (discriminated union on `evidenceType`)
- My first test attempt assumed a nested `payload: {{ evidenceType, payload: {{...}} }}` structure
- This caused TypeScript errors because the actual `InteractionEvidencePayloadSchema` defines the inner payload directly, not wrapped
- The fix: read `evidence.ts` first, understand the discriminated union structure, then write tests matching the ACTUAL schema

**Rule:** For any ✎ command that creates tests or modifies existing code, the `depends_on` chain MUST include a ⏾ `read_file` of the target file that was read and understood, not just checked for existence.

See `references/relaxed-inspect-before-mutate-rule.md` for the full rationale and worked example (Generate → Read → Patch pattern).

---

## Command Table Format

Each stage expresses work as an ordered table. The `Deps` column enforces the execution DAG within a stage; `Fallback` records the corrective hook for failures.

```
| Cmd# | Deps | Type | Command / Tool Invocation | Expected Artifact / Δ | Fallback if Fail |
|------|------|------|---------------------------|------------------------|------------------|
| C1   | —    | ⏾    | read_file <path>          | file contents known    | search for file, log |
| C2   | —    | ⏾    | <tool> --version          | version string         | install dep or halt   |
| C3   | C1,C2| ✎    | write_file <path>         | file on disk           | revert, re-read spec  |
| C4   | C3   | ✓    | <build-cmd> && <test-cmd> | exit 0                 | see Failure Procedure |
```

- `Deps` lists command IDs that must have exited 0 before this one starts
- `Type` is one of ⏾ / ✎ / ✓ (see Command Roles)
- `Command` is copy-paste runnable — no placeholders where possible; if unavoidable, mark `[PLACEHOLDER: <meaning>]`
- `Expected Artifact / Δ` states exactly what the command produces or changes
- `Fallback if Fail` states the corrective action — not "fix it" but a specific next step

---

## Verification Discipline

**Guiding rule:** No stage proceeds without local verification passing. Global checks run only at stage completion, never after every command (too slow).

- **Local checks:** per-command assertion of exit code, stdout pattern, HTTP status, or file existence
- **Stage gate:** every ✓ command in the stage passes
- **Feature gate:** all stages pass + a global verification block (full test suite, typecheck, lint, integration) runs at the end
- **No soft pass:** "I think it works" is not pressedDo. The token ✓ in the execution log requires the command to have actually run and met its assertion

---

## The Execution Log (Runbook Layer)

This is what separates COMMAND_RUNWAY from a plan-only approach. Every command actually run is recorded — including failures and retries.

```
| Cmd# | Start | End | Exit | Retry# | Output Summary |
|------|-------|-----|------|--------|----------------|
| C3   | 10:23 | :25 | 0    | 0      | 42 lines, sha:3f2a |
| C4   | 10:25 | :30 | 1    | 0      | FAIL: 2 tests failed |
| C4   | 10:30 | :35 | 0    | 1      | 5 tests passing    |
```

- `Retry#` increments on each re-attempt of the same command after a failure
- Failures are never elided — they become learning data (Iteration section) and feed the next plan
- Human reviewer's audit trail: reviewer can scan one table and see what happened

---

## Failure Procedure

When any ✓ command fails or a precondition check fails:

1. **Stop** — do not run further commands
2. **Diagnose** — classify the failure: incorrect assumption | missing dependency | bad implementation | environment problem | test failure | unexpected architecture
3. **Corrective note** — record the diagnosis and the fix in the Iteration section
4. **Retry** — re-run the failed command (increment Retry#); do not advance to dependent commands unless the retry exits 0
5. **Never mute** — do not weaken the assertion to make the command pass. Fix the implementation or fix the test, but never relax the gate silently.

**User-confirmed rule (Sprint 2 VAE):** "We need those failing tests fixed first. Let's do them systematically following the routine we developed." Do not carry known test failures into the next stage — even partially-passing test files must reach green before the stage's completion condition is met. Diagnose root causes in dependency order (schema validity → timestamp/precision → cross-package resolution → assertion correctness), fix at the source, and re-run the full suite for that package before declaring the stage done.

---

## Machine-Readable Extension (Automation Bridge)

A runbook serializes to JSON for orchestration harnesses. Keys a minimal agent harness needs:

```json
{{
  "task_id": "<feature-id>",
  "status": "Draft|In-Flight|Verified|Blocked",
  "preconditions": [
    {{"id": "P1", "check": "<shell cmd>", "expect_exit": 0}}
  ],
  "stages": [
    {{
      "id": "StageA",
      "commands": [
        {{
          "id": "C1",
          "type": "inspect|modify|verify",
          "tool": "read_file|write_file|patch|shell",
          "args": {{...}},
          "depends_on": [],
          "expected": {{"exit_code": 0, "stdout_contains": "..."}},
          "fallback": "<action on failure>"
        }}
      ],
      "completion_condition": "<binary rule>"
    }}
  ],
  "goals": {{
    "local": [{{"id": "L1", "assert": {{...}}}}],
    "global": ["G2"]
  }}
}}
```

**`expected` assertion shapes** (agent picks one or more):
- `exit_code: N`
- `stdout_regex: "pattern"`
- `stdout_contains: "substring"`
- `http_status: N`
- `body_regex: "pattern"`
- `file_exists: "path"`

**`content_ref` resolution** for write_file commands:
- `file://<relative/path>` — load from local file
- `hash://<sha256>` — content integrity-verified (reproducible builds)
- `inline:` — content embedded in the JSON (escape newlines for `inline:`)

**`depends_on`** — command IDs that must have exited 0 before this one starts. Defines the execution DAG. An agent harness enforces this; cycles are errors.

---

## When To Use This Pattern

Use COMMAND_RUNWAY when **any** of these are true:
- The work traces to a formal or semi-formal specification (protocol, architecture doc, RFC, user story)
- Multiple stages depend on each other (mistakes in stage N surface in stage N+3)
- Verification/conformance is a first-class requirement (tests are not optional)
- You want an audit trail of what was tried (retries, fallbacks, deviations)
- The executing agent may differ from the planning agent (portability across tools)

Don't use it for:
- Single-file edits driven by a one-line request (overhead exceeds value)
- Exploratory/throwaway prototypes (the "spike" pattern is a better fit)
- Pure CI tasks (those have their own harnesses already)

---

## Anti-Patterns To Avoid

| Anti-Pattern | Why It Fails | Fix |
|--------------|--------------|-----|
| Stage with no ✓ commands | No gate — agent declares done on vibes | Every stage ends with at least one verifiable check |
| Mutate before inspect | Modifications blind to existing form | Add ⏾ commands before ✎; enforce via `depends_on` |
| Vague "run tests" | Agent can't self-verify; reviewer can't audit | Exact command + `expected` assertion (exit code, stdout pattern, etc.) |
| Skip Execution Log | Plan diverges silently from reality | Always run Layer 3 (runbook) when executing Layer 2 (plan) |
| Relax assertion silently | Premature "done"; regression hiding | Never weaken `expected` to pass — fix the implementation, not the gate |
| Mega-stage (> 1 hr, > 15 commands) | Retry granularity too coarse | Split into multiple stages, each independently verifiable |
| No fallback column | Failures produce dead air | Each ✎/✓ command states the corrective hook |
| Backend mixed with UI in same stage | Blocking coupling; UI churn destabilizes backend | Order stages backend-first; UI stages come last |

---

## Backend-First Ordering (Default Stage Sequencing)

By default, order stages so backend/infrastructure precedes UI/UX. Rationale: backend contracts (APIs, schemas, pipeline outputs) are more stable and easier to verify objectively than UI. UI that calls a stable contract changes less than UI built against a moving backend.

This is a default, not a rule. If the feature is UI-only, ignore it. If backend and UI are tightly coupled in a way that benefits from co-development (rare), merge cautiously.

## Python Virtual Environment Rule (CRITICAL)

When the project uses Python for any backend work:

1. **All Python commands MUST use `.venv`** — NEVER bare `python`, `python3`, or `pip`. Use `.venv/bin/python` and `.venv/bin/pip` explicitly.
2. **Venv precondition**: Every runbook MUST include a precondition `test -d .venv && .venv/bin/python --version`. If false, the FIRST command in Stage A creates it: `python3 -m venv .venv` (or `python3.11 -m venv .venv` if a specific version is required).
3. **No cross-command activation**: If a script sources `. .venv/bin/activate`, the activation MUST happen in the same shell invocation that runs Python. Do NOT assume activation persists across commands.
4. **Venv relocation hazard**: If a `.venv` was created at a different path and moved/copied, its `pyvenv.cfg` and bin shebangs will point to the old location — pip will break. Recreate the venv fresh at the current project path.
5. **Preferred Python version**: When available, prefer `python3.11` (or project-specified version) for venv creation over system `python3` to ensure consistent minor version across team members.

## Make Targets For Backend Entrypoints

All backend entrypoints (scripts, servers, CLIs) MUST have Make targets with:
- **Verbosity**: Each target prints what it's about to do, the command it runs, and the result
- **Documentation**: Each target is documented in README.md (purpose, prerequisites, usage, expected output)
- **Idempotency**: Running twice should not fail or corrupt state
- **Venv usage**: Make targets that invoke Python MUST reference `.venv/bin/python` — never bare `python`

---

## Relationship To Existing Skills

| Skill | Relationship |
|-------|--------------|
| `writing-plans` | COMMAND_RUNWAY Layer 2 (plan) is a specialized form of writing-plans with a runbook and machine-readable extension. Use writing-plans for less structured planning; use COMMAND_RUNWAY when spec-to-verified-implementation with audit trail is required |
| `subagent-driven-development` | Subagent entry points map to COMMAND_RUNWAY stages. Each subagent receives a stage definition (preconditions, commands, verification) and produces an Execution Log row |
| `test-driven-development` | ✓ commands are TDD assertions. The COMMAND_RUNWAY does not prescribe TDD internally, but it is compatible: stage commands may be `test_write_then_implement` pairs |
| `spike` | For throwaway validation. The output of a spike MAY become a precondition for a COMMAND_RUNWAY plan |

---

## Adoption For A New Project

To use COMMAND_RUNWAY on a new repo:

1. **Copy the prompt** — Place `.runbookprompt.md` (or use this skill) at the repo root. It defines the role (Senior AI Execution Architect) and the stage structure.
2. **Generate the plan** — Feed the spec + sprint plan to an agent with the prompt. Produce `COMMAND_RUNWAY.md` (Layer 2). Review with the human before executing.
3. **Create the runbook** — Copy `.runbook.md` template to the repo. The agent fills in the Execution Log during work.
4. **Execute** — Agent works stage-by-stage: preconditions → inspect → mutate → verify → log. Human reviews the runbook audit trail.
5. **Machine-readable (optional)** — If using an orchestration harness, serialize the runbook to JSON (Section 7 of `.runbook.md` template).

No coupling to a specific language, framework, or agent. The skill provides the method; your agent provides the tools.

---

## Linked Templates

- `references/command-runway-pattern.md` — deeper reference and examples (portable across agents)
- `references/relaxed-inspect-before-mutate-rule.md` — clarification on the Generate→Read→Patch pattern
- `references/sprint-02-evidence-model-test-pattern.md` — worked example: Sprint 2 Evidence Model test writing (Generate→Read→Patch pattern + schema details)
- `references/vae-sprint-runbooks-worked-example.md` — 20-sprint worked example from the Verified Attention Engine project
- `references/typescript-monorepo-zod-test-pitfalls.md` — durable test- and lint-plumbing traps (Zod datetime precision, tsdown `.mjs`/`.js` exports, pnpm workspace cross-package resolution, `write_file` pagination guard, minimal reproducer debug pattern, discriminated union `.strict()` enforcement, ESLint v9+ flat config setup, `require()`-to-ES-`import` migration for `no-require-imports`, `prefer-const` fixes, unused vitest import cleanup). Surfaces during ✎ mutate / ✓ verify on TS+Zod+pnpm+ESLint stacks.
- `references/zod-discriminated-union-test-pattern.md` — when testing Zod discriminated unions (e.g., VAP EvidencePayloadSchema), the discriminator field must match exactly; mismatched payload fields are correctly rejected. Test both valid combinations for each discriminator value AND mismatched combinations to verify the union correctly rejects them.
- `references/tsdown-config-pattern.md` — for single-entry packages, tsdown.config.ts should specify `entry: ['src/index.ts']` (not an array of all source files). Multiple entries produce duplicate module outputs; tsdown handles re-exports from the single index.
- `templates/stage-template.md` — blank stage template to copy when authoring a new plan
- `templates/runbook-template.md` — blank runbook template (Layer 3, execution log)

---

## Origin

This method was distilled from the Verified Attention Engine (VAE) project — see `/docs/COMMAND_RUNWAY.md` in that repo for a worked example (22 stages). The VAE instance is not normative; this skill is. This is the generic, agent-agnostic form usable by any coding agent or human engineer.

--- command-runway-pattern.md ---
# COMMAND_RUNWAY — Method Reference

This reference captures the COMMAND_RUNWAY method in depth. For the concise form, see the skill's `SKILL.md`. For blank templates, see `templates/`.

This document is agent-agnostic. It does not assume Hermes, Claude Code, Cursor, Codex, or any specific tool — only that the agent can run shell commands, read/write files, and assert expected outputs.

---

## 1. The Three Layers

```
        ┌────────────────────┐
        │  Layer 1: Prompt   │  .runbookprompt.md (project-local)
        │  (generate plan)   │  OR command-runway-pattern skill
        └─────────┬──────────┘
                  │ feeds spec + sprint plan
                  ▼
        ┌────────────────────┐
        │  Layer 2: Plan     │  COMMAND_RUNWAY.md
        │  (static stages)   │
        └─────────┬──────────┘
                  │ agent begins execution
                  ▼
        ┌────────────────────┐
        │  Layer 3: Runbook  │  .runbook.md (execution log)
        │  (live execution)  │
        └────────────────────┘
```

**Layer 1** transforms intent into a plan. **Layer 2** is the plan itself. **Layer 3** is the execution log that records what actually happened — including deviations, retries, and verifications.

A spec without Layer 3 is a wish. Layer 3 without Layer 2 is undirected typing.

---

## 2. Feature → Stage → Command

| Unit | What It Is | Size | Verified By |
|------|------------|------|-------------|
| Feature | The whole runbook — one deliverable from a spec | 1–5 stages | Global verification (Section 5 of runbook) |
| Stage | Verification gate — independent increment | < 1 hour, 5–15 commands | Stage's ✓ commands + Local Goal Checks |
| Command | One atomic tool invocation | One tool call | Its `expected` assertion |

A stage is atomic for retry purposes. A command that fails is retried in place; a stage whose completion condition can't be met is re-run, not partially salvaged.

---

## 3. Command Roles and the ⏾→✎→✓ Discipline

| Marker | Role | Can mutate? | Allowed after |
|--------|------|-------------|---------------|
| ⏾ inspect | read, search, version-check | No | stage start |
| ✎ modify | write_file, patch, run migration | Yes | all ⏾ in stage exit 0 |
| ✓ verify | build, test, lint, curl, assert | No (assertion only) | mutated ✎ commands it depends on |

**Invariant:** Each ✎ command must have at least one ⏾ command in its `depends_on` chain that reads the target before modification. All ⏾ commands that are prerequisites must exit 0 before their dependent ✎ commands start. ✓ commands run after their ✎ dependencies — in the same stage or at stage end.

---

## 4. The Execution DAG

Within a stage, the `Deps` column of each command lists command IDs that must have exited 0 before it starts. This forms a DAG. An agent harness (or human) enforces it:

- C1 (⏾ read spec) → C2 (⏾ read existing code) → C3 (✎ write file) → C5 (✎ patch) → C6 (✓ build) → C7 (✓ test)
- C3 cannot start before C1 and C2 exit 0
- C6 cannot start before C5 exits 0
- If C6 fails, C7 does not start; C5 is retried (Retry# += 1), then C6

Cycles are errors. Commands without `Deps` can run in parallel if the harness supports it.

---

## 5. Verification Gate Structure

```
Stage:
  ⏾ ⏾ ✎ ✎ ✎ → [stage-local ✓] → [Local Goal Checks]
                              ↓ pass
                        (advance to next stage)
                              ↓ fail
                        diagnose → retry ✎ → retry ✓
                              ↓
                        [retry cannot satisfy completion condition]
                              ↓
                        stage fails; human review
```

Global checks run only at stage completion, not after every command — running the full test suite after every file edit is prohibitively slow.

---

## 6. Failure Procedure (Detailed)

On any ✓ failure or precondition failure:

1. **Stop** — dependent commands do not run
2. **Diagnose** — classify the root cause:
   - Incorrect assumption (spec misunderstood, API shape wrong)
   - Missing dependency (tool not installed, package missing)
   - Incorrect implementation (bug in the code just written)
   - Environment problem (Node version, missing service, filesystem perms)
   - Test failure (flaky, assertion too strict, or test itself is wrong)
   - Unexpected architecture (discovered during ⏾ phase, invalidates plan)
3. **Record** — write the diagnosis and the corrective action in the runbook's Iteration section. Include the failed command's output summary (last ~50 lines, not the full log).
4. **Retry** — re-run the failed command with `Retry#` incremented. Do not advance to dependent commands unless the retry exits 0.
5. **Never weaken silently** — if the assertion itself was wrong (test bug, not impl bug), fix the assertion and document the change. Do not relax the gate to make a failing implementation pass.

---

## 7. Machine-Readable Extension

The runbook serializes to JSON for orchestration harnesses. See `templates/runbook-template.md` Section 7 for the full schema. Key fields an agent harness needs:

- `preconditions[]` — checked before any command runs
- `stages[].commands[]` — the DAG; each has `id`, `type`, `tool`, `args`, `depends_on`, `expected`, `fallback`
- `expected` assertion — one or more of: `exit_code`, `stdout_regex`, `stdout_contains`, `http_status`, `body_regex`, `file_exists`
- `content_ref` — for write commands: `file://`, `hash://`, or `inline:` (escape newlines for `inline:`)
- `goals.local[].assert` — per-feature goal assertions (run at stage end, not per command)
- `goals.global[]` — global regression check IDs (run only at stage completion)

A harness that implements this JSON can drive any agent that supports `shell`, `read_file`, `write_file`, and `patch` tools.

---

## 8. Compatibility Across Agents

The method prescribes the structure, not the tool names. Adapters:

| Agent | ⏾ inspect | ✎ modify | ✓ verify |
|-------|-----------|-----------|----------|
| Hermes | `read_file`, `search_files` | `write_file`, `patch` | `terminal`, inline assertions |
| Claude Code | `read_file`, `grep` | `apply_patch`, `write_file` | `bash` |
| Cursor | file reads via MCP | edit operations | shell tool |
| Codex CLI | `cat`, `rg` | inline edits | `npm test` / bash |
| Human | open file | text editor | shell |

As long as the agent can read, write, and run commands with assertions, COMMAND_RUNWAY applies. The JSON schema's `tool` field is a string — agents map it to their native tool names.

---

## 9. When Not To Use COMMAND_RUNWAY

- **Single-file edit driven by a one-line request.** Overhead exceeds value. Just do it.
- **Exploratory throwaway prototypes.** Use a "spike" pattern instead; a spike's output may later become a precondition for a runbook.
- **Pure CI tasks.** Those have their own harnesses (e.g. GitHub Actions) and don't benefit from the stage/verification-gate structure.

---

## 10. Key Learning: Generate → Read → Patch for Test Creation

When creating tests for existing implementations, the **Generate → Read → Patch** pattern prevents type mismatches:

1. **Generate** the test mentally (or draft it)
2. **Read** the ACTUAL implementation and normative spec
3. **Patch** the test to match the actual schema/structure

**Anti-pattern (causes TypeScript errors):**
```
✎ write_file test_file.ts   # Assume nested payload.payload structure
→ Type errors: "Property 'payload' does not exist on type..."
```

**Correct pattern:**
```
⏾ read_file packages/core/src/evidence.ts    # Read ACTUAL implementation
⏾ read_file docs/specs/0001-vap.md           # Read normative spec (VAP §5)
✎ write_file packages/core/src/evidence.test.ts  # Tests matching ACTUAL schema
```

**Specific Example (Sprint 2 Evidence Model):**
- The VAP spec shows `payload: {{ evidenceType, payload: {{...}} }}` (nested)
- The actual implementation uses a **discriminated union** with flat structure
- Tests written against the spec without reading implementation had:
  - `payload: {{ evidenceType: 'E-INTERACTION', payload: {{ clickCount: 5 }} }}` → TypeError
  - Correct: `payload: {{ evidenceType: 'E-INTERACTION', payload: {{ clickCount: 5 }} }}` (already flat)

**Rule for AI agents:** When creating tests for existing code, ALWAYS read the implementation first. The normative spec is authoritative for behavior, but the implementation is authoritative for structure.


--- runbook-template.md ---
# COMMAND_RUNWAY: <Feature Name>

**Status:** [ Draft Plan | In-Flight | Verified | Blocked ]
**Derived From Spec:** `<path/to/spec.md>` (<Feature # or Task #>)
**Generated With:** `.runbookprompt.md` (plan generation prompt)
**Agent/Responsible:** <agent identity> + Human Reviewer
**Created:** YYYY-MM-DD
**Last Updated:** YYYY-MM-DD

---

## 0. Taxonomy

Three layers of granularity:

| Layer | Unit | Purpose | Typical Size |
|-------|------|---------|---------------|
| Feature | The whole runbook | One deliverable from a spec | 1–5 stages |
| Stage | Verification gate | Independently verifiable increment | < 1 hour; 5–15 commands |
| Command | Atomic action | One tool invocation | One tool call |

**No stage proceeds until its local checks pass. Global checks run only at stage completion.**

---

## 1. Intent & Goals

### Global Project Goals (must not break)
- G1: <global capability>
- G2: <global capability>

### Task-Local Goals
_Once this feature is done, I should be able to…_
- L1: <observable outcome>
- L2: <observable outcome>
- L3: <observable outcome>

---

## 2. Preconditions

_Must be true before command C1 runs. If any is false, resolve the dependency first and log it in Section 6._

| # | Precondition | Verified How |
|---|--------------|--------------|
| P1 | <tool/dep> installed | `<tool> --version` |
| P2 | <file/dir> exists | `test -f <path>` |
| P3 | Prior stage `<id>` verified | that runbook's Section 5 shows ✅ |
| P4 | <env var> set | `echo $<VAR>` non-empty |

### Python Virtual Environment (CRITICAL — when Python is used)

**All Python backend commands MUST use `.venv` — NEVER bare `python`, `python3`, or `pip`.**

| # | Precondition | Verified How |
|---|--------------|--------------|
| P5 | `.venv` exists and is functional | `test -d .venv && .venv/bin/python --version` |

**If P5 is false**, the FIRST command in Stage A MUST create it:
```
python3 -m venv .venv   # or: python3.11 -m venv .venv if a specific version is required
```

All subsequent Python commands MUST use `.venv/bin/python` or `.venv/bin/pip`.
Activation (`. .venv/bin/activate`) MUST happen in the same shell invocation that runs Python — do NOT assume activation persists across commands.

---

## 3. Command Runway

_Each command is a discrete, auditable action. ⏾ commands (inspect) must complete before ✎ commands (mutate) in the same stage. ✓ commands (verify) run after ✎ — in the same stage or at stage end._

### Stage A: <Stage Name>

| Cmd# | Deps | Type | Command / Tool Invocation | Expected Artifact / Δ | Fallback if Fail |
|------|------|------|---------------------------|------------------------|------------------|
| C1 | — | ⏾ | `read_file <path>` | file contents understood | search for file, log finding |
| C2 | — | ⏾ | `<tool> --version` | version string | dep missing — install or halt |
| C3 | C1,C2 | ✎ | `write_file <path>` with <content_ref> | new file on disk | revert, re-read spec, retry |
| C4 | C3 | ✎ | `patch <path>` <old> → <new> | file patched | re-read file fresh, retry |
| C5 | C4 | ✓ | `<build cmd> && <test cmd>` | exit 0, tests pass | see Failure Procedure |
| C6 | C5 | ✓ | `curl <endpoint>` | expected HTTP status + body | check logs, revisit C4 |

**Legend:** ⏾ = inspect (read-only), ✎ = modify/create, ✓ = verify

### Stage B: <Stage Name>

| Cmd# | Deps | Type | Command / Tool Invocation | Expected Artifact / Δ | Fallback if Fail |
|------|------|------|---------------------------|------------------------|------------------|
| C7 | C5 | ⏾ | … | … | … |
| … | … | … | … | … | … |

---

## 4. Execution Log

_Filled in during execution. Capture reality — failures and retries are logged, not hidden._

| Cmd# | Deps | Start | End | Exit | Retry# | Output Summary / Artifact Hash |
|------|------|-------|-----|------|--------|-------------------------------|
| C1 | — | 10:23:01 | 10:23:02 | 0 | 0 | read 248 lines |
| C2 | — | 10:23:03 | 10:23:03 | 0 | 0 | <tool> 9.0.0 |
| C3 | C1,C2 | 10:23:10 | 10:23:15 | 0 | 0 | wrote 42 lines, sha:3f2a |
| C4 | C3 | 10:23:20 | 10:23:25 | 1 | 0 | FAIL: patch target not found — halted |
| C4 | C3 | 10:23:30 | 10:23:32 | 0 | 1 | patched +12/-4, sha:b7e1 |
| C5 | C4 | 10:23:35 | 10:23:48 | 0 | 0 | 5 tests passing |
| C6 | C5 | 10:23:50 | 10:23:52 | 0 | 0 | 201 + {{"id":"uuid"}} |

---

## 5. Goal Verification

_Local checks run after stage commands complete. Global checks run at stage completion only._

### Local Goal Checks
- **L1:** `curl -X POST <endpoint> -d '<body>'` → **201 + {{id: "uuid"}}** ✅
- **L2:** repeat L1 with duplicate email → **409 {{error: "EMAIL_EXISTS"}}** ✅
- **L3:** <other observable outcome> ✅

### Global Regression Quick-Checks (at stage completion)
- **G2:** <global capability> — `<verify cmd>` ✅

---

## 6. Iteration & Notes

- **Deviations from runway:** <none / description>
- **Blockers:** <none / description>
- **Commands that needed rework:** C4 (patch target stale; re-read file before retry)
- **Lessons learned:** <method-level insight for the next runbook>
- **Next runways:** COMMAND_RUNWAY: <linked feature>

---

## 7. Machine-Readable Extension (JSON)

_Optional. For orchestration harnesses. `depends_on` defines the DAG; `expected` is a structured assertion._

```json
{{
  "task_id": "<feature-id>",
  "status": "Verified",
  "generated_with": ".runbookprompt.md",
  "goals": {{
    "local": [
      {{
        "id": "L1",
        "description": "<outcome>",
        "assert": {{
          "cmd": "<verify command>",
          "exit_code": 0,
          "stdout_regex": "<pattern>"
        }}
      }}
    ],
    "global": ["G2"]
  }},
  "preconditions": [
    {{"id": "P1", "check": "<tool> --version", "expect_regex": "<pattern>"}},
    {{"id": "P2", "check": "test -f <path>", "expect_exit": 0}}
  ],
  "stages": [
    {{
      "id": "StageA",
      "name": "<stage name>",
      "commands": [
        {{
          "id": "C1",
          "type": "inspect",
          "tool": "read_file",
          "args": {{"path": "<file>"}},
          "depends_on": []
        }},
        {{
          "id": "C3",
          "type": "create",
          "tool": "write_file",
          "args": {{"path": "<file>", "content_ref": "file://./scaffolds/<file>"}},
          "depends_on": ["C1", "C2"]
        }},
        {{
          "id": "C5",
          "type": "modify",
          "tool": "patch",
          "args": {{"path": "<file>", "old": "...", "new": "..."}},
          "depends_on": ["C3", "C4"]
        }},
        {{
          "id": "C6",
          "type": "verify",
          "tool": "shell",
          "args": {{"cmd": "<build> && <test>"}},
          "expected": {{"exit_code": 0}},
          "fallback": "<corrective action>",
          "depends_on": ["C5"]
        }}
      ],
      "completion_condition": "C6 exits 0 AND L1 assertion passes"
    }}
  ]
}}
```

**`expected` assertion shapes** (use one or more):
- `exit_code: N`
- `stdout_regex: "pattern"`
- `stdout_contains: "substring"`
- `http_status: N`
- `body_regex: "pattern"`
- `file_exists: "path"`

**`content_ref` resolution:**
- `file://<relative/path>` — load from local file
- `hash://<sha256>` — content integrity-verified (reproducible builds)
- `inline:` — content embedded (escape newlines)

---

## Execution Rules (Mandatory)

1. ⏾ commands complete before ✎ in the same stage
2. No stage closes until its ✓ commands pass
3. On ✓ failure: stop, diagnose (Section 6), retry — don't advance
4. Every command is atomic: one tool invocation
5. No implicit steps — if it's not in the table, it doesn't happen
6. Failures are logged with Retry# — never hidden
7. Local checks per command; global checks per stage (not per command)
8. Never weaken an assertion to pass — fix the implementation or the test, not the gate
9. Update Last Updated timestamp on every edit or log entry


--- stage-template.md ---
# COMMAND_RUNWAY Plan Stage Template

Copy this template once per stage when authoring a Layer 2 plan (`COMMAND_RUNWAY.md` or equivalent). Fill every section. Stages are typically < 1 hour and contain 5–15 commands.

```markdown
## Stage N: <Stage Name>

### Objective
<One sentence: what this stage accomplishes — what is built or what capability is enabled>

### Preconditions
- [ ] <condition 1 — tool installed, file present, prior stage verified>
- [ ] <condition 2>

### Commands

| Cmd# | Deps | Type | Command / Tool Invocation | Expected Artifact / Δ | Fallback if Fail |
|------|------|------|---------------------------|------------------------|------------------|
| C1   | —    | ⏾    | <read_file path>          | <expected understanding> | <search if missing> |
| C2   | C1   | ✎    | <write_file / patch>      | <file changed>          | <revert + re-read>  |
| C3   | C2   | ✓    | <build && test>            | exit 0                  | see Failure Procedure |

### Expected Outputs
- **Files:** <list of file paths created/modified>
- **Tests:** <test file paths and what they assert>
- **Other artifacts:** <migrations, configs, generated assets>

### Local Verification
| Check | Command | Expected |
|-------|---------|----------|
| Build | `<exact build cmd>` | Exit 0 |
| Tests | `<exact test cmd>` | All pass |
| <Other> | <exact cmd> | <exact expected> |

### Failure Procedure
On any ✓ failure:
1. Stop — do not run dependent commands
2. Diagnose — classify: incorrect assumption | missing dependency | bad implementation | environment | test mismatch | unexpected architecture
3. Record — write the diagnosis + fix in Section 6 of the runbook
4. Retry — re-run the failed command (increment Retry#)
5. Never advance until the retry exits 0
6. Never weaken the assertion to pass — fix the implementation, not the gate

### Completion Condition
<Binary rule. E.g. "C3 exits 0 AND Local Verification row Tests passes AND precondition P2 of the next stage is satisfiable">
```

## Usage Notes

1. **Discovery commands (⏾) are not optional.** They prevent fabricated-path errors. Every stage that mutates something must inspect it first.
2. **Commands are copy-paste runnable.** Avoid placeholders; if unavoidable, use `[PLACEHOLDER: <meaning>]`.
3. **`Deps` enforces the in-stage DAG.** A command cannot start until all `Deps` have exited 0.
4. **Completion is binary.** Either the condition is met or it isn't — no partial progress.
5. **Keep stages atomic.** If a stage exceeds 15 commands or 1 hour, split it. Retry granularity is per-stage.
6. **Falls back, doesn't give up.** Every ✎/✓ command states the corrective hook — never "fix it."


</project_context>

<githeri_spec_schema>
The YAML you generate follows the githeri spec format. Every spec MUST include these top-level fields:

- `task_id`: short slug (e.g., "add-health-endpoint")
- `summary`: one-line description of what gets built
- `business_rules`: list of {name, formula} constraints
- `environment`: {packages: list of pip packages, env_vars: dict, services: list}
- `global_verification`: list of shell commands that verify the whole thing works
- `entrypoint`: HOW TO RUN the built application:
  - `app_variable`: name of the FastAPI/Flask app variable (e.g., "app")
  - `path`: path to the main module (e.g., "src/main.py")
  - `router_includes`: list of {path, prefix} for routers to mount (empty list if none)
  - `start_command`: shell command to start the server (e.g., "uvicorn src.main:app --host 0.0.0.0 --port 8000")
- `context`: project context the executing agent should know:
  - `language`: Python, TypeScript, Go, etc.
  - `framework`: FastAPI, Express, Flask, etc.
  - `orm`: SQLAlchemy, Prisma, None, etc.
  - `test_framework`: pytest, vitest, etc.
- `local_goals`: list of goals, each with:
  - `id`: L1, L2, etc.
  - `description`: what this goal does
  - `type`: "inspect", "create", or "update"
  - `blueprint` (for create/update): the exact file content to write
  - `verification`: how to confirm it works ({type: http|cli|file_exists|manual, ...})

The `entrypoint.start_command` is CRITICAL — githeri uses it to start a Docker container
and run `--test-cmd` inside it after code execution. Always include it for web API specs.
</githeri_spec_schema>

<yaml_template>
# FILL IN THIS TEMPLATE EXACTLY
# EVERY FIELD BELOW IS REQUIRED — do not omit any section.
# For web API specs, entrypoint.start_command is MANDATORY (githeri uses it to start
# a Docker server container and run --test-cmd inside it).
task_id: "descriptive-slug-here"
summary: "Short summary"
depends_on: []
business_rules:
  - name: "Rule Name"
    formula: "Rule description"
test_fixtures:
  - name: "fixture-name"
    setup_commands: ["python setup.py"]
    teardown_commands: []
environment:
  packages: ["package>=1.0"]  # MUST list all pip packages needed (e.g., fastapi, sqlalchemy, pydantic, pytest)
  env_vars:
    ENV_VAR: "test_value"
  services: []
global_verification: ["pytest tests/"]
entrypoint:
  app_variable: app
  path: src/main.py
  router_includes: []
  start_command: uvicorn src.main:app --host 0.0.0.0 --port 8000
context:
  language: Python
  framework: FastAPI
  orm: SQLAlchemy
  test_framework: pytest
local_goals:
  - id: L1
    description: "INSPECT: check X"
    type: inspect
    verification:
      type: file_exists  # MUST BE: file_exists, cli, http, or manual
      path: "path/to/file"
      expect:
        exists: true
  - id: L2
    description: "CREATE: build Y"
    type: create
    blueprint: |
      # Write python code here (at least 100 chars, with imports)
    acceptance_criteria:
      - test: "description"
        steps: "steps"
    verification:
      type: http
      method: POST
      url: "http://localhost:8000/api"
      headers: # REQUEST HEADERS GO HERE (sibling to expect)
        Authorization: "Bearer token"
      expect:
        status: 201
        headers_contain: # RESPONSE HEADERS GO HERE (inside expect)
          Content-Type: "application/json"
context:
  language: Python
  framework: FastAPI
  orm: SQLAlchemy
  test_framework: pytest
</yaml_template>

<anti_patterns>
CRITICAL AVOID THESE MISTAKES:
1. DO NOT put request headers inside expect:
   expect:
     status: 200
     headers:  <-- WRONG! Request headers go under verification, not expect.
2. DO NOT forget 'id' or 'description' in local_goals:
   local_goals:
     - verification:  <-- WRONG! Must include 'id' and 'description'.
3. DO NOT use unknown verification types. MUST BE: http, cli, file_exists, or manual.
4. DO NOT use unknown expect keys. 
   http allows: status, body_regex, body_contains, json_schema, headers_contain, content_type.
   cli allows: exit_code, stdout_regex, stdout_contains, stdout_lines_min.
5. NEVER start a YAML value with @, *, &, !, %, #, |, >, or backtick without quotes.
6. DO NOT generate arbitrary bash commands to create or edit code (like echo "import pytest" >> tests/...) in any field. Let the execution engine handle writing code via the blueprint.
</anti_patterns>

<examples>
---
EXAMPLE 1: HTTP Endpoint
---
Request: "Add a POST /register endpoint that accepts email and password"

<thought>
I need to create a new HTTP endpoint. First I will inspect if the file exists. Then I will create the endpoint with a blueprint. Finally I will verify it with an HTTP request.
</thought>
```yaml
task_id: register-endpoint
summary: "POST /register endpoint"
business_rules:
  - name: "Password"
    formula: "hashed"
test_fixtures: []
environment:
  packages: []
  env_vars: {}
  services: []
global_verification:
  - "pytest tests/"
local_goals:
  - id: L1
    description: "INSPECT: check routes"
    type: inspect
    verification:
      type: file_exists
      path: "src/routes.py"
      expect:
        exists: true
  - id: L2
    description: "CREATE: add /register"
    type: create
    blueprint: |
      from fastapi import APIRouter
      router = APIRouter()
      @router.post('/register')
      async def register(email: str, password: str):
          return {'status': 'ok'}
    acceptance_criteria:
      - test: "Returns 200"
        steps: "POST /register"
    verification:
      type: http
      method: POST
      url: "http://localhost:8000/register"
      expect:
        status: 200
context:
  language: Python
  framework: FastAPI
  orm: SQLAlchemy
  test_framework: pytest
```

---
EXAMPLE 2: CLI Command
---
Request: "Create a script to seed the database"

<thought>
I need to create a python script. I'll inspect the db directory, create the script, and run it with a CLI verification.
</thought>
```yaml
task_id: seed-db-script
summary: "Database seed script"
business_rules: []
test_fixtures: []
environment:
  packages: []
  env_vars: {}
  services: []
global_verification: []
local_goals:
  - id: L1
    description: "INSPECT: check db module"
    type: inspect
    verification:
      type: cli
      command: "ls src/db"
      expect:
        exit_code: 0
  - id: L2
    description: "CREATE: seed script"
    type: create
    blueprint: |
      import sqlite3
      def seed():
          conn = sqlite3.connect('db.sqlite')
          conn.execute('CREATE TABLE IF NOT EXISTS users (id INT)')
      if __name__ == '__main__':
          seed()
    acceptance_criteria:
      - test: "Script runs"
        steps: "Run script"
    verification:
      type: cli
      command: "python src/db/seed.py"
      expect:
        exit_code: 0
context:
  language: Python
  framework: None
  orm: None
  test_framework: pytest
```

</examples>

Write your reasoning in <thought>...</thought> then output the YAML in a ```yaml block.

