Metadata-Version: 2.4
Name: githeri
Version: 0.2.0
Summary: Contract Engine for AI Coding Agents. Deterministic specs, test plans, and drift detection.
Author-email: Karakana Labs <dev@karakanalabs.com>
License-Expression: Apache-2.0
Project-URL: Homepage, https://githeri.com
Project-URL: Documentation, https://githeri.com/docs
Project-URL: Repository, https://github.com/karakana-labs/githeri
Project-URL: Issues, https://github.com/karakana-labs/githeri/issues
Keywords: ai,agents,spec,fastapi,testing,openapi,contract-testing
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Software Development :: Build Tools
Classifier: Topic :: Software Development :: Testing
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Python: >=3.10
Description-Content-Type: text/markdown
Requires-Dist: typer>=0.12.0
Requires-Dist: rich>=13.0.0
Requires-Dist: requests>=2.28.0
Requires-Dist: pyyaml>=6.0
Requires-Dist: pydantic>=2.0.0
Requires-Dist: python-dotenv>=1.0.0

# githeri

**Spec-Forge**: a pipeline that trains an AI to bridge the gap between a human's natural-language feature request and a machine-executable specification that feeds the COMMAND_RUNWAY skill.

The central thesis: most time in AI-assisted development is lost before the AI and the human agree on what to build. Githeri attacks that by producing validated, structured YAML specs — which then feed into COMMAND_RUNWAY runbooks that an executor agent follows verbatim.

## Pipeline

```text
Human natural-language request
        │
        � ▼
Spec-Forge (LLM + validator)
Generates:
    data/training_data.jsonl   (valid prompt + spec_yaml pairs)
    data/failed_specs.jsonl    (invalid specs, saved for analysis)
        │
        � ▼
Runbook Scorer (standalone, post-generation)
Scores:
    5 categories — Intent, Preconditions, Structure, Testability, Coverage
    Hard gate: missing Inspect/Create/Verify = 0.0
        │
        � ▼
Human Review
"L1-L4 look good. Approve."
        │
        � ▼
COMMAND_RUNWAY Skill
Consumes:
    a validated spec (single-feature YAML)
Produces:
    COMMAND_RUNWAY.md
        • ordered implementation plan
        • exact file paths
        • code modifications
        • test skeletons
        • verification commands (translated from spec local_goals)
        • rollback guidance
        • completion criteria
        │
        � ▼
GRG Executor (with COMMAND_RUNWAY pattern integration)
Consumes:
    a validated spec OR COMMAND_RUNWAY plan JSON
Produces (all under foreign/ directory):
    • implementation source files
    • test files
    • RUNBOOK.md (human-readable execution log with GRG scores)
    • RUNBOOK.json (machine-readable execution data)
    • automatic ruff check --fix on generated code
        │
        � ▼
Completed Feature
Outputs:
    • implementation complete
    • all tests passing
    • OpenAPI updated
    • documentation synchronized
    • human notified
```

## Spec Enrichment (IMPROVE_SPEC)

Every generated spec now includes optional **enrichment fields** that make specs machine-executable:

| Field | Location | Purpose |
|-------|----------|---------|
| `business_rules` | top-level | Invariants & formulas (e.g., "JWT Secret: 256-bit random, rotated quarterly") |
| `test_fixtures` | top-level | Seed data & setup commands (e.g., `.venv/bin/python scripts/seed_admin.py`) |
| `environment` | top-level | Required packages + env vars (e.g., `pyyaml>=6.0`, `JWT_SECRET`) |
| `global_verification` | top-level | Post-execution gate commands (e.g., `pytest tests/`, `bandit -r src/`) |
| `blueprint` | per-goal | **Required for `type: create`** (≥100 chars). Code-level outline: class signatures, route decorators, SQLAlchemy models, business logic steps |
| `acceptance_criteria` | per-goal | List of `{test, steps}` — executable test cases in pseudo-code |
| `type` | per-goal | `create` \| `update` \| `delete` \| `inspect` \| `verify` — drives runbook stage classification |

These fields are validated by `scripts/validator.py` and consumed by downstream generators (plan, runbook, scorer).

## Runbook Scoring System

Every generated spec is scored against runbook-readiness criteria (see `docs/scoring_spec.md`). Scoring is decoupled from generation — specs are saved first, then scored in a separate pass via `make score`.

The scorer (`scripts/runbook_scorer.py`) evaluates five weighted categories:

| Category | Weight | Key Checks |
|----------|--------|------------|
| **Intent & Goals** | 20% | Summary present, goals have descriptions, endpoint tasks have HTTP verification |
| **Preconditions** | 15% | `depends_on` references valid globals/stages, CLI tools declared in context |
| **Command Runway Structure** | 30% | **Hard gate**: must have Inspect (file_exists/read CLI), Create/Modify (build CLI), Verify (HTTP/test CLI). Stage order: Inspect → Create → Verify |
| **Verification Testability** | 25% | Concrete commands, explicit assertions (status/exit_code/content), reproducible URLs |
| **Completion Coverage** | 10% | Prompt-mentioned status codes, tests, OpenAPI updates reflected in spec |

**Hard gate**: If any of the three runway stages (Inspect, Create/Modify, Verify) is missing, the spec scores **0.0** and is not runbook-ready.

The scorer now **honors explicit `type` on goals** — a goal with `type: create` counts as Create/Modify even if its verification is `file_exists` (executor will generate the file). Similarly `type: inspect` and `type: verify` map directly to stages.

## Quick Start

### 1. Setup

```bash
git clone <repo-url> && cd githeri
python3.11 -m venv .venv && .venv/bin/pip install -r requirements.txt

# For local generation (Ollama):
ollama pull specgen:latest
ollama pull qwen2.5-coder:7b-instruct   # base model (fallback + code execution)

# For cloud generation (NVIDIA NIM - host any model like minimaxai/minimax-m3):
export NVIDIA_API_KEY=your_nvidia_api_key

# For LM Studio local execution (GRG pipeline):
# In LM Studio: enable "OpenAI Compatible Server" on port 1234
# Load qwen2.5-coder-14b-instruct-uncensored model

# For training: install on a GPU machine
pip install "unsloth[colab-new] @ git+https://github.com/unslothai/unsloth.git"
```

### 2. Generate specs

```bash
# Generate a validated spec from a fresh NL prompt (Ollama, specgen:latest)
make spec PROMPT="Add a POST /register endpoint that accepts email and password"

# Use NVIDIA NIM (default model: minimaxai/minimax-m3, base URL: https://integrate.api.nvidia.com/v1)
make spec PROMPT="Add a PATCH /users/{id}/settings endpoint" PROVIDER=nvidia API_KEY=$NVIDIA_API_KEY

# Use any OpenAI-compatible endpoint
make spec PROMPT="..." PROVIDER=openai-compat BASE_URL=https://api.fireworks.ai/inference/v1 API_KEY=... MODEL=accounts/nvidia/models/nemotron-3-ultra

# Generate N specs from random seed prompts (475 prompts across 21 categories)
make generate N=10
make generate N=100

# Generate ALL seed prompts in sequential order (full corpus generation)
make generate N=all
make generate-all

# Generate 10 random specs (explicit alias, good for short test runs)
make generate N=random
make generate-random

# Validate all specs in the corpus
make validate

# Score all specs against runbook criteria (standalone, post-generation)
make score
make score-failed          # score invalid specs in data/failed_specs.jsonl

# Generate COMMAND_RUNWAY plan from a validated spec (two-step: prompt → LLM)
make plan-from-spec SPEC=data/training_data.jsonl#8
# Default: PLAN_MODEL=specgen:latest, PLAN_PROVIDER=ollama, OUTPUT_DIR=./out
# Output: ./out/PLAN.md — use PLAN_MODEL=qwen2.5-coder:7b-instruct for stronger planning
```

### 3. Convert + fine-tune + export

```bash
# Convert training data to chat format (filters by runbook score)
make convert-chat              # default MIN_SCORE=0.75
make convert-chat MIN_SCORE=0.9  # only high-quality specs

# GRG Agent code training (generates verified code with real GRG scores)
make generate-code          # cycles through all 475 SEED_PROMPTS
make convert-code-chat      # converts passing examples to chat format

# LoRA fine-tuning (requires Unsloth + GPU with 8GB VRAM)
make train                     # outputs models/qwen2.5-coder-7b-specforge/

# Merge adapter + export GGUF for Ollama
make merge                       # outputs models/qwen2.5-coder-7b-specforge-gguf/

# Evaluate fine-tuned vs base model on held-out prompts
make eval-model                  # saves data/eval_results.json

# Register fine-tuned model in Ollama
ollama create specgen -f models/specgen/Modelfile
```

### 4. Upload to HuggingFace Hub

Model weights are stored in git (not LFS). The upload sends standard model files to HF Hub.

```bash
# Set up auth
echo "HF_TOKEN=hf_your_token_here" > .env

# Upload
make upload-hf REPO=githeri/specgen
make upload-hf REPO=githeri/specgen PRIVATE=1  # private repo
```

### 5. Install skill for agent use

The `skills/spec-forge/` directory is a self-contained Spec-Forge skill. Install it into your Hermes skills directory to use it from any session.

```bash
# Install to ~/.hermes/skills/ (default Hermes skills directory)
make install-skill             # copies skills/spec-forge/ to ~/.hermes/skills/spec-forge/

# Install to a custom path (e.g. new server, project-local skills dir)
make install-skill-to-server HERMES_SKILLS_DIR=/path/to/hermes/skills

# Remove
make uninstall-skill
```

After install, load the skill in any Hermes session:

```bash
skill_view(name='spec-forge')
```

The skill provides: `make spec PROMPT="..."` (fresh NL → validated spec), `make spec-and-plan PROMPT="..."` (spec + plan prompt), and the bundled validator (`scripts/validator.py`) + plan assembler (`scripts/plan_from_spec.py`).

### 6. Plan from existing specs

Generate a COMMAND_RUNWAY plan from any validated spec in the corpus (two-step: extract + LLM plan generation):

```bash
# Default: specgen:latest, ollama, ./out/PLAN.md
make plan-from-spec SPEC=data/training_data.jsonl#8

# Stronger planning model (if available locally)
PLAN_MODEL=qwen2.5-coder:7b-instruct make plan-from-spec SPEC=data/training_data.jsonl#8
```

### 7. Interactive coding agent (`scripts/githeri.py`)

The project root ships a real interactive coding agent (like Codex/Claude Code). It has a persistent conversation, file operations, command execution, and model/provider switching inside the app.

```bash
# Interactive mode — start a conversation
.venv/bin/python scripts/githeri.py

# In-session commands:
#   /model <name>         - Change model (e.g. /model qwen2.5-coder:7b-instruct)
#   /provider <name>      - Change provider (e.g. /provider ollama)
#   /clear                - Clear conversation
#   /help                 - Show help
#   /quit, /exit          - Exit
```

```bash
# Non-interactive: generate spec from NL prompt via specgen, then execute it
# Only two knobs: --prompt and exec model via env vars. Docker + spec model hardcoded.
GITHERI_MODEL=qwen2.5-coder:7b-instruct \
GITHERI_PROVIDER=ollama \
.venv/bin/python scripts/githeri.py \
  --prompt "Add a PATCH endpoint to update user displayName and bio"

# Execute an existing spec.yaml (skip spec generation)
.venv/bin/python scripts/githeri.py --spec output/my-feature/spec.yaml

# Switch exec model — same prompt, different execution model
GITHERI_MODEL=qwen2.5-coder-14b-instruct-uncensored \
GITHERI_PROVIDER=lmstudio \
.venv/bin/python scripts/githeri.py --prompt "Add a /health endpoint"

# Overrides (rare):
#   --no-docker        - Disable Docker (not recommended)
#   --workdir <path>   - Custom workdir (default: output/sandbox)
#   --docker-image <img> - Custom Docker image (default: githeri-sandbox)
#   --server           - Start server container after agent loop
#   --test-cmd "<cmd>" - Run test command inside container via docker exec
```

**Defaults (hardcoded, never change between tests):**
- `--docker: True` — every command runs in throwaway `githeri-sandbox` container
- `--workdir: output/sandbox` — all file ops jailed here
- `--spec-model: specgen:latest` — hardcoded (fine-tuned for spec generation)
- `--spec-provider: ollama` — hardcoded
- Exec model/provider: `GITHERI_MODEL` / `GITHERI_PROVIDER` env vars only

**Pipeline flow:** 1) specgen:latest (Ollama) → NL prompt → spec.yaml; 2) exec model reads spec, creates files, installs deps, runs verifications — all inside Docker; 3) optional server container; 4) optional test command via docker exec.

See [docs/SPECGEN_HARNESS.md](docs/SPECGEN_HARNESS.md) for full architecture, troubleshooting, and output structure.

## Provider Configuration

All generation targets accept provider overrides:

```bash
# NVIDIA NIM (default: minimaxai/minimax-m3 at integrate.api.nvidia.com/v1)
make spec PROMPT="..." PROVIDER=nvidia API_KEY=$NVIDIA_API_KEY

# OpenAI GPT-4o
make spec PROMPT="..." PROVIDER=openai API_KEY=$OPENAI_API_KEY MODEL=gpt-4o

# Anthropic Claude
make spec PROMPT="..." PROVIDER=anthropic API_KEY=$ANTHROPIC_API_KEY MODEL=claude-3-5-sonnet-20241022

# LM Studio (local, OpenAI-compatible on port 1234)
make spec PROMPT="..." PROVIDER=lmstudio MODEL="qwen2.5-coder-14b-instruct-uncensored" BASE_URL="http://localhost:1234/v1"

# Any OpenAI-compatible endpoint (Fireworks, Together, vLLM, self-hosted NIM)
make spec PROMPT="..." PROVIDER=openai-compat BASE_URL=https://api.fireworks.ai/inference/v1 API_KEY=... MODEL=accounts/nvidia/models/nemotron-3-ultra

# Override sampling
make generate N=5 PROVIDER=nvidia TEMPERATURE=0.3 MAX_TOKENS=4096

# NVIDIA NIM models with longer cold start: use bigger TIMEOUT
make generate N=1 PROVIDER=nvidia TIMEOUT=600
```

Environment variable fallbacks:
- `PROVIDER` → `ollama` (default)
- `MODEL` → provider-specific default
- `API_KEY` → `OPENAI_API_KEY`, `NVIDIA_API_KEY`, `ANTHROPIC_API_KEY`
- `BASE_URL` → `OPENAI_BASE_URL`, `NVIDIA_BASE_URL` (default `https://integrate.api.nvidia.com/v1`)
- `TEMPERATURE` → `0.2`
- `MAX_TOKENS` → `2048`

## Autonomous Execution System

The autonomous execution system takes a natural language prompt and produces a working implementation without human intervention in the loop. It uses the specgen pipeline to generate a validated spec, then uses the Command Runway skills to generate a plan and runbook, and finally executes the runbook using the GRG executor (with self-healing capabilities via GRG quality gates).

For detailed instructions, see [docs/SPRINTS_AUTONOMOUS.md](docs/SPRINTS_AUTONOMOUS.md).

### Quick Reference

```bash
# Basic local execution
.venv/bin/python scripts/spexecutor.py --prompt "Add a POST /notifications endpoint"

# Docker-isolated execution
.venv/bin/python scripts/spexecutor.py --prompt "Add user authentication" --docker

# Using Hermes as the execution backend
.venv/bin/python scripts/spexecutor.py --prompt "Implement rate limiting" --executor hermes

# Makefile: generate spec and plan
make spec-and-plan PROMPT="Add file upload endpoint"

# Makefile: full autonomous cycle (spec -> plan -> runbook -> execute -> report)
make autonomous-cycle SPEC=specs/test-endpoint.yaml

# GRG isolated execution (recommended for clean workspace)
make grg-full PROMPT="Add a POST /webhook endpoint that validates signature" PROVIDER=ollama
# Outputs: foreign/src/, foreign/tests/, foreign/RUNBOOK.md, foreign/RUNBOOK.json
```

## Recent Enhancements

### 2026-08-29 (This Session)

#### COMMAND_RUNWAY Plan Generation from Existing Specs (`make plan-from-spec`)

Added `make plan-from-spec SPEC=data/training_data.jsonl#<index>` — generates a full
COMMAND_RUNWAY plan from any validated spec in the corpus, in two steps:

1. **Extract** — load spec_yaml from JSONL by index (0-based) via `scripts/plan_from_spec.py`
2. **Generate** — build plan prompt + call LLM (`scripts/plan_from_spec_file.py` → `generate_plan()`)

Default: `PLAN_MODEL=specgen:latest`, `PLAN_PROVIDER=ollama`, `OUTPUT_DIR=./out`.
Output: `./out/PLAN.md` — full COMMAND_RUNWAY document with Feature, Target Environment,
and ordered Execution Stages (Objective, Verification, Completion Condition per stage).

Override for stronger plans: `PLAN_MODEL=qwen2.5-coder:7b-instruct PLAN_PROVIDER=ollama`.

Tested: 168-line plan with 3 execution stages generated for session-notes spec
(4176 chars, specgen:latest, ~60s). Requires `specgen:latest` (or any Ollama model)
to be pulled locally.

This unblocks the Sprint 9 workflow (prompt synthesis → bulk spec generation →
bulk plan generation → consolidated COMMAND_RUNWAY document) by providing the
per-spec plan generation step as a reusable Makefile target.

#### Scorer Fix: Accept `verification.type: content` as `file_exists` Alias

Both `qwen2.5-coder:7b-instruct` and `specgen:latest` emit `verification.type: "content"`
with `path` + `expect.content_contains` — structurally identical to `file_exists` but
not recognized by the scorer's canonical vocab. Fixed `scripts/runbook_scorer.py`:

- Added `"content"` to `REQUIRED_EXPECT_KEYS` (same keys as `file_exists`)
- `_check_inspect()` now treats `content` as inspect (read-only check)
- `_check_verify()` treats `content` with content_contains as verify evidence
- Verification testability section handles `content` same as `file_exists`

Before: avg score 0.104, 1/9 above 0.75.
After: avg score 0.972, 9/9 above 0.75.

#### Sprint 9 — Prompt Synthesis & Decomposition Pipeline (Planning)

Created `docs/SPRINTS.md` Sprint 9 (9a-9d) — four sub-sprints to accept raw unfocused
user text end-to-end:

- **9a** `scripts/prompt_synthesizer.py` — raw text → decomposed clean prompts (JSONL)
- **9b** `scripts/bulk_generate.py` — batch spec generation from prompt file
- **9c** `scripts/bulk_plan.py` — loop over specs → plans → consolidated doc
- **9d** `make text-to-plan TEXT="..."` — single target: synthesize → generate → score → plan

Independent of Sprints 5-8. Default models: specgen:latest for generation + planning.

#### Sprints 5, 6, 7 Marked Complete

- **Sprint 5** (Model provider switch) — already built: run_pipeline.py supports 7 providers
- **Sprint 6** (Skill bundling/install) — `install-skill` target body added to Makefile
- **Sprint 7** (Fine-tuning pipeline) — train*.py, merge*.py, eval*.py all exist; Makefile targets wired

`scripts/githeri.py` now supports a full end-to-end pipeline from natural language to tested code. **Spec generation is always handled by `specgen:latest` on Ollama** (finetuned for structured YAML output). **All execution and tests run inside a Docker sandbox by default** — the container is throwaway, so the host `.venv`/`.git` are never polluted across multiple test runs.

```bash
# Minimal invocation — only the prompt and exec model vary between tests.
# Everything else (docker, workdir, spec model/provider) is defaulted.
GITHERI_MODEL=qwen2.5-coder:7b-instruct \
GITHERI_PROVIDER=ollama \
.venv/bin/python scripts/githeri.py \
  --prompt "Add a PATCH endpoint to update user displayName and bio"
```

**Defaults (set once, never touched between tests):**

- --docker: True — every run_command runs in throwaway githeri-sandbox container
- --workdir: output/sandbox — all file ops jailed here; created under repo root
- --spec-model: specgen:latest — **hardcoded** (finetuned for spec generation only)
- --spec-provider: ollama — **hardcoded**
- Exec model/provider: env vars GITHERI_MODEL / GITHERI_PROVIDER — only things you change between tests
- --docker-image: githeri-sandbox

**Only two knobs per test:** --prompt (or --spec) and the exec model via env vars. No flags for docker, workdir, or spec model — they're locked in.

```bash
# Switch exec model — same prompt, same defaults, different model
GITHERI_MODEL=qwen2.5-coder:7b-instruct \
GITHERI_PROVIDER=ollama \
.venv/bin/python scripts/githeri.py --prompt "Add a /health endpoint"

# Different provider for execution (specgen stays ollama/specgen)
GITHERI_MODEL=qwen2.5-14b-instruct-latest:2 \
GITHERI_PROVIDER=lmstudio \
.venv/bin/python scripts/githeri.py --prompt "Add a /health endpoint"

# Execute an existing spec.yaml instead of generating one
.venv/bin/python scripts/githeri.py --spec output/my-feature/spec.yaml
```

**Overrides (rare):**

```bash
# Disable docker (e.g. when Docker daemon is down) — not recommended for tests
.venv/bin/python scripts/githeri.py --prompt "..." --no-docker

# Custom workdir
.venv/bin/python scripts/githeri.py --prompt "..." --workdir output/custom-sandbox

# Custom docker image
.venv/bin/python scripts/githeri.py --prompt "..." --docker-image my-sandbox
```

**Pipeline flow:**
1. Specgen LLM (specgen:latest / Ollama) converts NL prompt → spec.yaml (with entrypoint, packages, goals)
2. Execution model (set via env vars) reads spec, creates files in output/sandbox, installs deps, runs verifications — all inside Docker
3. Server container starts from entrypoint.start_command (if --server)
4. Test command runs inside container via docker exec (if --test-cmd)

**Environment isolation (critical):**
- Docker is the default and recommended mode — every `run_command` (including `pip install`) runs in a throwaway container. The host `.venv` is never touched by spec-generated dependencies.
- `--no-docker` disables this isolation. When used, pip installs and code execution happen in the host environment (the project `.venv`), which can pollute it with spec-generated packages (Flask, FastAPI, etc.). Use `--no-docker` only when you understand this risk.
- If you must run without Docker, create a separate virtualenv for the workdir:
  ```bash
  python -m venv /tmp/githeri-isolated
  source /tmp/githeri-isolated/bin/activate
  pip install -r requirements.txt   # only githeri's own deps
  GITHERI_MODEL=... GITHERI_PROVIDER=... \
    scripts/githeri.py --prompt "..." --no-docker --workdir /tmp/githeri-workdir
  ```

**Model separation:** Spec generation uses specgen:latest (trained for structured YAML output, hardcoded). Execution uses whatever you set via GITHERI_MODEL / GITHERI_PROVIDER. Override via env vars only — the --spec-model / --spec-provider flags are suppressed (backwards-compatible but ignored).

**Server lifecycle:** After the agent loop completes, githeri parses entrypoint.start_command from the generated spec, starts a named Docker container (docker run -d), and runs --test-cmd inside it. Container stays running for manual inspection: docker exec <name> bash.

See [docs/SPECGEN_HARNESS.md](docs/SPECGEN_HARNESS.md) for full architecture, troubleshooting, and output structure.

### 2026-08-15 (This Session)

#### GRG Agent Skill — Hermes Native Integration
The GRG agent skill is now installed as a native Hermes skill (`~/.hermes/skills/autonomous-ai-agents/grg_agent/`) with full multi-provider support:

1. **Skill Installation** — Copied from `skills/grg_agent/` and installed via editable pip install
2. **Lightweight GRG Dependency** — Uses local `grg-0.1.0-py3-none-any.whl` wheel (no karakana dependency) providing `AlphaMomentumTracker` and `compute_structural_alpha`
3. **Hermes Proxy Support** — Skill accepts `provider` argument: `ollama` | `hermes` | `auto` — uses Hermes's configured providers (NVIDIA Nemotron, Nous Portal, xAI Grok, etc.) instead of local models
4. **Direct GRG Agent Execution** — `grg:execute` command now calls `self.agent.solve()` directly instead of legacy `run_pipeline.py` subprocess
5. **Make Target Integration** — `scripts/grg_make_spec.py` updated to use the skill with provider argument: `make grg-spec PROMPT="..." PROVIDER=ollama|hermes|auto`
6. **Project Virtual Environment** — Runs in project's own `.venv/` (not external karakana venv)

**Key Benefit**: You can now use cloud models via Hermes proxy (`hermes proxy start`) instead of relying on locally installed Ollama models. The skill routes through Hermes's provider config which supports NVIDIA Nemotron, Nous Portal, xAI Grok, and any OpenAI-compatible endpoint.

#### LM Studio Local Model — First End-to-End Working Pipeline
**LM Studio** with **`qwen2.5-coder-14b-instruct-uncensored`** is the first local model to complete the full GRG pipeline end-to-end:

```bash
# Full pipeline from NL prompt → validated spec → plan → execution → runbook
make grg-full PROMPT="Implement a FastAPI POST /api/health-check endpoint..." \
    PROVIDER=lmstudio MODEL="qwen2.5-coder-14b-instruct-uncensored" BASE_URL="http://localhost:1234/v1"
```

**What works:**
- **GRG Agent solving** — Generates implementation code with GRG quality gates (composite scoring, diversity control, convergence detection)
- **Code verification** — Execution-based verification (syntax + runtime) in isolated temp files
- **Multi-strategy generation** — Standard, decompose, test_first, refine strategies with adaptive temperature
- **Health-check endpoint example** — Generated FastAPI code with SQLAlchemy DB check + Redis cache check, returns 200/503
- **All artifacts isolated in `foreign/`** — Clean workspace separation

**Configuration for LM Studio:**
```bash
# In LM Studio: enable "OpenAI Compatible Server" on port 1234
# Load qwen2.5-coder-14b-instruct-uncensored model
# Then run:
make grg-spec PROMPT="..." PROVIDER=lmstudio MODEL="qwen2.5-coder-14b-instruct-uncensored" BASE_URL="http://localhost:1234/v1"
```

**Why this model works:**
- 14B parameter coder model fine-tuned for code generation
- Uncensored variant removes alignment filters that can interfere with code structure
- OpenAI-compatible API in LM Studio works with GRG's `OllamaClient` (custom `api_key` support)
- Sufficient context window for spec + plan generation tasks
- Produces valid imports, proper error handling, and correct HTTP status codes

### 2026-07-30 (Latest)

#### THIRD_IMPROVE_SPEC — 7 Pipeline Fixes
After analyzing a 10-spec batch run, identified 7 recurring failure patterns and fixed all of them:

1. **Minimal structural skeleton in SYSTEM_PROMPT** — eliminated top-level field confusion
2. **task_id validator check** — rejects L11/G18-style IDs, requires descriptive slug like `jwt-auth-login`
3. **Verification types & expect keys table** — explicit vocabulary reduces hallucinated keys
4. **Placeholder ban strengthened** — concrete examples in prompt (`JWT_SECRET: "test-secret-..."`, `DATABASE_URL: "postgresql://user:***@..."`) + validator hint
5. **Helpful error hints** — "These fields belong inside a goal under `local_goals`, not at the spec root"
6. **depends_on validator hints** — rejects L/G refs, redirects to task_ids/stage names
7. **Structural acceptance_criteria template** — local goal field, not top-level

All 90 tests passing.

#### Generation Mode Aliases
Added `make generate N=random` and `make generate N=all` flags plus `generate-random` / `generate-all` aliases for short tests and full corpus runs. 475 seed prompts across 21 categories available.

#### Model Default Reverted
`OLLAMA_MODEL` reverted to `qwen2.5-coder:7b-instruct` (was regressed to `qwen3.5-4b-128k:latest`). Better structure compliance after prompt fixes.

### 2026-07-30 (Earlier)

#### Spec Enrichment (IMPROVE_SPEC)
Added five top-level enrichment fields (`business_rules`, `test_fixtures`, `environment`, `global_verification`) and three goal-level fields (`blueprint`, `acceptance_criteria`, `type`). All validated and scored.

#### Multi-Provider LLM Support
`run_pipeline.py` now supports Ollama, OpenAI, Anthropic, NVIDIA (via Together AI), and any OpenAI-compatible endpoint. Configured via `--provider` CLI arg or Makefile variables.

#### Runbook Scorer Stage Detection
Scorer now honors explicit `type: create|inspect|verify` on goals, fixing false "missing stage" penalties for `file_exists` verification on CREATE goals.

#### Validator Hardening
- Guarded against `expect` being a string instead of dict (prevents `AttributeError: 'str' object has no attribute 'get'`)
- Near-duplicate detection now handles malformed `expect` blocks defensively
- All 90 tests pass

## Model Compatibility Matrix

| Model | Context | Tools | Speed (M1 16GB) | Spec Gen | Recommended Use |
|-------|---------|-------|-----------------|----------|-----------------|
| **specgen:latest** | 128K | Yes | Fast | **Best first-attempt pass rate (fine-tuned)** | **Default local (Ollama) — spec generation** |
|| qwen2.5-coder:7b-instruct | 32K | Yes | Fast | Works (best structure compliance) | Fallback / base model for training |
|| qwen2.5-coder-14b-instruct-uncensored | 32K | Yes | Medium | **Works end-to-end (GRG pipeline)** | **LM Studio local** |
|| qwen3.5-4b-128k | 128K | Yes | Fast | Works | Larger context fallback |
|| qwen3.5-9b-code:128k | 128K | Yes | Slow | Excellent | Higher quality if time allows |
|| deepseek-r1:7b | 128K | TBD | Fast | YAML syntax errors | Not recommended for spec gen |
|| **Nemotron 3 Ultra** | 128K | Yes | Fast | Excellent | **Cloud via NVIDIA NIM (`--provider nvidia`)** |

**Current recommendation**: 
- **Spec generation (default)** → `specgen:latest` on Ollama (fine-tuned for structured YAML output, best first-attempt pass rate)
- **Bulk corpus** → `specgen:latest` on Ollama (fine-tuned, best first-attempt pass rate)
- **Specific features (GRG pipeline)** → `qwen2.5-coder-14b-instruct-uncensored` on LM Studio with `make grg-full` (100% success via execution verification)
- **Cloud** → `minimaxai/minimax-m3` via NVIDIA NIM (`--provider nvidia`)

## Training Data Generation Strategies

Based on empirical testing (2026-08-16), here are the reliable approaches:

### 1. Bulk Corpus Generation (Recommended for Training Data)
```bash
# Fast, decent success rate, fine-tuned for this pipeline
make generate N=100 PROVIDER=ollama MODEL="specgen:latest" BASE_URL="http://localhost:11434" TEMPERATURE=0.2 MAX_TOKENS=4096
```
- **Success rate**: ~70%+ (fine-tuned specgen:latest)
- **Speed**: ~50-100s/spec
- **Best for**: Generating large training corpora quickly

### 2. High-Quality Individual Specs (GRG Pipeline)
```bash
# Best for specific features - multi-strategy + execution verification
make grg-full PROMPT="Your feature" PROVIDER=lmstudio MODEL="qwen2.5-coder-14b-instruct-uncensored" BASE_URL="http://localhost:1234/v1"
```
- **Success rate**: ~100% (execution-verified)
- **Speed**: ~260s/spec
- **Best for**: Critical features requiring guaranteed working specs

### 3. GRG Agent Code Training (NEW — Real GRG Scores + Verified Code)
```bash
# Generates verified code with real GRG composite scores (not dummy 0.5)
# Uses qwen2.5-coder:7b-instruct on Ollama for real logprobs
make generate-code          # cycles through all 475 SEED_PROMPTS
make convert-code-chat      # converts passing examples to chat format
```
- **Success rate**: Variable (depends on prompt complexity)
- **Speed**: ~30-60s/prompt (optimized: 1 strategy, 1 candidate)
- **Output**: `data/training_data_code.jsonl` + `data/training_data_code_chat.jsonl`
- **Key difference**: Produces **executable code** with **real GRG composite scores** (0.47-0.49) because the model provides logprobs
- **Verification**: Syntax check + execution test (import + basic run)
- **Best for**: Fine-tuning code generation models with GRG quality signals

**Configuration** (in `scripts/generate_code_training_fast.py`):
```python
skill = create_skill(config={
    'llm_provider': 'ollama',
    'ollama_base_url': 'http://127.0.0.1:11434/v1',
    'ollama_default_model': 'qwen2.5-coder:7b-instruct',
    'max_iterations': 2,           # Must be >=2 for convergence check
    'temperature': 0.3,
    'top_p': 0.9,
    'max_tokens': 1024,
    'candidates_per_strategy': 1,  # Speed optimization
    'max_strategies': 1,           # Speed optimization
})
```

### 4. Hybrid Workflow (Best of Both)
```bash
# 1. Generate bulk corpus with specgen
make generate N=100 PROVIDER=ollama MODEL="specgen:latest" BASE_URL="http://localhost:11434"

# 2. Score and filter high-quality specs
make score MIN_SCORE=0.75

# 3. Re-generate failed critical specs with GRG pipeline
make grg-full PROMPT="..." PROVIDER=lmstudio MODEL="qwen2.5-coder-14b-instruct-uncensored" BASE_URL="http://localhost:1234/v1"

# 4. Generate verified code for fine-tuning
make generate-code
make convert-code-chat
```

### Observed Failure Patterns (specgen:latest)
The remaining failures (if any) are primarily:
1. **Near-duplicate HTTP verifications** - multiple goals hitting same endpoint with same method
2. **Missing `acceptance_criteria`** for CREATE goals
3. **Placeholder values** in headers (e.g., `Authorization: *** ***`)
4. **YAML block mapping errors** - CLI verification indentation issues

## For More Information

- [docs/SPRINTS_AUTONOMOUS.md](docs/SPRINTS_AUTONOMOUS.md) — sprint breakdown, model experiments, decisions, next steps
- [docs/scoring_spec.md](docs/scoring_spec.md) — runbook scoring specification
- [docs/IMPROVE_SPEC.md](docs/IMPROVE_SPEC.md) — spec enrichment field specification (v1)
- [docs/SECOND_IMPROVE_SPEC.md](docs/SECOND_IMPROVE_SPEC.md) — enrichment field enforcement (v2)
- [docs/THIRD_IMPROVE_SPEC.txt](docs/THIRD_IMPROVE_SPEC.txt) — 7 pipeline failure patterns + fixes (v3)
- [MODEL_CARD.md](MODEL_CARD.md) — model card (uploaded to HF Hub)
- [skills/spec-forge/SKILL.md](skills/spec-forge/SKILL.md) — Spec-Forge skill reference
- [skills/command-runway-pattern/SKILL.md](skills/command-runway-pattern/SKILL.md) — Command Runway pattern skill
- [skills/grg_agent/SKILL.md](skills/grg_agent/SKILL.md) — GRG Agent skill reference
