Metadata-Version: 2.4
Name: coding-os
Version: 0.3.21
Summary: The cognitive operating system that gives AI agents memory, structure, and discipline
Author-email: Kourosh Ebrahimzadeh <info@coding-os.dev>
License-Expression: Apache-2.0
Project-URL: Homepage, https://github.com/kouroshez/coding-os
Project-URL: Repository, https://github.com/kouroshez/coding-os
Project-URL: Issues, https://github.com/kouroshez/coding-os/issues
Keywords: ai,coding-agent,mcp,llm,cli,hexagonal,claude,codex
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: Operating System :: MacOS :: MacOS X
Classifier: Operating System :: POSIX :: Linux
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Software Development :: Code Generators
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: mcp>=1.27.0
Requires-Dist: pydantic>=2.12.5
Requires-Dist: click>=8.1.0
Requires-Dist: pyyaml>=6.0
Requires-Dist: jinja2>=3.1.0
Requires-Dist: jsonschema>=4.17.0
Requires-Dist: fastapi>=0.115.0
Requires-Dist: uvicorn[standard]>=0.30.0
Requires-Dist: sse-starlette>=2.1.0
Requires-Dist: claude-agent-sdk<0.3.0,>=0.2.110
Requires-Dist: anyio>=4.0.0
Requires-Dist: networkx>=3.0
Provides-Extra: rag
Requires-Dist: sentence-transformers>=2.7.0; extra == "rag"
Requires-Dist: numpy>=1.24.0; extra == "rag"
Provides-Extra: graph-os
Requires-Dist: sentence-transformers>=2.7.0; extra == "graph-os"
Requires-Dist: numpy<3.0.0,>=1.24.0; extra == "graph-os"
Requires-Dist: usearch>=2.16; extra == "graph-os"
Requires-Dist: sqlite-vec>=0.1.9; extra == "graph-os"
Requires-Dist: tree-sitter<0.27.0,>=0.25.0; extra == "graph-os"
Requires-Dist: tree-sitter-python<0.26.0,>=0.25.0; extra == "graph-os"
Requires-Dist: tree-sitter-typescript<0.26.0,>=0.23.0; extra == "graph-os"
Requires-Dist: tree-sitter-bash<0.26.0,>=0.25.0; extra == "graph-os"
Requires-Dist: tree-sitter-yaml<0.8.0,>=0.6.1; extra == "graph-os"
Requires-Dist: tree-sitter-go<0.26.0,>=0.23.0; extra == "graph-os"
Requires-Dist: tree-sitter-php<0.26.0,>=0.23.0; extra == "graph-os"
Requires-Dist: pathspec>=0.12.1; extra == "graph-os"
Requires-Dist: tree-sitter-rust<0.26.0,>=0.23.0; extra == "graph-os"
Requires-Dist: tree-sitter-ruby<0.26.0,>=0.23.0; extra == "graph-os"
Requires-Dist: tree-sitter-java<0.26.0,>=0.23.0; extra == "graph-os"
Requires-Dist: tree-sitter-c<0.26.0,>=0.23.0; extra == "graph-os"
Requires-Dist: tree-sitter-cpp<0.26.0,>=0.23.0; extra == "graph-os"
Requires-Dist: tree-sitter-c-sharp<0.26.0,>=0.23.0; extra == "graph-os"
Requires-Dist: tree-sitter-scala<0.27.0,>=0.23.0; extra == "graph-os"
Requires-Dist: tree-sitter-kotlin>=1.0.0; extra == "graph-os"
Requires-Dist: tree-sitter-lua>=0.4.0; extra == "graph-os"
Provides-Extra: board-os
Requires-Dist: aiohttp>=3.9.0; extra == "board-os"
Requires-Dist: watchdog>=4.0.0; extra == "board-os"
Provides-Extra: web
Provides-Extra: claude-sdk
Provides-Extra: codex-sdk
Requires-Dist: openai-codex<0.145.0,>=0.144.4; extra == "codex-sdk"
Dynamic: license-file

# coding-os

[![PyPI](https://img.shields.io/pypi/v/coding-os?color=blue&label=pypi)](https://pypi.org/project/coding-os/)
[![license](https://img.shields.io/badge/license-Apache--2.0-blue.svg)](./LICENSE)
[![python](https://img.shields.io/badge/python-3.10%20%7C%203.11%20%7C%203.12-blue)](./pyproject.toml)
[![CI](https://github.com/kouroshez/coding-os/actions/workflows/ci.yml/badge.svg)](https://github.com/kouroshez/coding-os/actions/workflows/ci.yml)
[![OpenSSF Scorecard](https://api.scorecard.dev/projects/github.com/kouroshez/coding-os/badge)](https://scorecard.dev/viewer/?uri=github.com/kouroshez/coding-os)
[![cli](https://img.shields.io/badge/cli-cos-informational)](./docs/architecture/meta-project.md)

> **Coding OS — the cognitive operating system that gives AI agents memory, structure, and discipline.**
> Teaches AI agents *how to think* (thinking_os) and *how to code*
> (workflow, hooks, skills, rules) — agent-agnostic so the same kernel
> serves Claude Code and OpenAI Codex without rewriting. **Modular by
> design**: take only what you need — even just the knowledge graph
> (`--profile lite --enable-module graph`).
>
> Website: <https://coding-os.dev> · Community: <https://community.coding-os.dev>

![Coding OS — knowledge graph, Scrumban board, and an enforced engineering loop for AI coding agents](docs/assets/hero-card.webp)

<table>
<tr>
<td align="center" width="34%">
<h3>💛 Sponsor</h3>
<a href="https://github.com/sponsors/kouroshez"><b>GitHub Sponsors</b></a><br/>
<a href="https://coding-os.dev/donate">One-off donation</a>
<br/><sub>Keeps a single maintainer shipping</sub>
</td>
<td align="center" width="33%">
<h3>💬 Community</h3>
<a href="https://community.coding-os.dev"><b>Forum</b></a><br/>
<a href="https://github.com/kouroshez/coding-os/discussions">Discussions</a>
<br/><sub>Questions, setups, rough edges</sub>
</td>
<td align="center" width="33%">
<h3>⭐ Star</h3>
<a href="https://github.com/kouroshez/coding-os"><b>Follow the repo</b></a><br/>
<a href="./CHANGELOG.md">Read the changelog</a>
<br/><sub>The cheapest way to help</sub>
</td>
</tr>
</table>

**Free, and staying that way.** Apache-2.0, no paid tier, no telemetry, no
gated features — the kernel you run is the whole thing. Sponsorship buys
maintenance time, not access.

### Right-sized by construction

Loading a whole skill catalogue into every prompt is how agents run out of
context before they run out of task. coding-os narrows it twice.

**At `cos init`** — a project installs only the stacks it declares. A
WordPress project gets the `wordpress` skill and never sees `go-patterns`,
`rails` or `spring-boot`. Alongside them sit the stack-agnostic ones —
`clean-code`, `testing-strategy`, `api-design`, `a11y`, `security-web` —
which apply to any codebase.

**On every turn** — what is on disk is not what is in context. A generated
rules table maps file globs to skills, and the matching skill loads only
when you are about to write a file it governs. Editing a React component
loads `nextjs-react`; nothing else comes with it.

Three real installs, each from a `cos init` run, showing what the declared
stacks actually add on top of the seven stack-agnostic rules everyone gets:

| Preset | Always-on rules | The stack-specific ones | Stack skills on disk |
|---|---:|---|---|
| `wordpress-cms` | 8 | `wordpress-backend` | `wordpress` |
| `mern` | 9 | `nextjs-frontend`, `node-express-backend` | `nextjs-react`, `node-express` |
| `hexagonal-product` | 11 | `fastapi-backend`, `go-backend`, `go-fiber-backend`, `react-native-mobile` | `go-fiber`, `go-patterns`, `python-fastapi`, `react-native-mobile`, `react-native-patterns` |

In the WordPress install, **zero files under `.claude/` mention `go-fiber`,
`rails` or `spring-boot`**, and the polyglot install has zero mentioning
WordPress anywhere. (Two governance contract docs do name other stacks — as
examples of the naming convention, not as instructions.) That is the difference
between a scoped install and a prompt directory that grows with every stack the
tool has ever supported.

**What that actually costs — measured, not asserted.** A scaffold writes 395–461
files; **9 to 12 of them ever enter the prompt**. Running the real `cos init`
for all 21 shipped presets and summing only what is resident:

| Project shape | Always-on tokens | Share of a 200k window |
|---|---:|---:|
| `t3-style` (Next.js) — the leanest | 12,704 | 6.4% |
| `wordpress-cms` | 12,738 | 6.4% |
| `mern` / `pern` (Express + Next.js) | 13,158 | 6.6% |
| `hexagonal-product` (Go + Fiber + FastAPI + React Native) | 13,972 | 7.0% |

```bash
uv run python src/scripts/context_budget.py --all-presets
```

The spread across every profile is **12.7k–14.0k tokens**, because 9,266 of
those tokens are the stack-agnostic rules everyone gets and the per-stack
overlays are only 199–1,171 each. Skills, slash commands, hooks and MCP tool
schemas are **not** in that number — they load on demand.

Two things this figure is not: it is not free (7% of the window is 7% of the
window), and it is not the **15,634 tokens (7.8%)** this meta-repo itself carries
— it pays for four kernel-only `meta-*` rules a consumer never receives
(`uv run python src/scripts/context_budget.py --project .` reproduces it). The
full accounting, the cache economics, and what is still *unmeasured* about
instruction density are in
[context-budget.md](docs/engineering/context-budget.md).

Disable a module and its rules, tools and slash commands leave with it — see
[Modular by design](#modular-by-design--take-only-what-you-need).

---

## Prerequisites

| Tool | Min version | Why | macOS install |
|---|---|---|---|
| Python | 3.10 | CLI, MCP server, extractors | `brew install python@3.12` |
| [uv](https://docs.astral.sh/uv/) | 0.5 | Fast Python installer + tool runner | `curl -LsSf https://astral.sh/uv/install.sh \| sh` |
| Bash | 4 | Hook scripts use 4.x features (macOS ships 3.2) | `brew install bash` |
| Node.js | 20 | **Only** if rebuilding the Hub UI under `src/core/web/ui/` | `brew install node@20` |
| Docker | 24 | **Only** for the Docker quickstart below | `brew install --cask docker` |
| `jq`, `perl` | any | **Optional** — faster hook path. Both degrade to Python, so an image without them is fully enforced, just slower per tool call. | preinstalled on macOS |

Linux: replace `brew install …` with your distro's package manager
(`apt`, `dnf`, `pacman`). Windows: WSL 2 + the same Linux steps.
`cos doctor --bootstrap` checks all of the above, including whether the hook
layer has a JSON/stdin parser at all — without one every gate fails closed.

## Install

Two paths to the same install — pick by preference, not capability. Both end
with the `cos` CLI on your PATH and the Hub reachable at
`http://127.0.0.1:9188`.

### Path A — one command, then click

Preflights prerequisites, installs the CLI, and boots the Hub:

```bash
curl -fsSL https://raw.githubusercontent.com/kouroshez/coding-os/main/install.sh | bash
# …or, from a checkout:  bash install.sh
```

Then open the Hub and press **New project**. The Composer picks a preset (or your
own stack mix), asks one sentence about the project, and scaffolds it — docs,
board, knowledge graph, and agent setup included.
([ADR-0007](docs/architecture/adr/0007-gui-first-install-path.md))

### Path B — native `uv`, scripted

```bash
# 1. Install the cos CLI globally, from PyPI
uv tool install coding-os              # upgrade later: uv tool upgrade coding-os

#    …or from a checkout, if you intend to work ON coding-os itself:
#    git clone https://github.com/kouroshez/coding-os.git && cd coding-os
#    uv tool install --editable .      # upgrade later: git pull (editable is live)

# 2. Verify
cos --version                          # → coding-os, version X.Y.Z
cos doctor --bootstrap                 # preflight: python/bash/git/uv/sed prerequisites
cos doctor                             # full health sweep (must be all-green)

# 3. Spawn a new project, scaffolded with an agent + a stack
#    --agent takes several at once: --agent claude,codex
cos init --agent claude --template django --name my-shop --yes
cd my-shop                             # adapter installer ran for you and wrote
                                       # .claude/, .mcp.json, .coding-os/

# 4. Boot the multi-project Web Hub: graph + board + cognition + search
cos hub start                          # → http://127.0.0.1:9188
```

Open `http://127.0.0.1:9188` in your browser. You will see the
knowledge graph of `my-shop`, the Scrumban board, the cognition
trace timeline, and unified search across all retrieval layers.

**The Hub is optional.** The CLI-only loop is complete on its own: `cd
my-shop`, open your agent (e.g. `claude` — MCP + hooks are already
wired), then `cos daily` / `cos task-create` / `cos task-start` — `cos init`
prints the exact first commands for your project. Every project ships its own
guide at `docs/workflow/workflow-guide.md`. For an existing repo, use
`cos adopt` instead of `init`.

For Codex, swap `--agent claude` for `--agent codex` (or pass
both — `--agent claude,codex`) — everything else is identical. Each agent's
installer is `src/adapters/<agent>/install.sh`; `cos init` runs it
for you and re-runs it on `cos update`.

### Modular by design — take only what you need

coding-os is not all-or-nothing. The **kernel** (session lifecycle + safety) is
always on; everything else is a **subsystem module** you switch on or off:
`docs` · `tasks` (Scrumban) · `graph` · `memory` · `cognition` ·
`observability` · `hub-extras` · `cicd`. Named **profiles** curate the set, so
the agent's MCP tool surface stays as small as you want it.

Wanted just the knowledge graph? That is the entire install:

```bash
cos init --agent claude --name my-app --profile lite --enable-module graph --yes
```

`lite` is kernel-only; `core` adds docs, tasks, graph and the Hub hooks;
`standard` *(the recommended default)* adds memory + observability; `full` is
everything. What each one actually delivers, from four `cos init` runs:

| Profile | Modules off | Always-on rules | Skills on disk | Hooks that self-skip |
|---|---:|---:|---:|---:|
| `lite` | 8 | 7 | 43 | 52 |
| `core` | 4 | 7 | 45 | 26 |
| `standard` | 2 | 8 | 46 | 12 |
| `full` | 0 | 9 | 46 | 0 |

The gate is a derived allowlist (`.coding-os/disabled-hook-scripts`) that a
disabled hook reads when it sources `cos-env.sh` — so a module you turned off
costs nothing at runtime. **The 14 safety-category hooks are never on that list,
at any profile** — secret-blocking, destructive-command and branch guards cannot
be switched off by choosing a leaner install. Verify on your own project:

```bash
cos module list                       # what you ended up with
wc -l .coding-os/disabled-hook-scripts   # what stopped firing
```

Start lean and change your mind later — from the project root,
`cos module enable memory`. `--profile` and `--disable-module` are **unioned**:
they can only remove. `--enable-module` is the escape that keeps one on and
pulls its dependencies with it — passing the same id to both flags is an error,
not an override. `cos init --help` lists the live ids straight from
`src/core/subsystems.yaml`, `cos module list` shows what you ended up with, and
Hub **Config → Modules** — or the Composer's *Advanced* section at create
time — flips any of it later. Full model:
[meta-project.md § subsystem modules](docs/architecture/meta-project.md#the-fourth-axis--subsystem-modules-which-kernel-surfaces-ship).

## Run with Docker (Hub layer; native for projects)

**Architecture split** — adopted because each layer wants a different
deploy shape:

| Layer | Runs where | Why |
|---|---|---|
| **Hub** (web panel: graph · board · cognition · search) | **Docker** (production-shaped) | Reproducible build · isolated runtime · same image dev → CI → prod |
| **Consumer projects** (each project's `.coding-os/`, MCP server, skills, adapters) | **Host (native)** | Agent runtimes (Claude Code / Codex CLI) live on the host filesystem · `cos init` factory writes alongside your source · IDE/editor needs direct paths |

The Hub container **reads** the host's projects via a read-only bind
mount and the host's registry file, so every absolute path stays
valid inside the container — no path translation.

### Quickstart

```bash
docker compose up
# → http://127.0.0.1:9188
```

By default, `docker-compose.yml` bind-mounts `$HOME` read-only at
the same path inside the container so `cos registry scan ~` finds
every `.coding-os/` directory below it. Hub state (SQLite, traces)
lives in the `cos-state` named volume and survives `down` / `up`.

Project auto-discovery, narrowing the mount for production, and manual
`docker run` (no compose): [docs/engineering/hub-architecture.md § Docker
deployment](./docs/engineering/hub-architecture.md#docker-deployment-hub-layer).

## MCP server wire-up (Claude / Codex)

`cos init` writes `.mcp.json` at the project root automatically. If
you ever need to register the MCP server manually (e.g. another tool
that reads MCP configs), this is the shape every adapter installs:

```json
{
  "mcpServers": {
    "coding-os": { "command": "cos", "args": ["server-start"] }
  }
}
```

Verify the wire is live in your agent runtime:

- **Claude Code:** `cos doctor` shows `mcp.coding-os = ok`; the CLI
  exposes `cos_*` tools via `ToolSearch("select:<tool>")`.
- **Codex CLI:** `codex --mcp-list` lists `coding-os`.

If the server isn't found, re-run `bash src/adapters/<agent>/install.sh`
from the project root, then restart the agent.

---

## What it is

`coding-os` is a three-layer composition (DNA → mRNA → phenotype):

```
src/core/  ──►  src/adapters/<agent>/  ──►  src/templates/<stack>/  ──►  consumer project
(DNA)         (mRNA)                       (phenotype)                 (organism)
```

| Layer            | What it owns                                                       |
| ---------------- | ------------------------------------------------------------------ |
| `src/core/`      | MCP server, hooks, rules, skills — **agent-agnostic, stack-agnostic** |
| `src/adapters/`  | Per-agent translation: `.claude/`, `.codex/` rendering             |
| `src/templates/` | Per-stack overlays: 27 stacks, 11 CI-verified — `cos list-stacks` marks each `verified` or `experimental` |
| `src/cli/`       | The `cos` factory CLI that composes the three layers               |

Adding a new stack or a new agent is a pure YAML + Markdown change.
No Python edits required.

## What it does

1. **Complexity Gate** — classifies problems before acting (Cynefin:
   CLEAR / COMPLICATED / COMPLEX / CHAOTIC / CONFUSION).
2. **Cognitive Cycle** — CLASSIFY → ORIENT → PLAN → EXECUTE → VERIFY.
   The kernel rule (`src/core/rules/thinking_os.md`) is always
   active; the deep skill loads only when the gate returns COMPLICATED
   or COMPLEX.
3. **Self-learning memory** — SQLite-backed observations, metrics,
   and learned patterns across sessions (`cos_search`, `cos_learn_*`).
4. **Hook enforcement** — hooks gate writes, edits, prompts,
   sessions, and stops (exact count in `src/core/hooks/registry.yaml`).
   Adapter parity matrix in `docs/engineering/`.
5. **Four-layer retrieval** — agent memory (`cos_search`) · doc RAG
   (`cos_doc_search`) · task graph (`cos_task_*`) · knowledge graph
   (`cos_graph_*`).
6. **Intent enforcement** — when the user uses exhaustive vocabulary
   ("all" / "every" / "completely" / "until done", matched in English
   and one additional language), the Stop hook refuses premature "done"
   until an evidence bundle is recorded.
7. **Upgrade path** — `cos update` keeps every consumer project in
   sync with `coding-os` without touching user content.

## The Web Hub (`http://127.0.0.1:9188`)

One FastAPI + Vite/React singleton serves **every** registered project through
`/api/p/<slug>/*` — `cos hub start` boots it, `cos hub status` reports health.
It is genuinely optional: the CLI loop is complete without it. What it adds is
one surface where the board, the knowledge graph, agent memory, live agent
presence and the hook stream are the *same* state the agent reads, not a
dashboard rendered beside it.

![Hub home listing four registered projects and the agents live in each](docs/assets/hub/hub-home.webp)

Hub home is the entry point: register an existing `.coding-os/` folder, scan a
directory for them, or scaffold a new project from a preset. Each project card
carries the agents attached to it right now — the model each is running, the
complexity gate it recorded, and the role it holds in the current chain.

### Overview — what the agent is doing, in one screen

Agents live, spend against budget, WIP against cap, and the hook stream as it
fires. `block-dangerous-commands 15s ago` is not a log line scraped after the
fact — it is the same enforcement record the gate wrote when it ran.

![Workspace Overview with live agent, cost, WIP counters and the recent hook stream](docs/assets/hub/workspace-overview.webp)

### The board — Scrumban with enforced WIP

Seven columns; the three that gate flow carry a cap — in progress, testing,
emergency. `1 / 1 WIP` is not decoration: `cos task-move` refuses to overfill a
capped column, so the board cannot drift from what is actually being worked on.
Columns holding nothing collapse to a labelled rail and hand their width to the
columns holding work — then expand back into full drop zones the moment you
start dragging a card.

![Scrumban board with three empty columns collapsed to rails and WIP counters on the rest](docs/assets/hub/board-flat.webp)

Group by swimlane instead of flat, and open any card for its outcome contract,
its lifecycle chips, and every commit that touched it with the real diff inline.
The task is the pointer; the code is the evidence.

### Chat — the agent, with the project already loaded

A session started here inherits the project's MCP server, hooks, rules and
board. Pick the model, the reasoning effort, and which semantic role should
answer; the quick actions under the composer are the openings that actually pay
off — *resume where we left off*, *map the subsystems*, *review my current
changes*.

![Chat landing with model, effort and role pickers above six quick actions](docs/assets/hub/chat-streamos.webp)

### Search — one query, four retrieval layers

Memory, docs, tasks and the graph answer the same query in parallel, each in its
own section with its own count. A layer still in flight shows `…`, never a zero —
an unanswered layer and an empty one are different facts, and a summary that
conflates them is worse than one that waits.

![Unified search returning 24 results for one query across docs, tasks and the graph](docs/assets/hub/search-query.webp)

### The graph — ask the codebase structurally

Sigma.js over the extracted code + doc graph. The left spine is containment, the
tabs switch between the blended overview, pure containment, dependencies and
detected communities, and the budget control decides how much of the graph to
draw. The badge is honest about coverage: it names the sample **and** the
whole-graph total, so a capped view can never read as a complete one. This is
the surface behind `cos_graph_*` — the reason an agent answers "who calls this?"
without grepping the tree.

![Knowledge graph canvas showing 800 of 78,128 nodes with view-mode tabs and a depth budget](docs/assets/hub/graph-explorer.webp)

### Memory — the lessons, and how far each one has travelled

Agent memory is not a log. Every lesson carries a confidence and a tier, and it
only becomes durable once separate sessions confirm it — so the page states the
exact rule it enforces (≥70% confidence *and* three confirmations) and how many
lessons have actually cleared it. Filter by type, source or confidence floor,
and run the distillation loop on demand instead of waiting for the nightly one.

![Agent Memory tab showing 102 lessons, the validation rule, and type and source filters](docs/assets/hub/memory.webp)

### Diagnostics — every hook that fired, and why

The log tail is structured, not a text blob: filter by level floor, scope glob,
message substring or a relative window, and follow it live. Alongside it sit the
doctor probes, observability charts, and per-session cognition traces.

![Diagnostics Logs tab with level, scope, substring and window filters over a structured event table](docs/assets/hub/diag-logs.webp)

### Modules — turn subsystems off and the tools go with them

The kernel is always on. Everything else is a switch, and disabling one gates
its MCP tools and self-skips its hooks. `Owns` is the blast radius of that
switch. `Depends on` is enforced, not advisory — `tasks` needs `docs`, so `docs`
cannot be disabled first, and the button that would break it is disabled with
the reason attached rather than failing after the click.

![Config Modules tab showing subsystem toggles, owned artifacts and dependency direction](docs/assets/hub/config-modules.webp)

Also in the Hub: per-project **Git** and supervision settings, and a
**Marketplace** for community skills and stacks (rolling out).
Source: `src/core/web/`; UI: `src/core/web/ui/` (`make ui-dev` for HMR).

## Architecture

```
coding-os/
├── src/                # All importable code (Python src-layout)
│   ├── cli/              # Factory entrypoint (`cos` command)
│   ├── core/             # Agent-agnostic brain (DNA)
│   │   ├── thinking_os/    # MCP server: memory, learning, metrics, cognition
│   │   ├── graph_os/       # Polyglot knowledge graph (SQLite backend)
│   │   ├── board_os/       # Scrumban task system
│   │   ├── web/            # Hub UI + FastAPI backbone
│   │   ├── hooks/          # Hook scripts (SSOT: registry.yaml)
│   │   ├── rules/          # Always-active rules + auto-generated artifacts
│   │   ├── skills/         # Universal skills
│   │   └── scripts/        # Kernel-internal regen tooling
│   ├── adapters/         # Per-agent translation (mRNA, adapter.yaml manifests)
│   │   ├── claude/         # Claude Code adapter
│   │   └── codex/          # OpenAI Codex CLI/Desktop adapter
│   ├── templates/        # Per-stack scaffolds (phenotype, stack.yaml-driven)
│   │   ├── _base/          # Generic base + fragments/
│   │   ├── django/         # Django + DRF + PostgreSQL
│   │   ├── nextjs/         # Next.js + React + TypeScript + Tailwind
│   │   ├── fastapi/        # FastAPI + Pydantic + SQLAlchemy
│   │   ├── go/             # Go stdlib + chi router
│   │   ├── go-fiber/       # Go + Fiber v3
│   │   ├── react-native/   # React Native + Expo
│   │   ├── python/         # Python library / CLI / MCP server
│   │   ├── meta/           # Meta-stack (for coding-os contributors)
│   │   └── …               # 27 stacks (11 CI-verified) — `cos list-stacks`
│   └── scripts/          # Maintenance + regen tooling
├── tests/              # cross-cutting tests
├── docs/               # Governance, engineering, playbooks, architecture
└── .coding-os/         # Per-project runtime state (gitignored)
```

## Command index (highlights · 99 `cos` subcommands total)

```
Project lifecycle    init · adopt · setup · add-adapter · add-stack · update · materialize · eject
Modules              module list · module enable · module disable   (per-project surface control)
Diagnostics          doctor · health · list-stacks · list-adapters · hooks-dir · hooks-log
Hub                  hub start · hub status · hub stop
Board                board · task-create · task-start · task-move · task-done · daily · retro · wip
Cognition            cognition trace · trace-replay · trace-summary
Supervision          supervision show · enable · disable · set   (per-role adapter/model/effort)
Graph                29 graph-* subcommands (build · find · deps · analysis · review);
                     22 mirror a cos_graph_* MCP tool one-for-one, enforced by a parity test
```

Full catalogue with flows: [docs/architecture/meta-project.md](./docs/architecture/meta-project.md).

## Slash commands (25 commands)

The `cos` CLI above is the *factory*. Inside an agent session you also get
**slash commands** — packaged workflows invoked by typing `/`: 11 workflow
commands (`/board`, `/daily`, `/retro`, `/task`, `/classify`, `/compose`,
`/memory-search`, `/verify`, `/review`, `/diagnose`, `/new-project`) and 14
`/role-*` commands (the semantic roles of the cognition chain). They ship in
`.claude/commands/` (and `.codex/commands/`) and are version-controlled, so
every teammate gets them on clone. Day-to-day usage:
[docs/workflow/workflow-guide.md](./docs/workflow/workflow-guide.md).

## MCP tools (`cos_*` family, all `ok / fail` envelope)

One MCP server (launched by `.mcp.json` → `cos server-start`) exposes every
`cos_*` tool across ten families: health, memory (`cos_search`), learning,
metrics, routing, docs (`cos_doc_search`), tasks (`cos_task_*`), graph
(`cos_graph_*`, 22 tools), cognition (`cos_compose_chain`), and retrieval.
Per-tool docs + envelope spec:
[docs/governance/mcp-tool-inventory.md](./docs/governance/mcp-tool-inventory.md).

## The knowledge graph — why it changes the economics

Most "AI coding" tools answer structural questions ("who calls this?",
"what breaks if I rename it?", "where does this data flow?") by *reading
files* until the agent guesses an answer. That burns tokens, slows the
loop, and produces hallucinations the moment a caller lives in a file
the agent didn't open.

coding-os ships a precomputed knowledge graph as the **third retrieval
layer** alongside memory and docs. Every commit refreshes 23 node kinds
(functions, methods, classes, modules, routes, MCP tools, docs,
headings, frontmatter, hooks, rules, skills, tasks, …) and 18 edge
types (`contains`, `calls`, `imports`, `inherits_from`,
`handles_route`, `has_param_type`, `references_doc`, `is_decorated_by`,
`links_to`, …). The agent then asks the graph — `cos_graph_references`,
`cos_graph_impact`, `cos_graph_rename_plan` — and gets a small,
high-confidence JSON envelope back.

### Benchmark — graph envelope vs a competent agent, on public repos

The number to beat is **not** "read every matching file" — no sensible agent
does that. It is what a good agent actually does: grep, then open a bounded
window around the matches in the few highest-hit files. That is the default
baseline, and every figure below is measured against it on public checkouts you
can reproduce.

```bash
uv run --extra graph_os python src/core/graph_os/bench/third_party.py \
    --repo https://github.com/django/django --ref 5.2 --queries 10
```

Median savings over the highest-degree symbols per repo — `--queries 10` on the
public checkouts, 8 on this one (`min` in brackets — the honest worst case):

| Repo | `.py` files | `references` | `impact` (3 hops) | `rename_plan` |
|---|---:|---:|---:|---:|
| psf/requests @ v2.32.5 | 36 | **77.7%** (41.9) | 24.2% (−53.8) | **74.8%** (43.7) |
| fastapi/fastapi @ 0.116.1 | 1,129 | **79.5%** (−3.4) | **−6.8%** (−85.6) | **82.4%** (11.0) |
| django/django @ 5.2 | 2,818 | **76.8%** (50.3) | 70.8% (18.5) | **77.1%** (51.1) |
| this repo | 3,317 | **79.7%** (65.9) | 74.0% (64.7) | **79.7%** (65.7) |

Read it as three findings, including the one that does not flatter us:

1. **"Who calls this?" and "what does a rename touch?" are a consistent ~75–82%
   cheaper**, across repos spanning two orders of magnitude in size. This is the
   robust win and the reason the graph-first rule exists.
2. **A 3-hop blast radius is size-dependent, and on mid-size repos it can cost
   *more* than reading.** `impact` is +71–74% on django and this repo, +24% on
   requests, and **−7% on fastapi**. A wide transitive envelope is not free;
   reach for `depth=3` when the codebase is large enough to make reading worse.
3. **Against bare `grep` output alone on a small repo, the graph loses badly**
   (−169% on requests). If match lines answer the question, they are the right
   tool. The graph earns its keep when you need the *complete* set.

That completeness is the part a token count cannot show. Every envelope carries
`total_count` and its own truncation flags, so the agent knows whether it has
the whole answer — grep never tells you what it missed. The harness enforces the
same discipline on itself: an envelope whose traversal was capped is reported as
incomplete and **never scored as a saving**. (The previous version of this table
did exactly that — it published "508 impacted, 98.3% saved" from a
`walk_truncated` envelope whose real count, at a sufficient budget, is 1,494.)

Method, the other two baselines, and the limits — including that
highest-degree probe selection favours the graph — are in
[third-party-token-bench.md](docs/engineering/third-party-token-bench.md).

### Coverage, budgets, health — the anti-hallucination contract

Every coverage-sensitive tool reports its own incompleteness
(`total_count` · `result_truncated` · `walk_truncated` — never silent),
all 23 node kinds answer end-to-end in 0–23 ms, `cos_graph_doctor`
sweeps stale nodes, and every Write/Edit re-indexes just the touched
file. The full contract — budget knobs, per-kind latency, Hub view
modes, and the probe-then-widen workflow — lives in
[graph_os-queries.md § Coverage, budgets, and benchmarks](./docs/engineering/graph_os-queries.md#coverage-budgets-and-benchmarks-readme-deep-dive).

Deep dive: [docs/engineering/graph_os-queries.md](./docs/engineering/graph_os-queries.md)
· [docs/engineering/graph-hallucination-cures.md](./docs/engineering/graph-hallucination-cures.md)
· [docs/governance/mcp-tool-inventory.md](./docs/governance/mcp-tool-inventory.md).

## Supported agents

| Agent | Hook coverage | Skills | MCP server | Notes |
| --- | --- | --- | --- | --- |
| Claude Code | Full for its native events ✅ | Native skills | ✅ | No native `SessionEnd`. |
| Codex CLI | Full for supported Codex events ✅ | Native agent skills | ✅ | Includes Bash, Read, `apply_patch`, MCP, prompt, compact, subagent, permission, Stop, and SessionEnd hooks. |
| Codex Desktop | Same project hook/config contract as Codex CLI ✅ | Native agent skills | ✅ | Project hooks require trust/review; Hub observability is native, while Hub interactive chat is still Claude-only. |

Parity matrix + reasoning: [docs/engineering/adapter-parity.md](./docs/engineering/adapter-parity.md)
(the 2026-04-25 workflow audit is a historical snapshot predating Codex parity).

## Agent supervision — pick the model per role (opt-in)

Off by default. Turn it on and every role — the 11 in the canonical chain plus
the `distiller`, `onboarder` and `repairer` specialists — can run on its own
adapter, model, and reasoning effort: a cheap model reviewing, an expensive one
architecting. It works with a **single** adapter too, since routing across one
provider's own model tiers is the common case rather than a fallback.

```bash
cos supervision enable
cos supervision set --orchestrator-model claude-sonnet-5     # project default
cos supervision set --role reviewer  --role-model claude-haiku-4-5 --role-effort low
cos supervision set --role architect --role-model claude-opus-4-8  --role-effort xhigh
cos supervision show
```

With more than one adapter installed, a role can cross runtimes entirely —
review on Codex while architecture stays on Claude:

```bash
cos supervision set --role reviewer --role-adapter codex
```

Eligibility is **probed, not declared**: an adapter is offered for a role when
its dispatcher resolves at load time, so an adapter whose CLI or SDK is missing
is never silently routed to. Model catalogs come from `adapter.yaml`, and an
adapter that publishes none (Codex takes a freeform `-m`) accepts the model you
type. A new runtime is therefore configurable the day it is installed.

When a provider reports a rate limit, only the **model pool** that hit it goes
into a persistent cooldown — providers meter each pool separately, so an
exhausted Opus pool must not stop a reviewer running on Haiku. Recovery is one
half-open probe, not a retry storm against a limit that cannot succeed.

Same policy from the Hub (**Config → Settings**), the CLI, or MCP — no Hub
required. Disabled means disabled: no probe, no state write, no tokens.

Full contract: [docs/engineering/agent-supervision.md](./docs/engineering/agent-supervision.md)
· operator guide: [docs/playbooks/agent-supervision-setup.md](./docs/playbooks/agent-supervision-setup.md)

## Configuration

`.coding-os.yaml` at every project root:

```yaml
version: "1.0"
agents: [claude, codex]
templates: [django, nextjs]
state_dir: .coding-os
code_extensions: [py, ts, tsx]
verify:
  backend: "make lint-backend && make test-backend"
  frontend: "cd src/frontend && npm run lint && npm test"
protected_files:
  - "*/migrations/*.py"
```

## Adding a new stack (zero Python changes)

Create `src/templates/<id>/stack.yaml` plus skills, rules, and scaffold
docs — the CLI auto-discovers it (`cos list-stacks`), then
`make manifest-regen && make regen-rules` refreshes the derived
artifacts. The same pattern works for new **adapters**
(`src/adapters/<id>/adapter.yaml` + `install.sh`). Step-by-step:
[docs/playbooks/template-authoring.md](./docs/playbooks/template-authoring.md)
· [docs/playbooks/adapter-authoring.md](./docs/playbooks/adapter-authoring.md).

## Project structure (for contributors)

```bash
make verify-hooks         # shellcheck + bash -n on every hook
make verify               # matrix-targeted tests for what changed
make test-mcp             # MCP self-test (cold start)
make docs-lint            # markdown structure + link integrity
cos health                # cross-project health summary
make manifest-regen       # refresh src/core/scaffold_manifest.json
make regen-rules          # refresh dimension-registry + skill-enforcement
```

CI runs the matrix on every PR. See `.github/workflows/ci.yml`.

## Documentation

| Doc                                                                                 | What's in it                                                |
| ----------------------------------------------------------------------------------- | ----------------------------------------------------------- |
| [AGENTS.md](./AGENTS.md)                                                            | **Agent entry point** — Core Loop, Critical Rules, Verification Matrix |
| [docs/architecture/meta-project.md](./docs/architecture/meta-project.md)            | Hexagonal design, DNA/mRNA/phenotype, propagation matrix    |
| [docs/governance/critical-rules.md](./docs/governance/critical-rules.md)            | 27 critical rules with rationale + repair steps             |
| [docs/governance/mcp-tool-inventory.md](./docs/governance/mcp-tool-inventory.md)    | Per-tool spec + envelope contract                           |
| [docs/governance/agent-workflow.md](./docs/governance/agent-workflow.md)            | Domain routing, task protocol, memory contract              |
| [docs/engineering/graph_os-queries.md](./docs/engineering/graph_os-queries.md)      | When to query the graph vs grep                             |
| [docs/engineering/hub-architecture.md](./docs/engineering/hub-architecture.md)      | Hub: FastAPI ↔ React SPA contract                           |
| [docs/engineering/agent-supervision.md](./docs/engineering/agent-supervision.md)    | Per-role adapter/model routing, capacity breaker, trigger modes |
| [docs/playbooks/](./docs/playbooks/)                                                | Hook authoring · adapter authoring · template authoring · MCP tool authoring |
| [docs/adapters/](./docs/adapters/)                                                  | Claude SDK · Codex CLI integration                          |
| [CONTRIBUTING.md](./CONTRIBUTING.md)                                                | Setup, contribution loop, PR checklist                      |
| [SECURITY.md](./SECURITY.md)                                                        | Vulnerability disclosure policy                             |
| [GOVERNANCE.md](./GOVERNANCE.md)                                                    | Decision model, quality gates, becoming a maintainer        |
| [KNOWN_LIMITATIONS.md](./KNOWN_LIMITATIONS.md)                                      | Honest constraints + the ratchets that shrink them          |
| [docs/governance/stability-contract.md](./docs/governance/stability-contract.md)    | What 1.0 freezes + the deprecation ladder                   |
| [CHANGELOG.md](./CHANGELOG.md)                                                      | Release notes                                               |

## Troubleshooting

| Symptom | Cause | Fix |
|---|---|---|
| `cos: command not found` after `uv tool install` | `~/.local/bin` (or uv's tool dir) not on `PATH` | `uv tool update-shell` then open a new shell |
| `cos doctor` reports `mcp.coding-os = absent` | Adapter installer hasn't run for this project | `bash src/adapters/<agent>/install.sh` from project root, then restart the agent runtime |
| `cos hub start` fails with `Address already in use :9188` | Port 9188 busy (likely an old Hub still running) | `lsof -ti:9188 \| xargs kill` then re-run; or `cos hub start --port 9999` |
| `make verify` complains `bash: declare -A …` | macOS default bash 3.2 doesn't have associative arrays | `brew install bash` (Makefile picks up `/opt/homebrew/bin/bash` automatically) |
| `cos init` fails on `npm ci` step | Node.js missing or below 20 | Install Node ≥20 (`brew install node@20`); only required if your template touches `src/core/web/ui/` |
| Docker build OOM on `npm ci` | Default Docker memory < 4 GB | Docker Desktop → Settings → Resources → bump memory to 4 GB+ |
| `ToolSearch` returns `InputValidationError` for a `cos_*` tool | First-call schema not loaded (Claude defers MCP schemas) | `ToolSearch("select:cos_<name>")` first, then call the tool |
| Codex hook is skipped | Project/hash trust is missing, the hooks feature is disabled, or the event/matcher is unsupported | Run `/hooks`, confirm `[features] hooks = true`, then inspect `cos hooks-list --agent codex` |
| Hub rejects the meta-repo checkout with `sits inside … already a coding-os project` | A stray `.coding-os/` exists higher up (e.g. `~/.coding-os/` from a test run) — fixed 2026-05-23: only **registered** ancestors block | Update + restart Hub: `git pull && cos hub stop && cos hub start`. If still blocking, the ancestor is genuinely registered: `cos registry remove <ancestor-path>` |

Still stuck? Run `cos doctor --verbose` and open a
[discussion](https://github.com/kouroshez/coding-os/discussions)
with the output attached.

## Support / Community

If coding-os saves you time, a star helps others find it. These links also
live in the Hub footer (never inside the new-project Composer).

- ★ Star / follow on GitHub: <https://github.com/kouroshez/coding-os>
- Sponsor the project: <https://github.com/sponsors/kouroshez> · <https://coding-os.dev/donate>
- Questions / ideas: <https://github.com/kouroshez/coding-os/discussions>
- Community forum: <https://community.coding-os.dev>

## License

Apache License 2.0 — see [LICENSE](./LICENSE). Copyright 2026
Kourosh Ebrahimzadeh and coding-os contributors.

Development began in April 2026; the full history is preserved in this
repository. Release automation (release-please) starts at the 0.3.0
baseline (2026-05-20) — see [CHANGELOG.md](./CHANGELOG.md).
