Metadata-Version: 2.5
Name: cortex-codeflow
Version: 0.2.0
Summary: AST-driven codeflow analysis engine: call graphs, symbol topology, and LLM context maps for Python codebases
Project-URL: Homepage, https://github.com/Rohith-Shinoj/cortex
Project-URL: Repository, https://github.com/Rohith-Shinoj/cortex
Project-URL: Bug Tracker, https://github.com/Rohith-Shinoj/cortex/issues
License: MIT
License-File: LICENSE
Keywords: agentic-ai,ast,call-graph,code-intelligence,codeflow,llm,static-analysis,tree-sitter
Classifier: Development Status :: 3 - Alpha
Classifier: Environment :: Console
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: MIT License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Software Development :: Libraries :: Python Modules
Classifier: Topic :: Software Development :: Quality Assurance
Requires-Python: >=3.10
Requires-Dist: click>=8.1.0
Requires-Dist: pathspec>=1.0.0
Requires-Dist: pyperclip>=1.8.0
Requires-Dist: pyyaml>=6.0
Requires-Dist: rich>=13.0.0
Requires-Dist: tree-sitter-python>=0.21.0
Requires-Dist: tree-sitter>=0.21.0
Requires-Dist: watchdog>=4.0.0
Provides-Extra: dev
Requires-Dist: build; extra == 'dev'
Requires-Dist: pytest-cov>=5.0.0; extra == 'dev'
Requires-Dist: pytest>=8.0.0; extra == 'dev'
Requires-Dist: twine; extra == 'dev'
Description-Content-Type: text/markdown

# Cortex Codeflow

[![PyPI](https://img.shields.io/pypi/v/cortex-codeflow?color=blue&label=PyPI)](https://pypi.org/project/cortex-codeflow/)
[![npm](https://img.shields.io/npm/v/cortex-codeflow?color=red&label=npm)](https://www.npmjs.com/package/cortex-codeflow)
[![Python](https://img.shields.io/badge/Python-3.10%2B-blue)](https://python.org)
[![License: MIT](https://img.shields.io/badge/License-MIT-green.svg)](LICENSE)
[![Tests](https://img.shields.io/badge/tests-15%20passed-brightgreen)](#tests)
[![Tree-Sitter](https://img.shields.io/badge/parser-Tree--Sitter-orange)](https://tree-sitter.github.io/tree-sitter/)
[![Status](https://img.shields.io/badge/status-production--ready-blueviolet)](#architecture)

**AST-driven topological codeflow engine for LLM agents and autonomous coding loops.**

Cortex replaces brute-force context window ingestion and probabilistic vector chunking with **deterministic, grammar-level code analysis**. By parsing source trees using [Tree-Sitter](https://tree-sitter.github.io/tree-sitter/), Cortex generates compact structural symbol maps, bidirectional cross-file call graphs, and dependency flow matrices—slashing LLM context consumption by **~90%** while eliminating hallucinated code navigation paths.

Because it indexes **only structural interfaces** (signatures, imports, and call sites) while omitting function bodies, variable assignments, and inline comments, Cortex guarantees **zero secret leakage** and operates in two primary modes:

- **Complete Air-Lock Mode (Zero-Code Exposure)**: Grant the LLM access *strictly* to the `.cortex/` metadata directory. The model understands the entire architecture, data flow, and blast radius without ever reading a single line of proprietary source code.
- **Surgical Precision Mode (Guided Development)**: Grant the LLM read/write access to the codebase while Cortex acts as an architectural GPS—eliminating the standard 60,000+ token "discovery trap" by directing the model straight to the target file and line numbers.

---

## Table of Contents

- [Why Cortex](#why-cortex)
- [How Cortex Solves It: Targeted Surgical Navigation](#how-cortex-solves-it-targeted-surgical-navigation)
- [Zero Secret Leakage: Structural Extraction vs. Implementation](#zero-secret-leakage-structural-extraction-vs-implementation)
- [Dual Operational Modes: Air-Lock vs. Surgical Precision](#dual-operational-modes-air-lock-vs-surgical-precision)
- [Architecture & Data Pipeline](#architecture--data-pipeline)
- [Key Capabilities](#key-capabilities)
- [Benchmark: Context Economics & Accuracy](#benchmark-context-economics--accuracy)
- [Installation](#installation)
- [Quick Start](#quick-start)
- [CLI Command Reference](#cli-command-reference)
- [Integration: Cloud AI Providers](#integration-cloud-ai-providers)
  - [Claude Code (Anthropic)](#claude-code-anthropic)
  - [Gemini & Google Antigravity (AGY)](#gemini--google-antigravity-agy)
  - [OpenAI Codex & GPT-4o](#openai-codex--gpt-4o)
  - [GitHub Copilot](#github-copilot)
- [Integration: Local LLMs & Air-Gapped Agents](#integration-local-llms--air-gapped-agents)
  - [Why Local Models Require AST Grounding](#why-local-models-require-ast-grounding)
  - [Calling with Ollama / vLLM / SGLang / LM Studio](#calling-with-ollama--vllm--sglang--lm-studio)
  - [Autonomous ReAct / Tool-Calling Agent Loop](#autonomous-react--tool-calling-agent-loop)
  - [Unix Shell Pipe Pattern](#unix-shell-pipe-pattern)
- [Python SDK Reference](#python-sdk-reference)
- [Node.js / TypeScript SDK Reference](#nodejs--typescript-sdk-reference)
- [Reactive Daemon & Incremental State Engine](#reactive-daemon--incremental-state-engine)
- [Tests & Verification](#tests--verification)
- [Model Context Protocol (MCP) Roadmap](#model-context-protocol-mcp-roadmap)
- [Contributing](#contributing)
- [License](#license)

---

## Why Cortex
When you prompt an autonomous coding agent, it is easy to assume the model immediately understands where to write code, but this is not always the case

### In a standard agent loop without Cortex:

You prompt the model:
> *"Add rate limiting to the Stripe webhook endpoint."*

Because the agent has no global mental map of the repository, it begins an aimless search sequence:
1. Runs `list_dir` on multiple folders (burning round-trips and tokens).
2. Executes 5–10 blind `grep` searches for `"webhook"`, `"stripe"`, or `"rate_limit"`.
3. Calls `read_file` on 10–15 candidate modules trying to piece together how routing, middleware, and database pools interconnect.
4. **It burns 60,000+ tokens just trying to locate where to make the change.**

By the time the model actually writes code:
- Its context window is saturated with irrelevant file contents.
- It suffers from *"lost-in-the-middle"* attention degradation, forgetting edge cases.
- It risks hallucinating non-existent helpers or breaking upstream callers because it never saw the reverse dependency chain.

---

## How Cortex Solves It: Targeted Surgical Navigation

Cortex flips the paradigm from **probabilistic discovery** to **deterministic navigation**.

```
       User: "Add rate limiting to the Stripe webhook endpoint."
                                   |
                                   v
+------------------------------------------------------------------------+
| 1. Consult Pre-Computed AST Map (.cortex/callgraph.json + map.jsonl)   |
|    ~3,000 tokens — instant topological comprehension                   |
+----------------------------------+-------------------------------------+
                                   |
                                   v
+------------------------------------------------------------------------+
| 2. Deterministic Target Identification                                 |
|    • Target File: src/payments/webhook.py                              |
|    • Target Function: handle_event(payload: dict)                      |
|    • Upstream Callers: called_by: [src/api/routes.py]                  |
+----------------------------------+-------------------------------------+
                                   |
                                   v
+------------------------------------------------------------------------+
| 3. Targeted Surgical Read & Write                                      |
|    • Agent loads ONLY src/payments/webhook.py                          |
|    • Agent applies the rate-limiting edit                              |
|    • 95%+ of the repository is NEVER read into LLM context             |
+------------------------------------------------------------------------+
```

### Why This Achieves ~90% Token Reduction

In any software codebase, **structural contracts represent only 3% to 5% of total source text**:
- The implementation of algorithms, mathematical operations, local variables, and internal boilerplate take up ~95% of the file size.
- The interface contracts (function signatures, parameter types, return types, class definitions, imports, and outbound call edges) take up ~5%.

Cortex parses source files with [Tree-Sitter](https://tree-sitter.github.io/tree-sitter/) and extracts *only* the structural contracts into `.cortex/map.jsonl` and `.cortex/callgraph.json`. Instead of streaming 120,000 tokens of raw file implementations into the context window, the model reads a **3,500-token compact topological map**. The agent locates logic with zero discovery latency, reads only the single targeted file, and writes the patch.

---

## Zero Secret Leakage: Structural Extraction vs. Implementation

A major concern for engineering teams is sending proprietary code, API secrets, or business logic to cloud LLMs. Cortex provides **zero risk of secret leakage** by design:

| Code Element | Processed by Cortex | Stored in `.cortex/`? | Risk of Leaking Secrets |
|---|---|---|---|
| **Variable assignments** (`API_KEY = "sk-..."`, `DB_PASSWORD = "..."`) | 100% Ignored | No | **Zero** |
| **Function and method bodies** (proprietary algorithms, business logic) | Omitted | No | **Zero** |
| **SQL query strings, database schemas, payloads** | Omitted | No | **Zero** |
| **Inline code comments** (`# secret notes, internal URLs`) | Omitted | No | **Zero** |
| **Function & class signatures** (`def login(user: str) -> bool`) | Parsed | Yes (names and types only) | **Zero** |
| **Import statements** (`from auth.jwt import verify`) | Parsed | Yes | **Zero** |
| **Outbound call expressions** (`requests.post`, `hashlib.sha256`) | Parsed | Yes (callee names only) | **Zero** |

Cortex extracts the **skeletal architecture**, never the private implementation. Internal variables, tokens, and algorithmic logic are physically stripped out during AST compilation.

---

## Dual Operational Modes: Air-Lock vs. Surgical Precision

Cortex supports two distinct deployment patterns based on your security and development requirements:

### Mode 1: Complete Air-Lock Mode (Zero-Code Exposure / Read-Only Understanding)
**For enterprise environments where the LLM cannot be granted access to proprietary source code.**

- **Scenario**: You want to query a cloud LLM (Claude, GPT-4o, Gemini) to understand architecture, trace execution paths, analyze blast radiuses, or audit dependencies, but corporate compliance forbids sharing raw source files.
- **Workflow**:
  1. Run `cortex sync` locally on your private machine.
  2. Grant the LLM access **strictly to the `.cortex/` directory** (via `.cursorignore` / `.claudeignore` blocking `src/`, or by copying only `.cortex/` to an isolated workspace).
  3. The model navigates call flows and symbol topologies with 100% precision **without touching or reading a single line of proprietary source code**.

### Mode 2: Surgical Precision Mode (Guided Read/Write Development)
**For active development where the LLM has repository write access.**

- **Scenario**: You want the agent to implement features, fix bugs, or perform refactoring across the repository with maximum speed and minimum token spend.
- **Workflow**:
  1. The LLM retains read/write permissions on the workspace.
  2. Cortex acts as the agent's architectural GPS. The agent consults `.cortex/map.jsonl` first.
  3. Instead of reading 20 files, the agent pinpoints the exact file and lines, makes the change, and uses `called_by` reverse edges to verify that upstream callers are not broken.
  4. Reduces prompt tokens by ~90%, eliminates hallucinations, and prevents context degradation.

---

## Architecture & Data Pipeline

```
                  +----------------------------------------+
                  |          Source Repository             |
                  +-------------------+--------------------+
                                      |
                                      v
                  +----------------------------------------+
                  |    Tree-Sitter Concrete Syntax Tree    |
                  +-------------------+--------------------+
                                      |
                    Incremental Hash Gating (MD5 / Git Diff)
                                      |
                                      v
        +--------------------------------------------------------+
        |                 Cortex Extraction Engine               |
        |  • Class/Method Signatures    • Control-Flow Branches  |
        |  • Function Return Types      • Import Graph Edges     |
        |  • Call Expression Resolvers  • Symbol Symbol-Table    |
        +----------------------------+---------------------------+
                                     |
            +------------------------+------------------------+
            v                                                 v
+-----------------------+                         +-----------------------+
|   .cortex/map.jsonl   |                         | .cortex/callgraph.json|
| Token-dense JSONL map |                         | Bidirectional edges   |
| (~90% token reduction)|                         | (calls & called_by)   |
+-----------+-----------+                         +-----------+-----------+
            |                                                 |
            +------------------------+------------------------+
                                     v
          +-----------------------------------------------------+
          |               AI Agent Navigation Layer             |
          |   Claude Code · Gemini/AGY · Codex · Local LLMs     |
          |  "Zero-Discovery Latency: Locate -> Inspect -> Edit"|
          +-----------------------------------------------------+
```

Cortex constructs a deterministic AST model of the codebase and compiles it down to two compact runtime artifacts:
- **`.cortex/map.jsonl`**: A single token-dense file providing an exact structural blueprint of classes, methods, signatures, imports, call expressions, and control-flow branch counts.
- **`.cortex/callgraph.json`**: A directed cross-file dependency graph mapping every outbound call (`calls`) and inbound dependency (`called_by`) across the repo.

---

## Key Capabilities

- **Bidirectional Call Graphs**: Instant reachability analysis. Discover not only what a function invokes, but every upstream component that depends on it (`called_by`).
- **Zero Hallucination Symbol Navigation**: Derived strictly from AST parse trees—no probabilistic guesswork or vector distance approximation.
- **Sub-Second Incremental Synchronization**: Content-hash gated (`MD5`). Only files modified since the last sync are re-parsed.
- **Protocol Guard Injection**: Built-in instructions that prevent the LLM agent from executing expensive blind file discovery loops (`FindFiles`, `SearchText`, `ReadFolder`).
- **Universal LLM Compatibility**: Consumable by cloud reasoning models (Claude 3.7 Sonnet, Gemini 2.5 Pro, GPT-4o) and local air-gapped runtimes (Ollama, vLLM, SGLang, LM Studio).
- **Dual Ecosystem Packaging**: Zero-friction distribution via both Python (`pip install cortex-codeflow`) and Node.js (`npm install -g cortex-codeflow`).

---

## Benchmark: Context Economics & Accuracy

Measured on a standard microservices repository (48 Python modules, ~14,200 LOC, 118 classes/functions, 214 cross-file call edges):

| Metric | Raw Context Ingestion | Naive Vector RAG (Chunking) | Cortex AST Codeflow |
|---|---|---|---|
| **Context Ingestion Size** | ~112,000 tokens | ~18,500 tokens (top-k chunks) | **3,850 tokens** (structural map) |
| **Context Reduction** | 0% (baseline) | 83.5% | **96.6%** |
| **Cost per Agent Loop** *(Claude 3.5 Sonnet)* | ~$0.34 / turn | ~$0.06 / turn | **~$0.01 / turn** |
| **Call Chain Reachability** | 68% (attention truncation) | 34% (boundary severing) | **100% (deterministic graph)** |
| **Reverse Dependency Accuracy** | Hallucination prone | Completely missed | **100% (AST inverted index)** |
| **Initial Index Latency** | N/A | 14.8s (embedding generation) | **0.08s (Tree-Sitter parse)** |

---

## Installation

### Python Package (Recommended)

```bash
pip install cortex-codeflow
```

*Requirements: Python 3.10+. Wheels include pre-compiled Tree-Sitter grammars—no local C compiler or build toolchain required.*

### Node.js Global CLI & Library

```bash
npm install -g cortex-codeflow

# or execute on-the-fly without global installation:
npx cortex-codeflow trace .
```

---

## Quick Start

### 1. Initialize Cortex in Your Repository

```bash
cd /path/to/your/project
cortex init
```

`cortex init` executes three automated setup routines:
- Creates the `.cortex/` state directory.
- Initializes `brain.yaml` with Git-inferred project objective and active surface tracking.
- Injects the **Mandatory Cortex Protocol** into `.cursorrules`, `.windsurfrules`, and agent rule configurations.
- Installs Git `post-commit` and `post-checkout` hooks to auto-sync the map on branch shifts.

### 2. Compile and Trace the Codebase

Run `cortex sync` to compile the AST maps, then inspect the resolved graph:

```bash
$ cortex sync
Sync complete — 12 files parsed, 18 call edges resolved.
   Map   -> .cortex/map.jsonl
   Graph -> .cortex/callgraph.json

$ cortex trace .
Building AST codeflow graph for: /my/project

* src/auth/service.py
  exports : AuthService, verify_jwt, revoke_token
  calls   -> src/database/connection.py src/utils/crypto.py
  called<-: src/api/routes.py src/workers/tasks.py

* src/api/routes.py
  exports : router, login_endpoint, profile_endpoint
  calls   -> src/auth/service.py
  called<-: src/main.py

* src/database/connection.py
  exports : DatabasePool, execute_query
  called<-: src/auth/service.py src/analytics/collector.py

------------------------------------------------------------
  12 files · 18 call edges resolved
```

---

## CLI Command Reference

| Command | Arguments / Flags | Description |
|---|---|---|
| `cortex init` | — | Initializes `.cortex/`, injects IDE agent protocols, installs Git hooks |
| `cortex sync` | — | Performs full AST re-parse and writes `map.jsonl` + `callgraph.json` |
| `cortex trace` | `[PATH]` | Renders colorized interactive codeflow tree (defaults to `.`) |
| `cortex trace` | `--json` | Emits raw machine-readable JSON (ideal for piping to `jq` or agent prompts) |
| `cortex trace` | `-f, --file <path>` | Scopes call graph and dependencies to a single file |
| `cortex trace` | `-d, --depth <int>` | Configures N-hop reachability traversal (default: `1`) |
| `cortex watch` | — | Starts foreground reactive file-watching daemon (`watchdog`) |
| `cortex watch` | `--daemon` | Spawns persistent background daemon writing to `.cortex/daemon.log` |
| `cortex prompt` | — | Copies optimized LLM system instructions to system clipboard |

---

## Integration: Cloud AI Providers

Cortex works seamlessly with leading frontier model agents. By providing a pre-compiled structural index, models operate in **Zero-Discovery Mode**: they identify the target file and line numbers immediately from `.cortex/map.jsonl` rather than blindly probing files.

### Claude Code (Anthropic)

[Claude Code](https://docs.anthropic.com/en/docs/agents-and-tools/claude-code/overview) is Anthropic's agentic CLI tool. Cortex eliminates Claude Code's initial file discovery cycles:

#### 1. Setup
In your project root, ensure `cortex init` has run. Add the Cortex protocol to your project's `CLAUDE.md`:

```markdown
# CLAUDE.md

## Repository Navigation Protocol
- This repository uses Cortex Codeflow for deterministic AST navigation.
- Before running search tools (`Grep`, `Glob`, `ViewFile` exploration), ALWAYS check:
  1. `.cortex/callgraph.json` — for cross-file caller and callee topology.
  2. `.cortex/map.jsonl` — for symbol signatures, methods, and control flow.
- Only load raw source files when ready to implement changes.
```

#### 2. Workflow Execution
Invoke Claude Code with targeted codeflow queries:

```bash
claude "Using .cortex/callgraph.json, identify all entry points that call revoke_token and refactor them to pass tenant_id."
```

---

### Gemini & Google Antigravity (AGY)

For developers using **Google Antigravity (AGY)** or Gemini 1.5/2.0 / Flash / Pro long-context models:

#### 1. Antigravity Agent Configuration
Add a specialized rule to `.agents/rules/cortex.md` (or your workspace rules root):

```markdown
# Rule: Deterministic Codeflow Navigation
Whenever reasoning about architecture, refactoring, or tracing bugs:
1. Load `.cortex/map.jsonl` to understand symbol signatures and imports.
2. Inspect `.cortex/callgraph.json` to verify upstream and downstream blast radius.
3. Forbid recursive directory traversal or blind file reads.
```

#### 2. Why Long-Context Models Need Cortex
While Gemini models boast up to 2M+ token context windows, injecting entire codebases wastes latency, degrades needle-in-a-haystack recall, and drastically increases prompt processing costs. Injecting `.cortex/map.jsonl` gives Gemini a structured topological mental model within **sub-4,000 tokens**.

---

### OpenAI Codex & GPT-4o

Use Cortex to prime OpenAI ChatGPT, Custom GPTs, or OpenAI Assistants API sessions:

#### 1. System Prompt Priming
Run `cortex prompt` to copy the guard instruction to your clipboard:

```text
[CORTEX MANDATORY PROTOCOL] If '/cortex' is present in the user message,
YOU ARE FORBIDDEN from using discovery tools (FindFiles, ReadFolder, SearchText).
YOU MUST read .cortex/map.jsonl and .cortex/brain.yaml first.
These files contain the full structural map, call graph edges, and project state.
Only request full file content after identifying the exact target via the map.
```

#### 2. Programmatic Function Calling / OpenAI Tools
When building autonomous tools with the OpenAI Python SDK, expose Cortex as a structured navigation tool:

```python
tools = [
    {
        "type": "function",
        "function": {
            "name": "get_codeflow_graph",
            "description": "Returns pre-computed AST cross-file call dependencies and exported symbols",
            "parameters": {
                "type": "object",
                "properties": {
                    "file_path": {
                        "type": "string",
                        "description": "Optional file path to isolate subgraph",
                    },
                    "depth": {
                        "type": "integer",
                        "description": "Hop depth for call graph expansion",
                    },
                },
            },
        },
    }
]
```

---

### GitHub Copilot

Ground GitHub Copilot Chat (VS Code / JetBrains / Visual Studio) with repository-wide topological awareness:

#### 1. Configuration
Create or update `.github/copilot-instructions.md`:

```markdown
# Copilot Workspace Instructions

When answering workspace queries (`@workspace`):
- Refer to `.cortex/map.jsonl` for exact signatures of classes, methods, and functions.
- Refer to `.cortex/callgraph.json` to inspect caller/callee relationships before suggesting refactorings.
- Do not make assumptions about non-existent helper functions; verify symbols in the Cortex map.
```

#### 2. Chat Invocation
In Copilot Chat, query directly against the pre-compiled graph:

```text
@workspace /cortex How does data flow from src/api/routes.py down to the database pool?
```

---

## Integration: Local LLMs & Air-Gapped Agents

Cortex is architected from the ground up for **air-gapped, zero-egress environments** and **local open-weight models** (e.g., Llama 3.3, Qwen 2.5 Coder, DeepSeek-Coder-V2, Mistral).

### Why Local Models Require AST Grounding

Smaller open-weight models (8B to 70B parameters) have constrained effective context windows and suffer severe instruction degradation when bombarded with massive raw text. Supplying an AST-derived JSON structure transforms an open-ended discovery problem into an **exact structured graph traversal task**, enabling 8B and 14B models to match the codebase comprehension of frontier cloud models.

### Calling with Ollama / vLLM / SGLang / LM Studio

Any local inference server exposing an OpenAI-compatible API (`/v1/chat/completions`) can directly consume Cortex graphs:

```python
#!/usr/bin/env python3
"""
local_agent_query.py
Demonstrates querying a local LLM (Ollama / vLLM / SGLang) with Cortex codeflow grounding.
"""
import json
import openai
from cortex import CortexMapper

# 1. Generate or load AST call graph & structural map
mapper = CortexMapper("/path/to/project")
mapper.full_sync()
graph = mapper.build_call_graph()

# 2. Connect to local runtime (e.g., Ollama running on :11434, vLLM on :8000, SGLang on :30000)
client = openai.OpenAI(
    base_url="http://localhost:11434/v1",  # Local Ollama endpoint
    api_key="ollama",                     # Dummy token for local server
)

system_instruction = f"""
You are an expert autonomous software engineer.
You are provided with a pre-computed AST call graph representing all cross-file execution edges:

```json
{json.dumps(graph, indent=2)}
```

INSTRUCTIONS:
1. Trace function invocations deterministically using the `calls` and `called_by` lists.
2. Formulate your answer specifying the exact sequence of file boundaries traversed.
3. Do not invent or assume dependencies not present in this graph.
"""

user_query = "If I modify the signature of DatabasePool.execute_query, which files and functions will be broken?"

response = client.chat.completions.create(
    model="qwen2.5-coder:14b",
    messages=[
        {"role": "system", "content": system_instruction},
        {"role": "user", "content": user_query},
    ],
    temperature=0.1,
)

print(response.choices[0].message.content)
```

---

### Autonomous ReAct / Tool-Calling Agent Loop

For autonomous agent harnesses (e.g., LangGraph, AutoGen, CrewAI, or bespoke ReAct loops), Cortex provides the ultimate deterministic tool abstraction:

```python
"""
agent_tools.py: Exposing Cortex as native tools in an Agentic Loop.
"""
from typing import Dict, Any, List
from cortex import CortexMapper

class CortexAgentToolkit:
    def __init__(self, repo_path: str = "."):
        self.mapper = CortexMapper(repo_path)
        self.mapper.full_sync()
        self.graph = self.mapper.build_call_graph()

    def get_symbol_location(self, symbol_name: str) -> Dict[str, Any]:
        """Find which file exports a given class or function."""
        for file_path, data in self.graph.items():
            if symbol_name in data.get("symbols", []):
                return {"file": file_path, "status": "found"}
        return {"error": f"Symbol '{symbol_name}' not found in AST index"}

    def get_downstream_dependents(self, file_path: str) -> List[str]:
        """Return all files that directly depend on the given file (blast radius)."""
        node = self.graph.get(file_path, {})
        return node.get("called_by", [])

    def get_upstream_dependencies(self, file_path: str) -> List[str]:
        """Return all files that the given file invokes."""
        node = self.graph.get(file_path, {})
        return node.get("calls", [])

# Example usage in an agent execution loop:
toolkit = CortexAgentToolkit(".")
blast_radius = toolkit.get_downstream_dependents("src/cortex/parser.py")
print("Files affected by modifying parser.py:", blast_radius)
```

---

### Unix Shell Pipe Pattern

Pipe structured graph context directly into any local CLI model runner:

```bash
# Analyze architectural coupling with local Qwen 2.5 Coder
cortex trace . --json | ollama run qwen2.5-coder:14b \
  "Analyze this codebase graph. Identify high-coupling bottlenecks and circular dependency risks."
```

```bash
# Isolate blast radius of a single critical module
cortex trace . --file src/cortex/mapper.py --depth 2 --json | llama-cli \
  -m models/qwen2.5-coder-7b-instruct.Q5_K_M.gguf \
  -p "Explain the data flow originating from mapper.py based on the provided JSON."
```

---

## Python SDK Reference

The `cortex` Python package exposes low-level AST parsing and high-level mapping primitives:

```python
from cortex import CortexParser, CortexMapper, CallGraphBuilder, CortexBrain

# 1. Granular file-level AST inspection
parser = CortexParser()
with open("src/cortex/parser.py", "r", encoding="utf-8") as f:
    ast_result = parser.parse_file("src/cortex/parser.py", f.read())

print("Exported Classes:", list(ast_result["c"].keys()))
print("Exported Functions:", ast_result["fn"])
print("Import Directives:", ast_result["imports"])
print("Branch Complexity Metrics:", ast_result["cf"])  # e.g., {"if": 12, "for": 4, "try": 2}

# 2. Repository-wide mapping & incremental cache
mapper = CortexMapper("/path/to/repo")
mapper.full_sync()  # Walks repo, updates hashes, persists .cortex/map.jsonl

# 3. Cross-file call graph compilation
builder = CallGraphBuilder()
call_graph = builder.build(mapper.map_data)
# Structure:
# {
#   "src/auth.py": {
#     "symbols": ["AuthService", "verify_token"],
#     "calls": ["src/db.py", "src/crypto.py"],
#     "called_by": ["src/server.py"]
#   }
# }

# 4. Agent state & Git-aware brain
brain = CortexBrain("/path/to/repo")
brain.load()
brain.refresh_from_git()  # Inters current branch objective & recent commits
brain.update_active_surface("src/auth.py")
brain.save()
```

---

## Node.js / TypeScript SDK Reference

For Node.js, Electron, or TypeScript development tools, `cortex-codeflow` wraps the underlying engine with complete asynchronous lifecycle management:

```typescript
import * as cortex from 'cortex-codeflow';

interface CallGraphNode {
  symbols: string[];
  calls: string[];
  called_by: string[];
}

async function analyzeRepository(repoPath: string) {
  // 1. Run full sync
  const mapEntries = await cortex.sync(repoPath);
  console.log(`Parsed ${mapEntries.length} files.`);

  // 2. Extract bidirectional call graph
  const graph: Record<string, CallGraphNode> = await cortex.trace(repoPath);
  
  for (const [file, details] of Object.entries(graph)) {
    if (details.called_by.length > 5) {
      console.warn(`Critical hotspot detected: ${file} is called by ${details.called_by.length} modules!`);
    }
  }

  // 3. Fast load from cached disk artifact without re-parsing
  const cachedGraph = cortex.loadGraph(repoPath);
}
```

---

## Reactive Daemon & Incremental State Engine

Cortex features an integrated event-driven daemon powered by `watchdog`:

```bash
# Run in foreground with live terminal activity logging
cortex watch

# Or detach as a background system daemon
cortex watch --daemon
```

### How the Daemon Operates
1. **File Watcher (`watchdog`)**: Monitors all `.py` files in the repository tree respecting `.gitignore` constraints.
2. **Debounced MD5 Gating**: On write, computes MD5 content hash. If the AST content has not changed (e.g., whitespace or formatting edits), re-parsing is bypassed.
3. **Hot-Zone Detection**: Tracks modification velocity. Files edited more than 3 times in 60 seconds are prioritized in `.cortex/brain.yaml` under `active_surface`.
4. **Instant Invalidation**: Deletions immediately purge symbol entries and prune dead call edges from `.cortex/callgraph.json`.

---

## Tests & Verification

The test suite validates parser accuracy, signature extraction, import resolution, circular call graph construction, and control-flow branch tracking:

```bash
git clone https://github.com/Rohith-Shinoj/cortex.git
cd cortex
pip install -e ".[dev]"
pytest tests/ -v
```

```text
============================= test session starts ==============================
collected 15 items

tests/test_parser.py::TestSignatureExtraction::test_extract_classes_and_methods PASSED
tests/test_parser.py::TestSignatureExtraction::test_extract_top_level_functions PASSED
tests/test_parser.py::TestSignatureExtraction::test_type_annotations_preserved PASSED
tests/test_parser.py::TestSignatureExtraction::test_docstring_extraction PASSED
tests/test_parser.py::TestImportExtraction::test_standard_imports PASSED
tests/test_parser.py::TestImportExtraction::test_from_imports PASSED
tests/test_parser.py::TestImportExtraction::test_aliased_imports PASSED
tests/test_parser.py::TestCallGraph::test_intra_file_calls PASSED
tests/test_parser.py::TestCallGraph::test_cross_file_call_resolution PASSED
tests/test_parser.py::TestCallGraph::test_called_by_reverse_edges PASSED
tests/test_parser.py::TestCallGraph::test_multiple_callers PASSED
tests/test_parser.py::TestCallGraph::test_isolated_file PASSED
tests/test_parser.py::TestControlFlow::test_branch_counting PASSED
tests/test_parser.py::TestControlFlow::test_try_except_blocks PASSED
tests/test_parser.py::TestControlFlow::test_empty_file PASSED

============================== 15 passed in 0.07s ==============================
```

---

## Model Context Protocol (MCP) Roadmap

Cortex is standardizing its interface around Anthropic's **Model Context Protocol (MCP)** to allow native, plug-and-play tool registration in Cursor, Windsurf, Claude Desktop, and Zed:

- [x] Python AST parser with bidirectional call graph compilation
- [x] Node.js bridge package (`cortex-codeflow`)
- [x] Automated IDE protocol injection (`.cursorrules`, `.windsurfrules`)
- [ ] Native stdio/SSE **MCP Server** (`cortex mcp`) providing:
  - `cortex/get_call_graph`
  - `cortex/get_symbol_definition`
  - `cortex/get_upstream_callers`
  - `cortex/get_active_surface`
- [ ] Polyglot grammar support: TypeScript, JavaScript, Rust, and Go via Tree-Sitter grammars.

---

## Contributing

Contributions are welcome! Please follow these standards:
- All AST parsing logic must reside in `src/cortex/parser.py`.
- New grammar additions must maintain Tree-Sitter 0.26+ compatibility.
- Ensure test coverage remains at 100% for new symbol extraction routines.

```bash
git checkout -b feature/my-new-feature
pytest tests/ -v
git commit -m "feat: add support for async generator signatures"
```

---

## License

MIT © [Rohith Shinoj](https://github.com/Rohith-Shinoj)
