Metadata-Version: 2.4
Name: reachscan
Version: 0.3.2
Summary: Reachability-based capability scanner for AI agent and MCP server codebases.
Author-email: Vinmay Nair <vinmay.nair@gmail.com>
License-Expression: Apache-2.0
Project-URL: Homepage, https://github.com/vinmay/reachscan
Project-URL: Website, https://reachscan.dev
Requires-Python: >=3.11
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: tree-sitter<1,>=0.23
Requires-Dist: tree-sitter-typescript<1,>=0.23
Requires-Dist: tree-sitter-javascript<1,>=0.23
Provides-Extra: dev
Requires-Dist: pytest; extra == "dev"
Requires-Dist: ruff; extra == "dev"
Requires-Dist: pipx; extra == "dev"
Requires-Dist: click; extra == "dev"
Requires-Dist: jsonschema; extra == "dev"
Dynamic: license-file

# reachscan

[![PyPI version](https://img.shields.io/pypi/v/reachscan)](https://pypi.org/project/reachscan/)
[![CI](https://github.com/vinmay/reachscan/actions/workflows/ci.yml/badge.svg)](https://github.com/vinmay/reachscan/actions/workflows/ci.yml)
[![Python 3.11+](https://img.shields.io/badge/python-3.11%2B-blue)](https://www.python.org/downloads/)
[![License](https://img.shields.io/badge/license-Apache%202.0-green)](LICENSE)
[![GitHub stars](https://img.shields.io/github/stars/vinmay/reachscan?style=social)](https://github.com/vinmay/reachscan)

> Static capability analysis for Python and TypeScript/JavaScript AI code.
> Know what it can do before it does it.

## Quick start

| You want to... | Use |
|---|---|
| Scan any agent or MCP server from your terminal | `pipx install reachscan`, then `reachscan <path \| github-url \| pypi:package>` |
| Check your own agent or MCP server on every pull request | The [GitHub Action](https://github.com/marketplace/actions/reachscan): `uses: vinmay/reachscan-action@v1` ([CI integration](#ci-integration)) |
| Vet someone else's MCP server before you install it | The plugin for Claude Code or Codex: ask "vet this MCP server before I add it: &lt;url&gt;" ([setup](#use-it-from-claude-code-or-codex)) |

---

## The problem

You're giving an LLM tools. Tools mean real-world access — files, shell, network, credentials.

Most developers add tools without a clear accounting of what permissions they're actually granting. The agent docs tell you what the tool is *for*. They don't tell you what it *can do*.

`reachscan` is the accounting.

It analyzes Python and TypeScript/JavaScript code and reports the actual capabilities present: what the code can read, write, execute, send, and access. Not what the README says. What the code does.

---

## What it detects

Seven capability classes, built from AST analysis of Python and TypeScript/JavaScript code:

| Capability | What it means |
|---|---|
| `EXECUTE` | Shell commands, subprocess, OS exec APIs |
| `READ` | Local file reads, path traversal |
| `WRITE` | File creation, modification, deletion |
| `SEND` | Outbound HTTP, websockets, raw sockets |
| `SECRETS` | Env vars, credential managers, secret stores |
| `DYNAMIC` | eval, exec, dynamic imports |
| `AUTONOMY` | Background tasks, schedulers, self-directed execution |

Cross-capability risks are also flagged when both capabilities are reachable: READ + SEND (secret leakage), SEND + WRITE (data exfiltration), EXECUTE + SEND (remote control), and EXECUTE + destructive WRITE (destructive agent).

---

## Reachability analysis

Knowing a capability exists in a codebase is useful. Knowing whether the LLM can actually trigger it is what matters.

`reachscan` detects the LLM-facing entry points in your codebase, builds an intra-project call graph, and traces which capabilities are reachable from those entry points. Every finding is tagged with one of five states:

| State | Meaning |
|---|---|
| `reachable` | Confirmed on a call path from an LLM entry point |
| `unreachable` | Exists in the codebase, not on any LLM call path |
| `module_level` | Runs on import — executes when the module loads, not via a function call |
| `unknown` | Inside code that can't be statically resolved (dynamic dispatch, parse failure) |
| `no_entry_points` | No entry points detected — full reachability analysis not possible |

The call graph follows up to 8 hops from each entry point. Call paths are shown in the report so you can see exactly how the LLM reaches a capability.

**Known limitations of the Python call graph.** It resolves:
- plain calls to project functions (`helper()`, including ones imported from another project file);
- `self.method()` calls to methods defined in the same class and file;
- `module.function()` calls.

It doesn't yet resolve method calls on object instances (`client.fetch()`, `self.db.query()`), chained attribute calls (`a.b.c()`), `super().method()`, or inherited methods. A capability reached only through such a call shows as unreachable or unknown, and can't produce an annotation mismatch.

### Entry point detection — Python

`reachscan` recognises LLM-callable functions across all major Python agent frameworks:

| Framework | Detection pattern |
|---|---|
| MCP (Python SDK / FastMCP) | `@mcp.tool()`, `@server.tool()` |
| MCP (lowlevel server API) | `@app.call_tool()`, `@app.list_tools()` |
| Pydantic AI | `@agent.tool`, `@agent.tool_plain` |
| LangChain / CrewAI | `@tool`, `class MyTool(BaseTool)`, `StructuredTool` |
| OpenAI Agents SDK | `@function_tool` |
| Semantic Kernel | `@kernel_function` |
| AutoGen | `@register_for_llm` |
| LlamaIndex | `FunctionTool.from_defaults(fn=...)`, `QueryEngineTool.from_defaults(...)` |
| DSPy | `dspy.Tool(func)` |
| Google ADK | `Agent(tools=[...])` |
| OpenAI Swarm | `Agent(functions=[...])` |
| CAMEL AI | `FunctionTool(func)` |
| smolagents, Strands, Haystack | `@tool` |
| Agno / Phidata | `@tool`, `class MyTools(Toolkit)` |
| Agency Swarm | `@function_tool` |
| MetaGPT | `@register_tool()` |
| Marvin | `@marvin.fn`, `@ai_model` |

Framework attribution uses a confidence-graded resolution chain: direct imports are resolved at 0.95 confidence, inferred instance variables (e.g. `weather_agent = Agent[Deps, T](...)`) at 0.80, and unresolvable decorator names fall back to the best available label at 0.60.

Python entry points feed into the reachability pass — the call graph is traced from each detected entry point to identify which capabilities the LLM can actually trigger.

### Entry point detection — TypeScript and JavaScript

`reachscan` parses `.ts`, `.tsx`, `.js`, `.jsx`, `.mts`, `.mjs`, `.cts`, and `.cjs` files with [tree-sitter](https://tree-sitter.github.io/), using prebuilt Python wheels, so no Node.js runtime is required. Comments and strings don't produce entry points. If a file can't be parsed (for example, TypeScript syntax newer than the bundled grammar), reachscan falls back to regex matching for that file.

| Pattern | What it detects | Confidence |
|---|---|---|
| `mcp_tool` | `server.tool("name", schema, handler)` — MCP SDK | 0.95 |
| `mcp_tool` | `server.registerTool("name", schema, handler)` — MCP SDK v1.6+ | 0.95 |
| `mcp_tool` | `server.addTool({ name: "...", ... })` — FastMCP | 0.90 |
| `mcp_tool_definition` | `{ name: "...", description: ..., inputSchema: ... }` objects | 0.85 |
| `langchain_tool` | `new DynamicTool({ name: "...", ... })` | 0.85 |
| `mcp_handler` | `server.setRequestHandler(Schema, ...)` | 0.80 |

Registration calls are matched however they're formatted. Declaration files (`.d.ts`), test files, minified bundles, and `node_modules`/`dist`/`build` directories are automatically excluded.

TypeScript and JavaScript code is also analyzed for capabilities: `child_process` and `execa` (EXECUTE), `fs` and `fs/promises` (READ/WRITE), `fetch`, `axios`, `got`, `undici`, `http(s)`, `net`, and WebSockets (SEND), `process.env`, `dotenv`, and `keytar` (SECRETS), `eval`, `new Function`, `vm`, and non-literal `import()`/`require()` (DYNAMIC), and `setInterval`, cron libraries, and worker threads (AUTONOMY). Imports are resolved first, so `regex.exec()` or a local `exec` helper isn't mistaken for `child_process.exec`.

TypeScript reachability works like Python's. Each tool's handler is an entry node, and the call graph follows direct calls to functions in the same file, relative imports (ESM, including `./x.js` → `x.ts`, CommonJS `require`, and namespace imports), `this.method()` within a class, and methods of object literals. Callbacks defined inside a function are treated as part of it. Method calls the graph can't resolve, such as `tool.execute()` on a class instance or on a tool object taken from a list, aren't followed: code that only such a call could reach is `unknown`. Other code that isn't on a path, including code reached only through computed calls (`table[name](...)`) or only from top-level code (for example the constructor of a module-level singleton), is `unreachable`. Path depth, states, and exit codes match the Python analysis.

Handlers are found for the patterns above, and also for: `addTool(toolObject)` with a tool object defined in the project; tool-definition objects using `schema`, `parameters`, or `args` instead of `inputSchema`, with a `handler` or `execute` property or method (for example `defineTool({ ..., handler })`); objects passed to a project wrapper that itself calls `registerTool` / `tool` / `addTool`; [xmcp](https://xmcp.dev) file-based tools (a file exporting `metadata` and a default function, in projects that depend on xmcp); and `server.tool(...)` / `registerTool(...)` calls whose name isn't a literal (reported with the name `unknown`, in files that import an MCP SDK).

### Verifying MCP tool annotations

MCP tools can declare [`ToolAnnotations`](https://modelcontextprotocol.io/specification/2026-07-28/schema#toolannotations) hints such as `readOnlyHint`, `openWorldHint`, and `destructiveHint`. Clients use them to decide what to auto-approve, but they're claims the server makes about itself. For Python MCP servers, reachscan checks each claim against what the tool can actually reach:

| Declared | Contradicted by a reachable... | Severity |
|---|---|---|
| `readOnlyHint: true` | WRITE, EXECUTE, or DYNAMIC | high |
| `openWorldHint: false` | outbound HTTP, websocket, or raw socket connect (not to a literal loopback host such as `localhost` or `127.0.0.1`; calls into the project's own modules, database drivers, and other protocol clients don't count) | high |
| `destructiveHint: false` (with `readOnlyHint: false`) | delete, move/rename, or truncating write | medium |

```text
Annotation Mismatches  —  MCP tool annotations contradicted by reachable code
-----------------------------------------------------------------------------
  [HIGH] get_report declares readOnlyHint: true
    but reaches WRITE via os.remove() (server.py:9)
    path: get_report → _cleanup → os.remove()
```

Every mismatch comes with the call path from that tool to the contradicting code. A contradiction without such a path isn't reported. Only explicitly declared hints are checked: absent hints fall back to the spec's conservative defaults, which claim nothing, and hints reachscan can't resolve statically (imported from outside the project, built by a helper function) are skipped and listed under `--explain`. FastMCP `@mcp.tool(annotations=...)` and lowlevel `types.Tool(...)` declarations are both supported. Lowlevel tools are linked to their branch in the `call_tool` handler when it dispatches with `if name == ...` or `match name:`. Mismatches appear in the text report, in JSON (`annotation_mismatches`, schema 1.1), and in SARIF as rule `mcp-risk-mismatch`. TypeScript support is planned.

---

## What it looks like

```text
Agent Capability Report
=======================

Python Entry Points (LLM-controlled surface)
----------------------------------------------
  • get_lat_lng  (pydantic_ai/decorator @ weather_agent.py:50)
  • get_weather  (pydantic_ai/decorator @ weather_agent.py:67)

Capabilities
------------
  • SEND

Combined Risks
--------------
  None inferred from combined-capability rules.

Reachability Summary
--------------------
     3 reachable     — LLM can trigger these directly
   117 unreachable   — exist in codebase, not on any LLM call path
     3 module-level  — execute on import, not on any call path

Reachable Findings  —  LLM can trigger these directly
------------------------------------------------------
  [HIGH] SEND via ctx.deps.client.get -> https://api.weather.example.com (network @ weather_agent.py:58)
    path: get_lat_lng
    explanation: This code can send data over the network to external services.
    impact: Sensitive local data could be transmitted to untrusted endpoints.

Other Findings  —  not on LLM call path
-----------------------------------------
  [HIGH] UNREACHABLE  SECRETS via os.getenv('ANTHROPIC_API_KEY') (secrets @ model_client.py:12)
    explanation: This code accesses secrets or credential sources.
    impact: Credentials may be disclosed and used for unauthorized access.

  [HIGH] MODULE_LEVEL  SECRETS via os.getenv('PYDANTIC_AI_MODEL') (secrets @ config.py:25)
    reachability: Executes on import — runs whenever this module loads
    explanation: This code accesses secrets or credential sources.
    impact: Credentials may be disclosed and used for unauthorized access.
```

You get file paths and line numbers. Not just "this repo uses subprocess" — you get exactly where, how, and whether the LLM can reach it.

---

## Who needs this

**Agent developers** — audit your own code before shipping. Know exactly what you're granting the LLM access to, and where those grants live in your codebase. Add the [GitHub Action](#ci-integration) to catch new reachable capabilities in every pull request.

**Security and platform teams** — you're deploying agents your developers wrote, or agents that use third-party frameworks. Before they hit production, run a scan. Get a fast, defensible answer to "what can this thing actually do?"

**Anyone integrating third-party tools** — tools, plugins, and MCP servers come with capabilities attached. Scan them *before* wiring them into your agent. `reachscan https://github.com/some-org/some-tool` takes seconds and requires nothing installed on that repo, or ask your coding agent to do it with the [Claude Code / Codex plugin](#use-it-from-claude-code-or-codex).

**MCP server authors** — show your users exactly what your server can and cannot do. A clean scan result is a trust signal, and the [GitHub Action](#ci-integration) keeps it clean as the server changes.

---

## It's not just for agents

The name is intentional but the scope is broader.

Any Python or TypeScript/JavaScript code that runs in an AI-adjacent context is a valid target — tool libraries, retrieval pipelines, memory modules, execution sandboxes. If an LLM can call it, you want to know what it can do.

---

## Precision

Detection quality was validated in a structured false positive audit across 10 major open-source agent repos (AutoGPT, LangChain, LlamaIndex, CrewAI, OpenAI Agents SDK, Autogen, pydantic-ai, agentops, anthropic-cookbook, python-sdk) — approximately 3,900 labeled findings:

| Detector | FP Rate |
|---|---|
| `file_access` | 0.0% |
| `secrets` | 0.0% |
| `dynamic_exec` | 0.0% |
| `network` | 0.7% |
| `autonomy` | 1.6% |
| `shell_exec` | 1.9% |
| **Overall** | **0.47%** |

Low noise by design. When it fires, it's real.

This audit covers the Python detectors. The TypeScript/JavaScript detectors are newer and will get their own audit.

---

## What this is NOT

- Not a vulnerability scanner
- Not a linter
- Not a dependency checker
- Not a compliance tool
- Not a prompt injection detector

**It is a capability audit.** Static analysis only — results describe what the code is capable of, not what it will do in any given execution.

---

## Installation

### Option 1 — Recommended (install as a CLI tool)

```bash
pipx install reachscan
```

Or with pip:

```bash
pip install reachscan
```

Then run:

```bash
reachscan .
```

### Option 2 — Install from source (development)

```bash
git clone https://github.com/vinmay/reachscan.git
cd reachscan

python -m venv .venv
source .venv/bin/activate      # Windows: .venv\Scripts\activate

pip install -e .[dev]
```

### Option 3 — Run without installing

```bash
python -m reachscan.cli examples/demo_agent
```

---

## Use it from Claude Code or Codex

The reachscan plugin adds a `vet-mcp-server` skill to your coding agent. Before you add an MCP server, ask:

```
vet this MCP server before I add it: https://github.com/org/some-mcp-server
```

The agent runs reachscan on the server and tells you what each tool can execute, read, write, and send, with call paths for high-risk findings and an **Install / Review first / Avoid** verdict based only on the scan results. The plugin needs the reachscan CLI on your `PATH` (`pipx install reachscan`) and a terminal, so it works in Claude Code and Codex.

**Claude Code** (inside a session):

```
/plugin marketplace add vinmay/reachscan
/plugin install reachscan@reachscan
```

**Codex:**

```bash
codex plugin marketplace add vinmay/reachscan
```

Then open the Plugins Directory, choose the **reachscan** marketplace, and install the plugin. Details are in [`integrations/agent-plugins`](integrations/agent-plugins/README.md).

---

## Requirements

- Python 3.11+
- pip or pipx

---

## Usage

```
reachscan [target] [--json | --sarif] [--severity {high,medium,none}] [--explain]
```

`target` accepts:

| Input | Example |
|---|---|
| Local path | `reachscan .` |
| Local path, JSON output | `reachscan ./my_agent --json` |
| Local path, SARIF output | `reachscan ./my_agent --sarif` |
| GitHub repository URL | `reachscan https://github.com/org/repo` |
| MCP HTTP endpoint | `reachscan mcp+https://mcp.example.com` |
| PyPI package (latest) | `reachscan pypi:requests` |
| PyPI package (pinned) | `reachscan pypi:requests==2.31.0` |

The GitHub URL path does a shallow clone — you don't need the repo checked out locally.

### Exit codes

| Code | Meaning |
|------|---------|
| `0` | Scan complete, threshold not exceeded |
| `1` | Scan complete, ≥1 reachable finding exceeds severity threshold |
| `2` | Scan failed (bad target, network error, unhandled exception) |

### `--severity` flag

Controls when the CLI exits 1:

| Value | Exit 1 when... |
|-------|----------------|
| `high` *(default)* | reachable finding with `risk_level == "high"` |
| `medium` | reachable finding with `risk_level in ("high", "medium")` |
| `none` | never — always exits 0 |

### `--explain` flag

Expands the call chain for every reachable finding, showing each hop with its source file. Use this when you want to understand exactly how the LLM reaches a capability — not just that it can, but through which functions.

Without `--explain`:
```
  [HIGH] DYNAMIC via exec() (dynamic_exec @ addon.py:431)
    path: execute_blender_code → … → execute_code
```

With `--explain`:
```
  [HIGH] DYNAMIC via exec() (dynamic_exec @ addon.py:431)
    call chain:
      execute_blender_code @ server.py
      → send_command @ server.py
      → execute_code @ addon.py
```

Only applies to the text report. Has no effect with `--json` or `--sarif`.

### `--sarif` flag

Writes [SARIF 2.1.0](https://docs.oasis-open.org/sarif/sarif/v2.1.0/sarif-v2.1.0.html) to stdout, for GitHub code scanning and other SARIF viewers. It can't be combined with `--json`, and exit codes are the same.

- One rule per capability (`reachscan/EXECUTE`, `reachscan/SEND`, ...) and one per combined risk (`reachscan/combined/remote_control`, ...).
- Levels: a reachable high-risk finding is `error`, a reachable medium-risk finding is `warning`, and everything else is `note`.
- Each reachable finding has a `codeFlow` that walks the call chain from the LLM entry point to the sink, so the code scanning UI shows the path step by step.
- Reachability state, confidence, and entry point are in each result's `properties`.
- By default only `reachable` and `module_level` findings are included, so the Security tab shows only what an LLM can trigger. Add `--sarif-include-unreachable` to include everything.
- When a language has findings but no detected entry points, reachability isn't evaluated for it and its findings are left out by default. The SARIF run then carries a warning notification (`invocations[].toolExecutionNotifications`) saying how many findings weren't shown, so an empty Security tab isn't mistaken for a clean scan.

---

## CI Integration

The easiest way is the [reachscan GitHub Action](https://github.com/marketplace/actions/reachscan). It installs reachscan, uploads findings to the GitHub Security tab with call chains, and fails the job when a reachable finding meets your severity threshold:

```yaml
name: reachscan
on: [push, pull_request]
permissions:
  contents: read
  security-events: write
jobs:
  reachscan:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v7
      - uses: vinmay/reachscan-action@v1
        # with:
        #   severity: medium   # or none for report-only
        #   path: servers/my-mcp-server
```

Inputs, outputs, and more examples are in the [action's README](https://github.com/vinmay/reachscan-action).

To run the CLI yourself instead, for example to keep a JSON report as a build artifact:

```yaml
- uses: actions/checkout@v7
- run: pipx install reachscan
- name: Run capability audit
  run: reachscan . --json > reachscan-report.json
  # Exits 1 if HIGH reachable capabilities found
- uses: actions/upload-artifact@v7
  if: always()
  with:
    name: reachscan-report
    path: reachscan-report.json
```

To audit without blocking the pipeline (report only):

```yaml
- run: reachscan . --json --severity none > reachscan-report.json
```

To upload SARIF to the GitHub Security tab without the action:

```yaml
permissions:
  security-events: write
  contents: read
steps:
  - uses: actions/checkout@v7
  - run: pipx install reachscan
  - name: Run reachscan
    run: reachscan . --sarif > reachscan.sarif
  - uses: github/codeql-action/upload-sarif@v4
    if: always()
    with:
      sarif_file: reachscan.sarif
      category: reachscan
```

---

## Project direction

The goal:

> Give AI systems a permission model they've never had.

Static capability detection is the foundation. Reachability analysis on top of it answers the harder question: not just *can* this code do something, but *can the LLM trigger it*.

---

## Status

What works today:

| Area | Status |
|---|---|
| Python | Capability detection, entry points for the frameworks above, call-graph reachability (up to 8 hops), MCP annotation verification |
| TypeScript / JavaScript | Parsed with tree-sitter (no Node.js needed). Entry point detection, capability detection for all seven classes, and reachability through TS call paths |
| Scan targets | Local paths, GitHub URLs, PyPI packages (`pypi:name[==version]`), MCP HTTP endpoints (`mcp+https://...`) |
| Output | Text report, JSON ([schema v1](docs/schema_v1.md)), SARIF 2.1.0 with call chains, `--explain` call traces |
| CI | [GitHub Action](https://github.com/marketplace/actions/reachscan) with Security tab upload and a severity gate; exit codes for any other CI |
| Coding agents | [Plugin for Claude Code and Codex](#use-it-from-claude-code-or-codex) that vets MCP servers before you install them |
| Precision | 0.47% false-positive rate across ~3,900 labeled Python findings ([details](#precision)); a TypeScript audit is planned |

The JSON output schema is stable at v1 — see [`docs/schema_v1.md`](docs/schema_v1.md) for the full field reference. Feedback, edge cases, and false positive reports are especially valuable — open an issue.

---

## Support

If reachscan is useful to you, [star the repo](https://github.com/vinmay/reachscan) — it helps others find it and tells us people care about this problem.
