Metadata-Version: 2.4
Name: megabrain
Version: 0.3.0
Summary: Local code-intelligence engine: one call returns all the code related to a question, explained with the real code spliced in.
Author-email: Berna Castro <bernacas@gmail.com>
License-Expression: MIT
Project-URL: Homepage, https://github.com/bernatch22/megabrain
Project-URL: Repository, https://github.com/bernatch22/megabrain
Keywords: code-intelligence,retrieval,rag,embeddings,code-search,mcp,ast,tree-sitter,developer-tools
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Operating System :: OS Independent
Classifier: Intended Audience :: Developers
Classifier: Topic :: Software Development :: Libraries
Classifier: Topic :: Software Development :: Documentation
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: numpy>=1.24
Requires-Dist: tree_sitter>=0.21
Requires-Dist: tree_sitter_typescript>=0.23
Provides-Extra: languages
Requires-Dist: tree_sitter_ruby>=0.23; extra == "languages"
Requires-Dist: tree_sitter_go>=0.23; extra == "languages"
Requires-Dist: tree_sitter_rust>=0.23; extra == "languages"
Requires-Dist: tree_sitter_php>=0.23; extra == "languages"
Dynamic: license-file

<p align="center">
  <img src="https://raw.githubusercontent.com/bernatch22/megabrain/master/assets/megabrain.png" alt="megabrain" width="180">
</p>

<h1 align="center">megabrain</h1>

<p align="center">
  <b>One call returns all the code related to a question</b><br>
  — explained like a senior engineer, with the real code spliced in.
</p>

<p align="center">
  <a href="https://pypi.org/project/megabrain/"><img src="https://img.shields.io/pypi/v/megabrain?style=flat-square&color=3776AB" alt="PyPI"></a>
  <img src="https://img.shields.io/badge/python-3.10+-3776AB?style=flat-square&logo=python&logoColor=white" alt="Python 3.10+">
  <img src="https://img.shields.io/badge/license-MIT-green?style=flat-square" alt="MIT">
  <img src="https://img.shields.io/badge/retrieval-no%20LLM%20·%20~200ms-2ea44f?style=flat-square" alt="No LLM in the retrieval path">
  <img src="https://img.shields.io/badge/code-zero%20hallucination-6f42c1?style=flat-square" alt="Zero code hallucination">
  <img src="https://img.shields.io/badge/MCP-ready-000000?style=flat-square" alt="MCP ready">
</p>

---

**megabrain** is a local code-intelligence engine. It replaces minutes of file-by-file
crawling — grep, read, explore-agent chains — with a single grounded answer. Index a repo
once; every later question retrieves *all* the related code and stitches it into a
walkthrough narrated by an LLM that can **only point at code, never rewrite it** — so
nothing is hallucinated. Retrieval itself uses **no LLM** (~200 ms); the one LLM call
just narrates.

## Languages

Code is chunked over its real AST (the [cAST](https://arxiv.org/abs/2506.15655)
split-then-merge recipe), so chunks are whole functions/classes with breadcrumbs — never
arbitrary line windows.

| | languages | how |
|---|---|---|
| **built-in** | **Python**, **TypeScript / JS / JSX / TSX / MJS / CJS**, **Markdown** | stdlib `ast` · tree-sitter · no-LLM doc chunker |
| **`[languages]` extra** | **Ruby**, **Go**, **Rust**, **PHP** | tree-sitter grammars |

Adding a language is a `LangSpec` entry + `pip install tree_sitter_<lang>` — a config
entry in a registry, not a branch in the indexer. Import/call **graph** edges are built
for Python and TS/JS today; other languages retrieve on dense+lexical signals (no graph
needed for correctness).

## Install

```bash
pip install megabrain                 # core: Python · TS/JS · Markdown
pip install 'megabrain[languages]'    # + Ruby · Go · Rust · PHP
```

From a clone, for development:

```bash
git clone https://github.com/bernatch22/megabrain.git && cd megabrain
pip install -e '.[languages]'
python3 -m pytest                      # offline test suite — no network, no key
```

## Setup

One key, read from the environment (with a `~/.zshrc` fallback):

```bash
export OPENROUTER_API_KEY=...          # embeddings + ask, all via OpenRouter
```

Everything runs through OpenRouter's OpenAI-compatible API, so **any model works** — pick
per role by env (the defaults reproduce the validated stack):

```bash
export MEGABRAIN_EMBED_MODEL=perplexity/pplx-embed-v1-0.6b   # embeddings (default)
export MEGABRAIN_ASK_MODEL=qwen/qwen3-coder                  # ask narrator (default; ~5x cheaper than Haiku, on par)
```

## Usage

```bash
megabrain index  ~/repo                                    # incremental (sha256), no daemon
megabrain ask    ~/repo "how does auth work end to end"    # walkthrough + real code (~6–20s)
megabrain ask    ~/repo/src/auth "how are tokens issued"   # scope to a sub-path (path-scope)
megabrain ask    ~/repo "how do I configure X" --docs      # explain the docs, not the code
megabrain query  ~/repo "request retry logic"              # raw code map, no LLM (~200ms)
megabrain get    ~/repo src/x.py --symbol Class.method     # one file or symbol
megabrain serve-api ~/repo --port 2134                     # long-running JSON API (warm state)
```

**Path-scope:** pass a sub-folder (`~/repo/src/auth`) to any of `ask` / `query` / `get`
and retrieval is confined to files under it — the repo root (where the index lives) is
auto-detected. Multi-repo works too: `megabrain query ~/a/src,~/b "..."`.

## Provider flexibility — cloud, native, local, hybrid

Embeddings and chat can **each** point at any OpenAI-compatible endpoint. localhost
servers (Ollama / LM Studio / vLLM) need **no API key**:

```bash
# native provider (e.g. A/B a model directly):
export MEGABRAIN_EMBED_BASE_URL=https://api.perplexity.ai/v1   # uses PERPLEXITY_API_KEY

# hybrid — local private embeddings + cheap OpenRouter narration:
export MEGABRAIN_EMBED_BASE_URL=http://localhost:11434/v1
export MEGABRAIN_EMBED_MODEL=embeddinggemma
export MEGABRAIN_EMBED_BATCH=8            # smaller requests for local servers

# fully local (decent GPU) — nothing leaves the machine:
export MEGABRAIN_CHAT_BASE_URL=http://localhost:11434/v1
export MEGABRAIN_ASK_MODEL=qwen3-coder:30b
```

Changing the embed model auto-triggers a full re-embed on the next `index` (or force it
with `--force`), so vectors never silently mismatch. Local-stack benchmarks live in
`evals/LOCAL_MODELS.md`.

## How it works

A three-stage pipeline. **Only `ask` calls an LLM — and only to narrate.**

| stage | what it does |
|---|---|
| **index** | cAST chunk → embed (`pplx-embed-v1-0.6b`, int8, L2-normalized) → SQLite. Incremental by `sha256`, no watcher. |
| **query** | No-LLM retrieval (~200 ms): dense-chunk + file-skeleton fusion, with import/call-graph candidates. Returns a map — **CORE** (full code of the top files) + **RELATED** (every connected file with its best chunk). |
| **ask** | One streamed chat call (qwen3-coder by default) writes the walkthrough and cites code as `[[k]]`; the engine **replaces each citation with the verbatim block** (real file, real line numbers). Non-cited files are listed at the end. Fail-open: any API error falls back to the full `query` bundle. |

Because the model only emits citations and the engine splices code from disk, **code
cannot be hallucinated or rewritten.**

## MCP

Use it from Claude Code or any MCP client:

```bash
claude mcp add megabrain -- python3 -m megabrain.mcp_server
```

Tools: `megabrain_ask` (primary), `megabrain_query`, `megabrain_get`, `megabrain_index` —
`ask`/`query` take an optional `scope_path` for sub-path retrieval. The server
auto-refreshes a stale index before answering, so results always match disk.

## HTTP API

`serve-api` keeps the index warm in memory and serves retrieval over HTTP (stdlib only —
no framework). Embed it in an app, or front a docs site with real semantic search.

```bash
megabrain serve-api ~/repo --port 2134 [--host 0.0.0.0] [--cors https://site] [--no-llm]
```

| route | returns |
|---|---|
| `POST /search` `{query}` | raw bundle (`tier1` / `tier2`), same as `query` |
| `GET /docsearch?q=` | doc-search hits — `{title, slug, snippet, context, score, group}` |
| `POST /ask` `{question}` | LLM walkthrough (`{text, …}`) |
| `GET /get?file=&symbol=` · `POST /index` · `GET /health` | one file/symbol · reindex · status |

Binds localhost by default (front it with a reverse proxy); `--cors` opts into a browser origin.

## Design

Every choice below is backed by an internal golden set (30 verified queries):

| decision | evidence |
|---|---|
| cAST chunking (4K nws chars, breadcrumbs, partition-guaranteed) | unit-tested; every line lands in exactly one chunk — no gaps, no overlaps |
| `pplx-embed-v1` via OpenRouter (1024-d, int8 wire, **L2-normalized**) | beat `openai-3-large` on code in a bakeoff; ~$0.0016/repo |
| dense chunk + 0.5 × file-skeleton score | dual-granularity; precision up, no downside |
| graph (import + call edges) for candidates only | PageRank-as-ranking **rejected** by data (Acc@1 0.91 → 0.73) |
| **no LLM in the retrieval path** | every LLM *prune* variant cost completeness; `ask` explains, it never prunes |

**Engine retrieval** (internal golden set): R@1 **0.86** · bundle\_full **1.00** · p50 **~10 ms** warm.
**SWE-bench Lite** localization (no training): retrieval Acc@1 ≈ 0.52 / @5 ≈ 0.83 — on par
with the trained CodeRankEmbed retriever.

## Project layout

```
megabrain/   engine — chunkers, providers, embeddings, SQLite store, graph, indexer, query, ask, serve, cli, mcp_server
tests/       offline suite (no network/key/corpus) — run with `python3 -m pytest`
evals/       golden set + model bakeoffs (maintainer-side, private corpus)
```

---

<p align="center"><sub>MIT · github.com/bernatch22/megabrain</sub></p>
