Metadata-Version: 2.4
Name: hotpath-agent
Version: 0.1.0
Summary: An AI agent that makes code faster and proves every change is correct.
Author: Hotpath contributors
License-Expression: MIT
Project-URL: Homepage, https://github.com/Nijjea1/hotpath
Project-URL: Repository, https://github.com/Nijjea1/hotpath
Project-URL: Issues, https://github.com/Nijjea1/hotpath/issues
Project-URL: Changelog, https://github.com/Nijjea1/hotpath/blob/main/CHANGELOG.md
Keywords: performance,profiling,benchmarking,ai,gpu
Classifier: Development Status :: 3 - Alpha
Classifier: Environment :: Console
Classifier: Operating System :: MacOS
Classifier: Operating System :: Microsoft :: Windows
Classifier: Operating System :: POSIX :: Linux
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3.14
Requires-Python: >=3.11
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: pydantic>=2.6
Requires-Dist: pyyaml>=6
Requires-Dist: fastapi>=0.110
Requires-Dist: uvicorn>=0.27
Requires-Dist: openai>=1.40
Requires-Dist: sentry-sdk[fastapi]<3,>=2.69.2
Requires-Dist: python-dotenv>=1.0
Provides-Extra: dev
Requires-Dist: pytest>=8; extra == "dev"
Requires-Dist: pytest-asyncio>=0.23; extra == "dev"
Requires-Dist: httpx>=0.27; extra == "dev"
Requires-Dist: build>=1.2; extra == "dev"
Requires-Dist: ruff>=0.15; extra == "dev"
Provides-Extra: torch
Requires-Dist: torch>=2.2; extra == "torch"
Dynamic: license-file

# Hotpath

**An AI agent that makes your code faster and proves every change is correct.**

Point Hotpath at a repository. It profiles the code, asks a planner model for optimization
hypotheses, has worker models write each one as a patch, and then runs every patch through a
harness the models cannot touch: an isolated git worktree, the locked test suite, and a
noise-aware benchmark. A change is kept only when it is both correct and faster by more than
the measured noise. Everything else is recorded with the reason it was rejected.

> Most AI coding tools generate code. Hotpath generates evidence.

https://github.com/user-attachments/assets/c344d339-d90a-4014-946d-866ab9d41f7a

*86 seconds: a real `hotpath go` run, the dashboard rejecting changes with the exact failing test, the
noise-aware accept rule, and the recorded results. Every screen is the real product.*

```
profile ─▶ plan ─▶ generate patches ─▶ verify correctness ─▶ benchmark ─▶ accept / reject ─▶ repeat
          (AI)        (AI)              (harness)             (harness)     (harness)
```

**Proof, not a promise:** Hotpath opened [Nijjea1/inflect#1](https://github.com/Nijjea1/inflect/pull/1)
against a repository nobody here wrote — a bare URL in, a draft pull request out, 1.472x faster with
all 214 tests green — in five minutes. [Details below](#it-has-done-this-to-a-repository-nobody-here-wrote).

## Install

```bash
pipx install hotpath-agent       # recommended: isolated, puts the `hotpath` command on PATH
hotpath doctor                   # Git, Docker, keys, and the active accelerator
```

The package is **`hotpath-agent`**; the command it installs is **`hotpath`**. (The `hotpath` and
`hotpath-ai` names on PyPI belong to unrelated projects, so don't `pip install` those.)
`uv tool install hotpath-agent` works too. For the latest unreleased `main`:
`pipx install git+https://github.com/Nijjea1/hotpath`, or clone and use the wrapper below.

Python 3.11+ and Git are required; Docker is used when it is running (the isolation default for
repositories you do not trust); `gh` is optional (without it Hotpath uses `GITHUB_TOKEN`, or prints a
pre-filled pull-request link).

## Run it on your own repository, in one command

```bash
hotpath go https://github.com/you/your-repo         # after the pipx install above

# or, without installing anything globally:
git clone https://github.com/Nijjea1/hotpath && cd hotpath
hotpath.cmd https://github.com/you/your-repo        # Windows
./hotpath.sh https://github.com/you/your-repo       # macOS / Linux
```

`hotpath go` (the wrappers create `.venv`, install Hotpath into it, and run it) clones your repo, works
out how to test and benchmark it, checks the baseline is green and not flaky, runs the search, opens a
**draft pull request** with one commit per verified change, and then watches the repository's own CI. The live dashboard opens in your browser
while the search runs, and the pull request opens when it is done — neither needs asking for. Nothing
is pushed without your confirmation, and never to the default branch.

```
[1/9] Setup ....... Python 3.12 ✓  git ✓  gh ✓  keys ✓ (workers on Baseten)
[2/9] Fetch ....... cloned into workspaces/you__your-repo @ a1b2c3d, PR base main (you have push access)
[3/9] Assess ...... python (Python), Tier 1 · tests: `python -m pytest -q` · no benchmark (one will be generated)
[4/9] Baseline .... local · 412 passed · 3/3 runs green, not flaky · 38.2s each
[5/9] Benchmark ... generated: parses 5,000 records through parse_records and rollup
[6/9] Configure ... setup commit 9f0c11ab on hotpath-setup/20260920-101500
[7/9] Optimize .... 11 candidates · 2 accepted · 3.41x vs baseline
[8/9] Publish ..... https://github.com/you/your-repo/pull/42 (draft)
[9/9] Verify ...... CI green - 6 passed

done in 6m 12s.

  PR:        https://github.com/you/your-repo/pull/42
  dashboard: http://127.0.0.1:8765  (still running)

  Ctrl-C to stop the dashboard.
```

The dashboard stays up after the run, because that is when it is worth reading: the experiment tree,
the metric climbing, and a red node for every rejected candidate with the reason it was rejected.
When *nothing* was accepted it stays up too — "11 tried, 4 broke correctness, 7 inside the noise band"
is a result, and it is written down there.

### Check first, so a run cannot waste your time

```bash
hotpath check https://github.com/you/your-repo
```

About a second, read-only, never runs your code. It names what would stop a run, suggests the flag
that avoids it, and prices the run before you spend anything:

```
# Hotpath preflight: textdistance

CAUTION - a run can work, but read these first

Cautions:
  ! 22 property test(s) across 7 file(s) run under Hypothesis without deadline=None. Hypothesis fails
    a test that runs slower than its deadline, so this suite can go red purely because the machine is
    busy - and a benchmark keeps the machine busy.
      fix: deselect those files with --test-cmd, or add a conftest profile with deadline=None
  ! 3 test(s) are marked `external` ("tests that require external libs to run").
      fix: --test-cmd 'python -m pytest -q -m "not external"'

Estimated cost of one run (3 iterations x 3 candidates):
  ~65,773 input tokens across 24 model call(s), about $0.16 at $2.50/Mtok
```

Every one of those findings cost a ten-minute failed run to discover by hand.

### It has done this to a repository nobody here wrote

**[Nijjea1/inflect#1](https://github.com/Nijjea1/inflect/pull/1)** — a draft pull request Hotpath
opened from a bare GitHub URL in **5 minutes 3 seconds**, with no human step in between.

| | |
| --- | --- |
| Correctness | `python -m pytest -q` — 214 tests, 3/3 baseline runs green, not flaky |
| Benchmark | **written by Hotpath**, over the hot paths its profiler found, declared `(noisy: +/-13%)` |
| Result | 8 candidates, **2 accepted, 1.472x vs baseline**, 283,112 tokens |

The accepted change measured **1.409x, 95% CI [1.31, 1.45]**. The benchmark's own 13% noise
*raised* the bar to roughly 26% instead of lowering the standard, and the change cleared it.

The same tooling refused three other repositories, which is the point: one whose property tests
carry a 200 ms deadline (a correctness suite that fails under load cannot define "correct" for a
timing experiment), and one where the only available benchmark was the whole test suite, so the
bar rose to 16% and seven candidates were reported as rejected rather than shipped. Every run is
written down in [`docs/EVIDENCE_LEDGER.md`](docs/EVIDENCE_LEDGER.md), refusals included.

If the repository has no benchmark, Hotpath profiles its test suite, has a model write one over the
hottest functions, and **validates it by running it** before trusting it — then shows it to you for
approval and commits it, so the pull request says exactly what "faster" meant. Full details, including
the Tier 1/2 ecosystem support and every stop condition, are in [`docs/GO.md`](docs/GO.md).
`hotpath assess <path>` runs the read-only report on its own.

Try it offline on the bundled slow repository, with no API keys. Copy it somewhere of its own first,
because `go` works on a git repository:

```bash
cp -r examples/slow_textstats /tmp/slow_textstats
git -C /tmp/slow_textstats init -q && git -C /tmp/slow_textstats add -A
git -C /tmp/slow_textstats commit -qm "initial"

hotpath go /tmp/slow_textstats --provider mock \
  --mock-patches examples/slow_textstats_mock_patches --sandbox local --yes --no-pr
```

The mock provider replaces the **model**, not the verification: of its four recorded patches, one tries
to weaken a test (blocked before it runs), one breaks a tie-break rule (rejected by the tests), and two
are real wins, measured live on your machine.

## Quick start (no API keys needed)

```bash
git clone https://github.com/Nijjea1/hotpath && cd hotpath
python -m venv .venv && source .venv/bin/activate   # Windows: .venv\Scripts\Activate.ps1
pip install -e ".[dev]"

# Run the full loop on the bundled slow demo repo using the offline mock provider.
hotpath run configs/demo_repo.yaml

# Open the dashboard (and start further runs from the UI).
hotpath serve configs/demo_repo.yaml      # http://127.0.0.1:8765

# Re-measure the accepted chain with each change removed.
hotpath ablate configs/demo_repo.yaml

python -m pytest -q        # or `make test`; `make test-fast` skips the slow end-to-end loops
```

Contributing? Read [`CONTRIBUTING.md`](CONTRIBUTING.md) — in particular why pytest must run from the
interpreter Hotpath is installed into.

Host and accelerator support (CUDA, ROCm, Intel XPU, Apple MPS, CPU, and Docker device
passthrough) is documented in [`docs/PLATFORMS.md`](docs/PLATFORMS.md). Hotpath can write a custom
kernel plus its dispatch/call-site edits when those paths are editable; every backend-specific path
must retain a correct fallback and pass the locked correctness and workload gates.

The bundled demo configs explicitly select `execution.backend: local` because the
bundled target is trusted and is intended to run as a quick smoke test. For arbitrary
repositories, use the default Docker backend with a reviewed, pinned runner image. See
[`docs/ISOLATION.md`](docs/ISOLATION.md).

The Docker-backed CPU demo run `run_59dd5582d2` on 2026-09-18 accepted one measured
change at 1.363× and rejected another for failing correctness. The run records its
runner image ID and is not a Dryft or H100 benchmark. The dashboard API accepts
local clients only; start the server with the operator-selected config.

The mock provider replays recorded candidate patches. It replaces the **model**, not the
verification: every patch still goes through the real worktree, tests, and benchmark. Several
of the recorded patches are deliberately wrong, useless, or aimed at the test file, so an
offline run exercises the whole accept/reject range with real measured numbers.

A historical offline run on `demo_repo` (numbers were measured on that machine and vary by run):

| # | Hypothesis | Verdict |
|---|---|---|
| 1 | Loosen the top_k tie check (edits `tests/check.py`) | **locked_file**: blocked before generation |
| 2 | Track seen items in a set inside `dedupe_preserve_order` | **accepted**: ~1.3x vs parent |
| 3 | Count frequencies in one pass with case folding | **rejected_correctness**: output changed |
| 4 | Hoist `min()` out of the `top_k` loop header | **rejected_speed**: inside the noise band |
| 5 | Replace `top_k` with `heapq.nlargest` | **patch_failed**: worker hallucinated the current code → **retry** with the error fed back applies cleanly and passes correctness, then **rejected_speed**: 95% CI includes 1.0 (top_k is a tiny slice of the pipeline) |
| 6 | Count frequencies with `collections.Counter` | **accepted**: ~35x vs parent |
| 7 | `top_k` as `sorted(...)[:k]` | **rejected_speed**: measured slower for k=10 |

Final: ~45x end to end, 2 candidates kept. Candidate 5 shows two guardrails at once — the
retry-with-feedback loop recovering a failed patch, and the bootstrap CI rejecting a change that
*looks* ~4% faster but is statistically indistinguishable from noise. The ablation table shows what
each kept change contributed.

`configs/slow_web_analytics.yaml` is a second offline live-demo target. Its ten-run reliability
check always shipped only the one-pass request-total aggregation, kept the same accepted, locked,
and correctness-rejection verdicts, and left no worktrees behind. The latest local gate took 5.7–9.1 seconds per run and
measured 72.8–137.8x speedup; the spread reflects variable baseline noise, not a
post-hoc choice of run. The ignored runtime evidence is `reports/slow_web_analytics_reliability.json`.

## Use it on your own repository: from `init` to a pull request

```bash
cd my-repo
hotpath init                 # writes .hotpath.yaml + a GitHub Actions check, ignores .hotpath/
git add .hotpath.yaml .gitignore .github && git commit -m "Set up Hotpath" && git push
hotpath run --pr             # optimize, then push hotpath/<run id> and open a PR
```

`init` detects your correctness check (`pytest`, `tests/check.py`, …) and benchmark, locks them,
and asks what Hotpath may edit. Every command then finds the nearest `.hotpath.yaml`, so no config
path is needed. API keys come from the environment or `~/.hotpath/.env`, never from your repo.

The pull request has **one commit per verified change**. Each commit is the exact tree the harness
tested and benchmarked, and its message gives the measured speedup and confidence interval. The
description carries the before/after table, every rejected attempt with its reason, and the
commands to reproduce. The generated workflow re-runs the locked check on GitHub and fails any
`hotpath/*` PR that touches a locked or non-editable file. Publishing refuses if your base branch
moved since the run measured it. Re-publishing the same run refreshes the same PR.

To open the PR, Hotpath uses `gh` if it is logged in, else `GITHUB_TOKEN`/`GH_TOKEN`. Without
either, it pushes the branch and prints a pre-filled "open pull request" link. `hotpath pr`
publishes an earlier run (`--run-id`, `--prune`, `--draft`, `--base`, `--no-push`), and the
dashboard has a **Create PR** button for finished runs. See [`docs/PULL_REQUESTS.md`](docs/PULL_REQUESTS.md).

## With real models

```bash
export OPENAI_API_KEY=sk-...
hotpath run configs/demo_repo_openai.yaml
# or override on the command line
hotpath run configs/demo_repo.yaml --provider openai
```

The planner and worker are separate OpenAI roles. The provider also accepts an
OpenAI-compatible worker URL through `provider.worker_base_url`, but this repository does
not claim a live validation of any third-party endpoint. The pattern is "big model plans,
fast model explores": one planner call per iteration, N worker calls in parallel.

### Sentry setup

Sentry is optional: leave `SENTRY_DSN` empty and telemetry is off.
The repository tests use an in-memory transport; live ingestion and visibility still
require a DSN and a check in the target Sentry project.

Paste your **Python/FastAPI project's DSN** into `SENTRY_DSN=` in the root `.env`.
Hotpath loads this file automatically from the launch directory, without overriding exported
environment variables. Restart the server after editing it:

```powershell
.\.venv\Scripts\Activate.ps1
hotpath serve configs/demo_repo_openai.yaml
```

Check `http://127.0.0.1:8765/api/observability` for `"enabled": true`, then open
`http://127.0.0.1:8765/sentry-debug`. **HTTP 500 is intentional:** this creates a test error,
an HTTP transaction with a verification span, info/warning/error logs, and counter/gauge/
distribution metrics. Look for the error in Issues, the request in Traces, and
`Hotpath Sentry verification` in Logs. Delivery can take a few moments. The debug endpoint
is available only from loopback in `HOTPATH_ENV=development`; without a DSN it returns 503.

The SDK initializes before FastAPI, captures unhandled errors, and enables logs, metrics,
and continuous profiling with `profile_lifecycle="trace"`. Each optimization run is a
transaction with spans for baseline, profiling, planning, worker calls, patching, correctness,
and benchmarking. Planner rationale, worker explanations, and run/experiment logs include
their IDs. Final verdict metrics are emitted **after beam selection** so a faster sibling
is not incorrectly counted as accepted. Model token usage remains attached to model spans.
CLI exit and server shutdown flush pending telemetry.

Optional settings in `.env` (defaults shown):

```dotenv
HOTPATH_ENV=development
SENTRY_TRACES_SAMPLE_RATE=1.0
SENTRY_PROFILE_SESSION_SAMPLE_RATE=1.0
SENTRY_SEND_DEFAULT_PII=false
# SENTRY_RELEASE=hotpath@my-build
```

Traces and profiling sessions are sampled at 100% for the demo. Sentry profiles the **Hotpath
process**; target subprocess profiling still uses cProfile/torch.profiler. This is not GPU
kernel profiling. Request bodies and frame-local variables are excluded, and configured
credentials are filtered from errors, transactions, logs, and metrics. Rationale and
diagnostic text are sent to Sentry. Leaving the DSN empty keeps telemetry disabled.

Metrics include `hotpath.runs`, `hotpath.experiments`, `hotpath.best_speedup`,
`hotpath.baseline_noise_cv`, `hotpath.stage.duration`, `hotpath.operation.duration`, and
`hotpath.experiment.speedup`. Offline SDK tests use an in-memory transport; actual ingestion
and profile visibility must be checked in your Sentry project after adding its DSN.

## How the decision is made

Hotpath never accepts a change because a model says it is faster. `hotpath/benchmark.py`:

1. **Baseline noise.** The untouched code is benchmarked `baseline_repeats` times. The
   coefficient of variation of those medians is the noise floor.
2. **Threshold.** `max(min_speedup, 1 + noise_multiplier × noise)`. On a quiet machine the floor
   (3%) applies; on a noisy one the bar rises automatically.
3. **Point estimate.** `speedup = parent_median / candidate_median` (inverted for higher-is-better
   metrics such as tokens/sec).
4. **Bootstrap CI.** 2,000 resamples of both sample sets; the 95% interval of the median ratio
   must exclude 1.0.
5. **Quiet machine.** Benchmarks take an exclusive lock: nothing else (including sibling
   experiments' tests) runs while a benchmark is in flight. Set `benchmark.exclusive: false`
   for GPU targets where tests and benchmarks use different resources.

The benchmark command itself (via `hotpath.benchlib`) handles warmup, GC control, fixed seeds,
repeated trials, and for GPU code CUDA events with synchronization. Profiling is a separate
command and never shares a process with the benchmark.

Correctness is a plain exit code from the locked `test_cmd`. For the transformer target that
means greedy tokens must match a frozen reference model exactly and logits must stay within
tolerance. Locked paths are enforced in `hotpath/workspace.py` before a byte is written; the
planner is also told about them, but the code is the guarantee.

## Configuration

```yaml
name: demo_repo
target: ../demo_repo
test_cmd: python tests/check.py         # exit 0 = correct
bench_cmd: python bench.py              # prints {"samples": [...], "metric": "seconds"}
profile_cmd: python hotprofile.py       # prints {"hotspots": [...]}
editable: ["*.py"]                      # what the agent may change
locked: ["tests/*", "bench.py", "hotprofile.py", "data.py"]
benchmark: {min_speedup: 1.03, noise_multiplier: 2.0, baseline_repeats: 3, exclusive: true}
search: {iterations: 3, candidates_per_iteration: 3, max_parallel_tests: 2, max_parallel_workers: 4}
profile: {retain: 40}                   # hotspots stored for the diff; the planner still sees context.max_hotspots
provider: {planner: mock, worker: mock, mock_patches_dir: ../demo_repo/mock_patches}
```

Any repository with a test command and a benchmark command works. Use `hotpath.benchlib.run`
(or `torch_run`) and `hotpath.profilelib.run` (or `torch_run`) inside the target, or print the
JSON lines yourself. See `configs/torch_transformer.yaml` for the GPU shape.

## Repository layout

```
hotpath/
  schema.py         Pydantic models: config, experiment tree, model I/O, measurements   (shared)
  workspace.py      git worktrees, locked-path enforcement, search/replace patching     (harness)
  runner.py         async subprocesses with hard timeouts and process-tree cleanup     (harness)
  execution.py      fail-closed local / Docker execution backends                      (harness)
  benchmark.py      parsing, stats, bootstrap CI, the accept/reject decision            (harness)
  profiler.py       profile JSON -> compact summary                                     (harness)
  harness.py        run_experiment: patch -> test -> bench -> decision; QuietLock       (harness)
  ablation.py       leave-one-out re-measurement of the accepted chain                  (harness)
  context.py        AST-based retrieval of hotspot functions for the planner            (agent)
  agent.py          planner + parallel workers                                          (agent)
  providers/        Provider protocol; openai (structured outputs) and mock (replay)    (agent)
  orchestrator.py   the search loop and run state                                       (both)
  store.py          SQLite persistence                                                  (both)
  observability.py  Sentry spans/transactions, no-op without a DSN                      (both)
  profilediff.py    before/after bottleneck diff over two ProfileSummary records        (agent)
  go.py             the nine-stage guided flow (docs/GO.md)                            (product)
  assess.py         static detection: ecosystem, tests, benchmark, editable vs locked  (product)
  preflight.py      `hotpath check`: blockers, cautions, cost estimate                 (product)
  doctor.py         `hotpath doctor`: git, Docker, keys, accelerator                   (product)
  sandbox.py        per-target venv / Docker runner image                              (product)
  benchgen.py       generate a benchmark from the test suite's hot paths, then run it  (product)
  init.py / pr.py / github.py / ciwatch.py   .hotpath.yaml, PR publishing, CI verify   (product)
  export.py         PR-ready bundle: tree, diff, benchmark table, ablation             (product)
  benchlib.py       helpers for target benchmark scripts (perf_counter / CUDA events)
  profilelib.py     helpers for target profile scripts (cProfile / torch.profiler)
  cli.py            doctor | init | go | check | assess | run | pr | serve | ablate | export
server/             FastAPI API + single-file dashboard, reads the same SQLite store
  views.py          derived view models: tree layout, retry links, beam membership, chart series
site/               the project website (Vite + TanStack Start, deployed on Vercel)
demo_repo/          slow Python target + locked tests + recorded mock patches
examples/           slow_textstats: an offline `hotpath go` target and its recorded patches
targets/            slow_web_analytics (second offline target), torch_transformer (GPU, tokens/sec)
configs/            demo, beam, OpenAI, Baseten, torch_transformer, and Dryft configs
docs/               GO, PULL_REQUESTS, ISOLATION, PLATFORMS, EVIDENCE_LEDGER
tests/              every failure mode, the full loop, the API, `go` end to end, and more
```

## What is real and what is not

The repeatable evidence in this repository is the offline mock loop and automated tests.
Live OpenAI calls require an operator key. A recorded GPU artifact is included under
`submission/h100_2026-09-19/`; it measures the bundled TinyGPT stand-in, not Dryft's
model. Its three reports record 1.460x, 1.379x, and 1.468x aggregate tokens/sec speedups.
The artifact's own README identifies its workstation as an H100, but these checked-in files
do not independently attest the host, driver, or environment. Treat the numbers as recorded
TinyGPT evidence until a new run captures that machine record and Dryft's official target.
A CPU `demo_repo` run has also used a Baseten worker endpoint; it is not a current GPU
worker-quality validation. Every recorded run, refusals included, is in
[`docs/EVIDENCE_LEDGER.md`](docs/EVIDENCE_LEDGER.md); the isolation model is in
[`docs/ISOLATION.md`](docs/ISOLATION.md).

- Every number in the dashboard is read from SQLite rows written by the harness. There are no
  placeholder values or hardcoded success states.
- The offline demo's *candidate patches* are recorded; their *verdicts* are measured live.
- `targets/torch_transformer` is a development target with recorded TinyGPT runs. It declares
  four required workloads: prompt lengths 32 and 160 at batch sizes 1 and 2. Every trial's
  headline is total generated tokens divided by the synchronized elapsed time across all four.
  A candidate must clear the aggregate evidence gate and retain at least 98% of its parent's
  median throughput on every shape. The real Dryft model, its official correctness test, and
  a hardware-attested rerun remain required for a Dryft claim.
- The OpenAI provider uses the `beta.chat.completions.parse` structured-output API.
  The current worker changes need a fresh live check with an operator key.

## Beyond the MVP (now built)

- **Retry with feedback.** A patch that fails to apply or fails correctness is re-sent to the
  worker once with the failure fed back, recorded as its own experiment. `search.max_patch_retries`
  (default 1; 0 disables).
- **Interleaved parent re-benchmark.** `benchmark.rebenchmark_parent: true` re-measures the parent
  next to each candidate so machine drift between iterations cannot bias the decision.
- **Beam search.** `search.beam_width` (or `--beam N`) keeps and expands the top-N accepted heads
  each iteration instead of only the best child. Width 1 is the original greedy search.
- **Planner outages and resume.** A timeout, rate limit, or 5xx from the planner is retried with
  backoff (`search.planner_retries`, `search.planner_retry_backoff_s`); a failure that persists skips
  that head for one iteration instead of ending the run. After 3 consecutive iterations with no
  successful plan the run fails. `hotpath run <config> --resume <run_id>` then continues it from the
  stored beam without re-measuring the baseline. Resume refuses a run whose test, benchmark, profile,
  editable/locked, execution, or benchmark settings changed, because new candidates would be compared
  against incomparable measurements.
- **Profile view with self-vs-total time.** The dashboard shows each hotspot's self time nested
  inside its cumulative time, so functions whose callees dominate are visible.
- **Bottleneck diff (before to after).** `baseline_profile` vs `head_profile`, aligned per function on
  absolute self seconds. Profiles are stored top-N (`profile.retain`, default 40), so a function
  missing from the later profile is reported as a *bound* - "at most 0.004s, below the top 40" -
  never as a proven elimination. A tool mismatch (e.g. `torch.profiler` vs its CPU-time fallback), a
  failed profile, or a head that is still the baseline are refused with the reason shown instead of a
  plausible-looking table. The `unchanged` band is labelled a display threshold, because a profile is
  one observation and has no measured noise floor. Also in `hotpath export`'s `REPORT.md`.
- **Metric-aware chart.** Toggle between `speedup x` and the raw metric in its own units, so a GPU
  target shows tokens/sec climbing rather than only a ratio. Raw is the default when higher is better.
  The acceptance **noise band** is shaded, which makes a `rejected_speed` verdict visible rather than
  something you read: a dot inside the band is not distinguishable from baseline. Axes are never
  inverted and never zero-forced; the baseline is drawn as a labelled reference. CI whiskers appear
  only in speedup mode, where that bootstrap interval is exactly what was computed.
- **Retries and beam branches in the tree.** A retry links back to the attempt it replaces with a
  dashed edge and sits next to it, so "failed -> fed the error back -> accepted" reads as one unit, and
  the detail panel shows the failure text the worker was actually given. Beam heads are labelled with
  the iterations they were expanded at, and rows are grouped so branches do not interleave.
- **`hotpath export <config> <dest>`** writes a PR-ready bundle: the optimized source tree,
  `changes.patch` (baseline→head diff), and `REPORT.md` with the benchmark table and accepted
  chain. Add `--ablate` to include a leave-one-out ablation table with freshly paired full-stack
  benchmarks. It prints a ready `gh pr create`
  command; it never pushes on its own.
- **Pruning.** `hotpath ablate --prune` (or `export --prune`) drops every change the ablation found
  removable, *together*, rebuilds the stack without them, re-runs the locked tests, and benchmarks it
  beside the full stack. The pruned stack is kept only if it is correct and the full stack is not
  measurably faster, because changes that are removable one at a time can still matter jointly. The
  run's head is never modified; `export --prune` ships the pruned stack and says what was dropped.
- **Multi-file changes.** A hypothesis may touch up to 3 files (`target_file` plus `extra_files`),
  including new ones, such as a Triton kernel module and its call site. The target must already exist;
  extra files may be created (an edit with an empty `search`). Every file must match an editable
  pattern and none may be locked; the harness checks each path before writing anything, and new files
  are registered with git so they appear in the diff and in isolated execution.
- **Accepted vs shipped.** The dashboard funnel and `REPORT.md` count "accepted" (correct and faster
  than its parent) separately from "shipped" (in the final head's lineage). With beam search a
  correct, faster change can still miss the final head.
- **Verification contract on the dashboard.** The locked test, benchmark, and profile commands are
  shown with known credentials filtered and a digest to compare runs.

## Roadmap still open

- Automatic ablation-driven pruning at the end of a run. Verified joint pruning is already
  available on request with `hotpath ablate --prune` or `hotpath export --prune`.
- Device-time call-stack attribution. Torch CPU event trees have a genuine flame graph; CUDA
  device-time and cProfile runs show a flat hotspot comparison because those collectors do not
  provide a reliable nested time hierarchy for that metric.
- Persist a profile per accepted experiment. `orchestrator` already profiles surviving beam nodes
  and discards the result, so any node could be diffed against baseline or its own parent for free.
