Metadata-Version: 2.4
Name: cleanvibe
Version: 1.18.0
Summary: A workflow accelerator that scaffolds AI-assisted coding projects with opinionated documentation and launches Claude Code.
Author: Immanuelle
License-Expression: MIT
Project-URL: Homepage, https://github.com/EmmaLeonhart/cleanvibe
Project-URL: Issues, https://github.com/EmmaLeonhart/cleanvibe/issues
Keywords: claude,ai,scaffold,workflow,vibe-coding
Classifier: Development Status :: 3 - Alpha
Classifier: Environment :: Console
Classifier: Intended Audience :: Developers
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Software Development
Classifier: Topic :: Utilities
Requires-Python: >=3.9
Description-Content-Type: text/markdown
License-File: LICENSE
Dynamic: license-file

# cleanvibe

**Website · [cleanvibe.emmaleonhart.com](https://cleanvibe.emmaleonhart.com)**

A tiny Python CLI that scaffolds AI-assisted coding projects and launches Claude Code.

`cleanvibe` is not a coding tool. It's a **state initializer** -- it removes the friction between "I want to build something" and "Claude is working inside a well-structured environment." The real value is the working contract it installs: a short `CLAUDE.md` plus six workflow skills in `.claude/skills/` that enforce documentation discipline, meaningful commits, and iterative file-based thinking. Repos are private by default, and Claude launches with a first message explaining which mode it is in.

## Install

```
pip install cleanvibe
```

### Developer install (working on cleanvibe itself)

```
git clone https://github.com/EmmaLeonhart/cleanvibe
cd cleanvibe
pip install .         # or use !dev-install.bat on Windows
```

**Use `pip install .`, not `pip install -e .`.** The repository directory is itself named `cleanvibe`, which means an editable install collides with Python's namespace-package CWD scanning when you run `python` from the repo's *parent* directory — Python finds the repo dir as a namespace package and beats the editable finder, so `import cleanvibe` returns a module with no `__version__`. The non-editable install copies the package into site-packages where it always wins. After any source edit, run `pip install .` again (or `!dev-install.bat`). The console-script entry point (`cleanvibe` on PATH) works correctly under both install modes — this quirk only affects programmatic `import cleanvibe`.

## Usage

### Create a new project

```
cleanvibe new my-project
```

This will:
1. Create the directory `my-project/`
2. Write `CLAUDE.md` (a short pointer to the skills + project-specific notes)
3. Write `README.md` (starter documentation)
4. Write `queue.md` (active work queue, pre-seeded with a first-session bootstrap sequence that walks Claude through triaging dropped-in files, inferring the project, interviewing the user, creating `todo.md`, populating the real queue, and pushing to a private GitHub repo)
5. Write `devlog.md` (where "done" lives) and `.gitignore` (sensible Python defaults)
6. Create `data_lake/` (drop files in before the first session) and, on Windows, `!runClaude.bat`
7. Vendor `.claude/skills/` (the six workflow skills — see below)
8. Initialize a git repo on `main` with an initial commit
9. Launch Claude Code inside the project, with the `new` starting prompt

### Skills (v1.14.0+)

The workflow behaviors that used to be inlined into `CLAUDE.md` now ship as six
standalone **skills**, auto-discovered by Claude Code from `.claude/skills/`:

| Skill | Fires when |
|---|---|
| `emergency-stop` | you repeatedly say "stop" / demand an immediate halt |
| `cron-is-local` | you mention "cron" / "schedule" (means local `CronCreate`) |
| `autonomous-loop` | starting extensive autonomous work (the three-cron playbook) |
| `queue-driven-workflow` | any multi-step work (plan into `queue.md` first; the `todo`→`queue`→`devlog` flow) |
| `writing-style` | writing any prose (avoid the "honest"/"frank" tic) |
| `cleanvibe-update-check` | session start, weekly (refresh skills from cleanvibe) |

They're vendored into every `new` / `convert` / `clone` / `research` / `original` /
`chat` project (not `replicate`, which is a bounded workflow) and
kept current by the `cleanvibe-update-check` skill (which reads
<https://cleanvibe.emmaleonhart.com/updates.md>). The single source of truth is
`cleanvibe/skills.py`; `CLAUDE.md` keeps only a short `## Skills` pointer. To
back-fill these into an existing repo, run `migrate_repos_to_skills.py`.

### Research a question — your own investigation

```
cleanvibe research reservoiragent
cleanvibe research reservoiragent --question "What is the memory capacity of a reservoir-computing agent?"
cleanvibe new reservoiragent --research      # equivalent alias
```

`research` is for an **original-research project** — *your own* investigation,
not a [replication](#replicate-a-paper) of someone else's paper. It is `new`
plus the two things that make research legible: an up-front **literature
review** and a **published, themed report**. It scaffolds everything `new`
does (`CLAUDE.md`, `README.md`, `queue.md`, `devlog.md`, `.gitignore`,
`data_lake/`, the three-cron playbook) and adds:

- **`literature/`** — the literature review, built *before* any code. The
  bootstrap queue's distinctive step uses agentic RAG (web search, `WebFetch`,
  the `deep-research` skill if present) to survey prior work, collect sources
  with citations, and synthesize `literature/REVIEW.md` (what's known, the
  gaps, what *this* project adds). This grounds the project in the field
  instead of reinventing it — and is what separates `research` from `new`.
- **`docs/`** — a **published GitHub Pages report site**, pre-styled with a
  warm "paper" light theme + dark-mode variant (the look of
  [latent-space.emmaleonhart.com](http://latent-space.emmaleonhart.com/)), plus
  a transportable PDF built from `FINDINGS.md`. `.github/workflows/pages.yml`
  deploys it. The agent edits the content; the theme stays.

The bootstrap sequence is **literature-review-first**: start the crons →
triage `data_lake/` → **define the research question with you** → **literature
review (agentic RAG)** → write the long-horizon `todo.md` → push to a
**private** GitHub repo (going public for Pages is your call) → replace the bootstrap queue with the real
experiment/build queue → work it, keeping `FINDINGS.md` + the `docs/` report
current. Pass `--question` if you already know the question; otherwise the
bootstrap pins it down with you.

### Original research — when you don't have a topic yet

```
cleanvibe original driftprobe
cleanvibe original driftprobe --area "reservoir computing"
cleanvibe new driftprobe --original                          # equivalent alias
```

`original` is `research` for an **uncertain topic**: you don't yet have a fixed
research question. It keeps everything `research` has — `literature/`,
`data_lake/`, the three-cron playbook, the themed `docs/` report — and prepends
one distinctive bootstrap step:

- **`topics/`** — the **topic-finding loop**, run *before* the literature review.
  The bootstrap explores the focus area (agentic search / RAG), drafts a slate of
  candidate research questions, scores them (novelty, tractability, interest,
  available data/compute, what a result is worth), confirms the shortlist with
  you, and converges on ONE — recording the candidates + scoring + the chosen
  question + rationale in `topics/TOPICS.md`. Then it proceeds exactly like
  `research`.

The seed is `--area` (a field to explore), **not** `--question` — the question is
what the loop discovers. The bootstrap sequence is **topic-finding-first**: start
the crons → triage `data_lake/` → **topic-finding loop (pick the question)** →
**literature review (agentic RAG)** → write `todo.md` → push to a **private** repo → replace
the bootstrap queue → work it. Use `original` when you want to investigate *some*
area but haven't settled on the precise question; use [`research`](#research-a-question--your-own-investigation)
when you already know what you're asking.

### Chat — a git-tracked conversation

```
cleanvibe chat                                   # -> chat-YYYY-MM-DD/
cleanvibe chat tea-notes --topic "oolong vs pu-erh"
```

`chat` is for a **conversation about one topic** rather than a software project:
research-heavy, light on code, kept in a **private** git repo so you can resume,
search and share it. The session opens by asking you what you are trying to do
(AskUserQuestion) before it plans or researches anything. Conclusions and
sources go into `notes/`, and the README keeps a running "where things stand".

**Session logs are git-tracked.** The scaffold's `.claude/settings.json` runs a
small stdlib script (`.claude/hooks/save_session_log.py`) after every response
and at session end. It copies the transcript into `sessions/` as raw `.jsonl`
plus a readable `.md`, and commits only `sessions/`. At session end it also
pushes if the repo has a remote. Transcripts contain everything in the session,
including tool output, which is one reason the repo stays private.

**It starts with Remote Control on** (`claude "<prompt>" --remote-control`,
unnamed), so you can pick the conversation up from the Claude app or web.
`!runClaude.bat` does the same. NAME is optional; without one you get
`chat-YYYY-MM-DD` in the current directory, auto-suffixed `-2`/`-3` if it exists. Chat mode has no three-cron
playbook and no Pages report.

### Every mode: private repo, starting prompt

- **Private by default.** Every mode that creates a GitHub repo creates it with
  `gh repo create --private`. Going public is your call. On a private repo the
  research/replication Pages workflows upload the report as a workflow artifact
  instead of deploying (GitHub's free plan cannot publish Pages from a private
  repo). Make the repo public, or set the repo variable `CLEANVIBE_PAGES=true` on
  a paid plan, to deploy the site.
- **Starting prompt.** Claude launches with a first message saying the project
  was started with cleanvibe, which mode, what that mode is for, and to work
  `queue.md` item 1, asking you first if the goal is unclear. `!runClaude.bat`
  relaunches with the same prompt.

### Doctor — audit a project for drift

```
cleanvibe doctor            # audit the current directory
cleanvibe doctor path/to/project
```

A **read-only** check of a cleanvibe project for the drift that builds up over
time. It changes nothing, and exits `1` if it finds anything (so it can run in CI):

| Check | Flags |
|---|---|
| `files` | a missing `CLAUDE.md`, `README.md`, `queue.md`, or `devlog.md` |
| `skills` | a vendored skill that is missing or differs from this cleanvibe's copy |
| `queue-done` | ticked boxes, check marks, `DONE`, or strikethrough left in `queue.md` |
| `version` | `queue.md`'s "Current version" not matching `pyproject.toml` |
| `devlog-tags` | a `v*` git tag with no `devlog.md` entry |
| `section-refs` | a reference to a `CLAUDE.md` section heading that doesn't exist |
| `ci` | a `tests/` directory with no GitHub Actions workflow |
| `pages-gate` | a pre-v1.18.0 Pages workflow that fails on a private repo |

### Clone an existing repo — codebase onboarding

```
cleanvibe clone https://github.com/user/repo
```

`clone` is for **onboarding an existing codebase**, not bootstrapping a blank
one. It is deliberately different from `new`:

1. `git clone` the repository
2. Create and check out a dedicated `cleanvibe-onboarding` branch — **the
   default branch is left untouched**
3. *Prepend-or-write* an onboarding `CLAUDE.md` and `queue.md`: if the repo
   already has them, the fresh block goes on top (newest first) and the
   original content is preserved below — re-running just layers another block
4. Inject `.gitignore` only if missing. **No `data_lake/`** (it is a real
   codebase, nothing was dropped in) and **no README overwrite**
5. Commit the onboarding scaffold on the branch
6. Launch Claude Code inside the project

The onboarding `queue.md` is small and focused: read & document the repo,
make existing docs accurate, **rewrite `CLAUDE.md` to the repo's real
development practices**, add tests/CI if sparse, then synthesize any existing
planning artifacts and hand off to the repo's own `todo.md`.

### Replicate a paper

`cleanvibe replicate` takes a **clawRxiv** reference, an **arXiv/alphaxiv**
reference, a **plain URL** to non-arXiv research, **or** a folder name:

**From a clawRxiv paper (skill-first):**

```
cleanvibe replicate https://www.clawrxiv.io/abs/2605.02609
cleanvibe replicate clawrxiv:2605.02609
```

[clawRxiv](https://www.clawrxiv.io/) publishes papers authored autonomously by
AI agents and exposes a JSON API (`/api/abs/<id>`) that **differentiates the
paper content, abstract, and skill file** (an agent-runnable replication
recipe). That separation is the purest recipe-first case, so clawRxiv gets its
own dedicated mode. The scaffold fetches all three up front: the paper content
is written **locally** to `replication_target/source/paper.md` (gitignored —
the paper is copyrighted and is **never committed**), and when clawRxiv ships a
separate skill file it lands at `replication_skill.md` at the root (otherwise
the recipe is embedded in the content and the queue tells the agent to extract
it). A `download_paper.py` re-fetches the content from the clawRxiv API if
`replication_target/` is ever empty (e.g. a fresh clone). The generated
`queue.md`/`SKILL.md` are
**skill-first**: go live early, run the recipe, verify it against the paper,
check all references, then fill only the gaps. clawRxiv ids look arXiv-shaped,
so a bare id stays arXiv — use a `clawrxiv.io` URL or `clawrxiv:<id>` to select
clawRxiv mode.

**From an arXiv / alphaxiv paper:**

```
cleanvibe replicate https://arxiv.org/abs/1706.03762
cleanvibe replicate https://www.alphaxiv.org/overview/2201.02177
cleanvibe replicate https://doi.org/10.48550/arXiv.1706.03762
cleanvibe replicate 1706.03762v5
```

Any arXiv/alphaxiv id or URL is accepted — `/abs/`, `/pdf/`, `/html/`,
`/src/`, alphaxiv's primary `/overview/`, `/audio/`, `/forum/`, the arXiv
**DOI** form (`doi.org/10.48550/arXiv.<id>`), `arXiv:<id>` citation style,
trailing slugs and query strings all resolve. A pinned `vN` **version** is
preserved (recorded in `paper.json` and used for the download), not silently
dropped. This will:
1. Fetch the paper's metadata from the arXiv API (with **429-aware
   retry/backoff** — arXiv rate-limits, so requests honour `Retry-After`
   and back off rather than crashing)
2. Create `replicating-<paper-slug>/` (silently `-2`/`-3` if it already exists)
3. Scaffold a standalone replication project: cleanvibe conventions
   (`CLAUDE.md`, `queue.md`, `data_lake/`) **plus** the replication structure —
   `SKILL.md` (the agent-executable replication plan), `download_paper.py`
   (fetches the arXiv **LaTeX/e-print source** and extracts it to
   `replication_target/source/`, with the PDF as a fallback), the paper's home
   `replication_target/` (gitignored — never in `data_lake/`; the authors' code
   is cloned here as a git submodule), `paper.json`, and `.github/workflows/`
   that build a GitHub Pages findings site, a transportable PDF report, and a
   downloadable ZIP replication package
4. Initialize a git repo with an initial commit
5. Launch Claude Code inside the project

The generated scaffold is built around the **efficient, recipe-first path**:

- **Source, not HTML.** `download_paper.py` downloads the arXiv **e-print
  source** (`arxiv.org/src/<id>`) and extracts the `.tex` to
  `replication_target/source/`. The `.tex` is far more token-efficient than the
  rendered HTML, which embeds figures as huge base64 data-URIs you'd otherwise
  have to strip. **The paper is never committed:** the whole `replication_target/`
  tree is gitignored (papers are copyrighted), so the download is local context
  only — run `python download_paper.py` to (re)populate it whenever it's empty.
  The `cleanvibe replicate` command runs this download itself **before launching
  Claude**, so the agent opens onto an already-extracted paper — but nothing
  under `replication_target/` ever enters a commit.
- **Consent before running code.** Because a replication runs code you didn't
  write (the recipe / cloned scripts / a downloaded zip), the generated
  `queue.md`'s **first step** makes the agent stop and get your explicit consent
  before executing any external/cloned code. Reading the paper, source, and
  recipe is fine; *running* third-party code is the gated action.
- **Find the recipe FIRST.** Authors very often ship a reproduction recipe
  right in the paper source (usually near the end): a `SKILL.md`/`AGENTS.md`, a
  `reproduce.*`/`replicate.*`/`run.sh` script, a Makefile target, a Dockerfile,
  or a downloadable **replication zip**. The generated `queue.md`/`SKILL.md`
  tell the agent to find it (copying a recipe to `replication_skill.md`,
  extracting a zip into `replication/`) and **run it first**, *before* any deep
  paper analysis — then verify its output against the paper, check **all** the
  paper's references, and only reimplement the gaps the recipe didn't cover.
- **Go live early.** The agent is told to create a PRIVATE GitHub repo and push
  near the start, so every commit pushes and CI builds as the work goes — not
  left local-only. Every mode defaults to private; on a private repo the Pages
  workflow uploads the report as a workflow artifact instead of deploying it,
  until you make the repo public.
- **Themed report with a status badge.** The GitHub Pages findings site is
  rendered with the **shared cleanvibe report theme** (`report-theme.css` — the
  same warm "paper" + dark-mode theme `cleanvibe research` uses) and topped with
  a big color-coded **replication status badge** — 🟢 replicated / 🔴 failed /
  🟠 insufficient hardware / 🔵 in progress — driven by a `status` field in
  `paper.json` (defaults to in-progress). A transportable PDF is built too.

**From a plain URL (research that isn't on arXiv):**

```
cleanvibe replicate https://some-lab.org/papers/cool-thing.pdf
cleanvibe replicate https://openreview.net/forum?id=XXXX
```

When the argument is a plain `http(s)` URL that isn't an arXiv/clawRxiv
reference, cleanvibe **downloads it** as the replication source — the page or
PDF lands **locally** in `replication_target/source/` (`paper.pdf` or
`paper.html`, detected automatically) — gitignored, **never committed** — and
provenance is recorded in `source.json`. A `download_paper.py` re-downloads
from that recorded URL if `replication_target/` is ever empty. Same 429-aware
retry/backoff as arXiv mode. Use it for research hosted on lab sites,
OpenReview, journal pages, or anywhere that isn't arXiv/clawRxiv.

**From a folder you fill yourself (manual drop-in mode):**

```
cleanvibe replicate my-paper-replication
```

When the argument is **not** an arXiv/alphaxiv reference (and not a URL) it is
treated as a folder name and a *manual drop-in* project is scaffolded — no
metadata
fetch, no `download_paper.py`, no `paper.json`, no network. You drop the
paper PDF(s) into `replication_target/` and any datasets/notes into
`data_lake/` yourself; the scaffolded `CLAUDE.md` / `queue.md` / `SKILL.md`
/ `README.md` say so up front, and the first queue step makes the agent
**stop and ask you for the paper** if `replication_target/` is empty rather
than invent one. Injection is non-destructive: you can create the folder,
drop your PDF in, *then* run `cleanvibe replicate ./that-folder` — nothing
you put there is overwritten.

Every replication produces three compounding artifacts: the runnable
replication, a published findings report, and the reusable `SKILL.md`
methodology. See `docs/replication_framing.md` for the full vision.

### Options

```
cleanvibe new my-project --dry-run        # Preview what would be created
cleanvibe new my-project --no-claude      # Skip launching Claude Code
cleanvibe research my-study --dry-run     # Preview the research scaffold
cleanvibe research my-study --no-claude   # Scaffold a research project without launching Claude
cleanvibe chat --dry-run                  # Preview the chat scaffold
cleanvibe doctor                          # Audit the current project for drift (read-only)
cleanvibe clone REPO path --dry-run       # Preview what would be done
cleanvibe replicate URL --dry-run         # Preview the arXiv replication scaffold
cleanvibe replicate FOLDER --dry-run      # Preview the manual drop-in scaffold
cleanvibe replicate URL --no-claude       # Scaffold without launching Claude
cleanvibe --version                       # Show version
```

## Why?

Most people struggle with blank repo paralysis, poor commit hygiene, and AI assistants that ramble without producing durable artifacts. `cleanvibe` solves this by injecting a disciplined thinking contract into every project from the start.

The `CLAUDE.md` template enforces:
- Commit early and often with meaningful messages
- No planning-only modes -- all thinking produces files and commits
- Keep documentation up to date as the project evolves
- Use `planning/` directories for exploration instead of internal planning modes

## Cross-platform

Works on Windows, Linux, and macOS. Zero dependencies beyond Python 3.9+.

## Website

Full walkthrough — what cleanvibe is and what each subcommand does — at the
project site (built from `pages/` and deployed by GitHub Actions):
**https://cleanvibe.emmaleonhart.com/**

## Stability

As of **v1.0.0**, cleanvibe commits to the following contract (semantic
versioning from here on):

- **Subcommands** `new`, `research`, `original`, `chat`, `clone`, `convert`,
  `replicate`, and `doctor` are stable. Their core behavior will not change incompatibly
  within the 1.x line.
- **Injected files**: `new` guarantees `CLAUDE.md`, `README.md`, `queue.md`,
  `.gitignore`, and `data_lake/.gitkeep`. `research` guarantees all of those
  **plus** `literature/.gitkeep`, `docs/index.html` (the themed report site),
  and `.github/workflows/pages.yml`. `chat` guarantees `CLAUDE.md`, `README.md`,
  `queue.md`, `devlog.md`, `.gitignore`, `notes/`, `sessions/`, `data_lake/`,
  `.claude/settings.json` and `.claude/hooks/save_session_log.py`. `replicate` always guarantees
  `SKILL.md`, `CLAUDE.md`, `queue.md`, and a **gitignored** `replication_target/`
  (the paper lives here, local-only, and is **never committed** — papers are
  copyrighted); in arXiv mode it additionally guarantees `paper.json` and
  `download_paper.py`; in clawRxiv mode it guarantees `paper.json`, a
  `download_paper.py` (re-fetches the content from the clawRxiv API), and a
  local `replication_target/source/paper.md` (plus `replication_skill.md` when
  clawRxiv ships a separate skill file); URL mode guarantees `source.json` and a
  `download_paper.py` (re-downloads from the recorded URL). `download_paper.py`
  is absent only in manual drop-in mode — you supply the paper by hand, so there
  is nothing to fetch.
- **Non-destructive by contract**: `clone` and `convert` never overwrite
  existing files — `clone` prepends; `convert` only injects what is missing.
  `replicate` in arXiv mode never errors on a name collision (silent
  `-2`/`-3` suffix); in folder mode it injects only what is missing so a
  pre-dropped paper is never clobbered.
- **Template wording** may evolve (improvements to the workflow contract are
  not breaking); the *set* of guaranteed files and the subcommand contracts
  above are what 1.x holds stable.
- **Zero runtime dependencies** remains a hard guarantee for the 1.x line.

## License

MIT
