Metadata-Version: 2.4
Name: reprove
Version: 0.1.0
Summary: Verify a pull request with evidence, not opinions.
Author: Miguel Jardim
License: MIT
Project-URL: Homepage, https://github.com/ping-dev-ui/reprove
Classifier: Programming Language :: Python :: 3
Classifier: License :: OSI Approved :: MIT License
Classifier: Operating System :: OS Independent
Classifier: Topic :: Software Development :: Testing
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: coverage>=7.0
Provides-Extra: llm
Requires-Dist: anthropic; extra == "llm"
Provides-Extra: test
Requires-Dist: pytest; extra == "test"
Requires-Dist: pytest-cov; extra == "test"
Requires-Dist: ruff; extra == "test"
Dynamic: license-file

# reprove

**Verify a pull request with evidence, not opinions.**

`reprove` runs on a repository and a pull request and answers one question:
*what evidence exists that this change is safe to release, and what evidence
is missing?* It is built for the case where the change was written quickly,
often by an AI tool, and nobody has checked whether the tests actually bite.

Every finding carries a status:

| status | meaning |
|---|---|
| `reproduced` | reprove ran it and saw it |
| `inferred` | reasoned from the diff, not executed |
| `needs-evidence` | could not be checked here; the report says what would be needed |

reprove never claims more than it ran.

## What it does

1. **Diff.** `git diff base...head`. Each changed file is classified as source, test or config. Changed source lines are the unit everything else measures against.
2. **Tests.** Detects and runs the project's test command (`pytest` if `pyproject.toml`/`pytest.ini`/`tests/` exist; `npm test`, `vitest` or `jest` from `package.json`). Captures pass/fail counts, duration and the tail of the output, with a configurable timeout.
3. **Diff coverage.** Python: the suite runs once under `coverage` with per-test contexts, and executed lines are mapped onto the changed lines: *changed lines executed by tests: X of Y*, listing the uncovered ones. JS/TS: `c8` or `nyc` if the project already has one, otherwise `needs-evidence` with the exact command that would produce it.
4. **Mutation of changed lines only** (Python). One AST mutator at a time on the changed lines: comparison swap (`<`/`<=`, `==`/`!=`, `>`/`>=`), arithmetic swap (`+`/`-`, `*`/`/`), numeric boundary (`n` to `n+1`, `n-1`), boolean flip, negation removal, `and`/`or`. Each mutant is written into a temp copy of the repository and the tests that cover that line are run (the full suite when coverage is unavailable). Result: `killed` or `survived`, with the exact mutation and line. **A surviving mutant on a changed line is the headline finding: the tests do not detect that change.** Budgets: `--max-mutants` (default 50) and `--time-budget` seconds.
5. **Hypotheses (optional, bring your own key).** Only if `ANTHROPIC_API_KEY` is set and `--no-llm` is not passed: the diff, test summary and surviving mutants go to the model (default `claude-sonnet-5`, `--model` to change), which proposes a few specific risks, each with a pytest reproduction. reprove writes each reproduction into a temp copy and runs it. A hypothesis is shown **only with its run result**: `reproduced` (the test failed as predicted), `not-reproduced`, or `could-not-run`. Without a key this stage is skipped and the report says so in one line.
6. **Report.** Markdown sized for a PR comment, and JSON. A verdict line (`release-safe evidence: strong | partial | weak` with the specific gaps), findings ordered by severity (file:line, what was checked, the command run, an output excerpt, status), then the coverage table, the mutant table, the hypotheses table, and *what reprove could not check here*.

## Honesty rules

- reprove proves what it ran. It does not prove absence of bugs.
- Coverage of a line is not correctness of a line; a killed mutant is one specific change the tests noticed.
- Surviving mutants are the point: they are places where a change could ship unnoticed.
- The hypothesis stage is the only part that uses a model, it is off by default without a key, and every hypothesis is executed before it is shown.

## Install

Python 3.10 or newer. Install reprove into the **same environment as your project's test dependencies** (it runs `python -m pytest` with its own interpreter).

```bash
pip install reprove              # core: coverage only
pip install "reprove[llm]"       # + anthropic, for the optional hypothesis stage
```

From a checkout:

```bash
git clone https://github.com/ping-dev-ui/reprove
cd reprove
python -m venv .venv
# Windows: .venv\Scripts\activate    Linux/macOS: source .venv/bin/activate
pip install -e ".[llm,test]"
```

## Usage

```bash
reprove verify [--base REF] [--head REF] [--format md|json|both] [--out PATH]
               [--max-mutants N] [--time-budget S] [--test-timeout S]
               [--no-llm] [--model NAME] [--max-hypotheses N]
               [--fail-on weak|partial|never] [--repo PATH]
```

Defaults: base is `origin/main` if it exists, else `main`; head is `HEAD`; `--fail-on never`; Markdown to stdout.

```bash
# the last commit, no model
reprove verify --base HEAD~1 --head HEAD --no-llm

# a branch against main, both formats, fail the job on a weak verdict
reprove verify --base origin/main --format both --out reprove-report --fail-on weak
```

Exit codes: `0` ok, `1` the verdict hit `--fail-on`, `2` reprove could not run (not a git repository, unknown ref).

Tests, coverage and mutation run against the **working tree**. If the working tree is not at `--head`, or has uncommitted changes, the report says so under *what reprove could not check here*.

## GitHub Action

```yaml
name: reprove
on:
  pull_request:

permissions:
  contents: read
  pull-requests: write   # to create/update the sticky comment

jobs:
  verify:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
        with:
          ref: ${{ github.event.pull_request.head.sha }}   # run against the PR head, not the merge commit
          fetch-depth: 0
      - uses: actions/setup-python@v5
        with:
          python-version: "3.12"
      - run: pip install -e ".[test]"        # your project's test dependencies
      - uses: ping-dev-ui/reprove@main
        with:
          max-mutants: "50"
          time-budget: "300"
          fail-on: never                     # weak | partial | never
          no-llm: "false"
        env:
          ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}   # optional (BYOK)
```

The action installs reprove, runs `verify` with base `${{ github.event.pull_request.base.sha }}` and head `${{ github.event.pull_request.head.sha }}`, writes `reprove-report.md`/`.json`, uploads them as an artifact, and creates or updates **one sticky PR comment** (marker `<!-- reprove -->`) using the `gh` CLI with `GITHUB_TOKEN`. Inputs: `max-mutants`, `time-budget`, `fail-on`, `no-llm`.

**Bring your own key.** The hypothesis stage runs only if you add your own `ANTHROPIC_API_KEY` secret. reprove reads the key from the environment only; it never writes, logs or stores it, and never sends anything anywhere else. Without the secret the stage is skipped and the report says so.

## What it proves, and what it cannot

Proves (status `reproduced`):

- the test suite passes or fails on this working tree, with counts and the command;
- which changed lines the tests execute, and which they never touch;
- for each mutant on a changed line, whether the covering tests noticed it;
- for each model hypothesis, whether its reproduction actually fails.

Cannot prove (and says so):

- absence of bugs, or correctness of any line: only that specific changes were or were not noticed;
- anything the tests do not exercise, including code paths reachable only in production;
- whether a **test-only** change made the tests weaker: a diff with no changed source lines gets at most a `partial` verdict;
- mutation and hypothesis execution for JS/TS in v1 (detected, reported as `needs-evidence`);
- a mutant survives trivially when no test executes its line: the report labels these *no test executes this line*;
- a file that the tests import from somewhere other than the working tree (for example an installed copy of a `src/` package). reprove checks this with an import sentinel: it breaks the file in the temp copy first, and if the covering tests still pass, that file's mutants are reported as `needs-evidence` instead of `survived`.

Model-written reproductions are executed like any other test in your project, inside a temp copy: only enable the hypothesis stage on repositories whose tests you would run anyway.

## Verdict

- **strong**: tests pass, every executable changed line is executed, every mutant run was killed, nothing skipped, no reproduced hypothesis.
- **partial**: tests pass but something is missing: uncovered changed lines, a few survivors, mutants skipped by a budget, coverage unavailable, or a test-only change.
- **weak**: no test command, failing or timed-out tests, fewer than half the changed lines executed, at least half the mutants surviving, or a reproduced hypothesis.

The gaps are always listed next to the verdict.

## Developing

```bash
pip install -e ".[llm,test]"
python -m pytest -q          # ~40 s; builds real git repos under tmp_path
ruff check .
reprove verify --base HEAD~1 --head HEAD --no-llm   # run reprove on itself
```

Tests never call the network; the model is mocked. See `docs/demo-self.md` for reprove run on its own repository.

## License

MIT, copyright 2026 Miguel Jardim.
