Metadata-Version: 2.4
Name: todoscope
Version: 0.12.0
Summary: Find maintenance comments in source code and optionally interpret them with AI, without ever sending source code.
Keywords: todo,comments,scanner,cli,maintenance
Author: Zelmari
Author-email: Zelmari <240836003+Zelmari@users.noreply.github.com>
License-Expression: MIT
Classifier: Development Status :: 4 - Beta
Classifier: Environment :: Console
Classifier: Intended Audience :: Developers
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Software Development :: Quality Assurance
Classifier: Topic :: Utilities
Requires-Dist: openai>=3.0.0
Requires-Dist: pathspec>=1.1.1
Requires-Dist: tree-sitter>=0.25.2,<0.26
Requires-Dist: tree-sitter-c>=0.24.2
Requires-Dist: tree-sitter-c-sharp>=0.23.5
Requires-Dist: tree-sitter-cpp>=0.23.4
Requires-Dist: tree-sitter-go>=0.25.0
Requires-Dist: tree-sitter-java>=0.23.5
Requires-Dist: tree-sitter-javascript>=0.25.0
Requires-Dist: tree-sitter-rust>=0.24.2
Requires-Dist: tree-sitter-typescript>=0.23.2
Requires-Python: >=3.12
Project-URL: Repository, https://github.com/Zelmari/todoscope
Description-Content-Type: text/markdown

# TodoScope

TodoScope finds maintenance comments (`TODO`, `FIXME`, ...) in your code and
prints a clean report. Optionally, it asks an AI to interpret each comment and
estimate its priority — **without ever sending your source code anywhere**.

```bash
todoscope src/
```

## What it does

- Scans Python, JavaScript, TypeScript, JSX/TSX, Rust, Java, Go, C, C++,
  and C# files for comments that start with your markers (`TODO` by
  default). The default enabled set is `.py .js .jsx .ts .tsx .rs`; enable
  more extensions (`.java .go .c .h .cpp .cc .cxx .hpp .cs`) through
  `extensions` in `.todoscope.json`.
- Only real comments count: `TODO` inside strings, template literals, JSX
  text, or raw strings is ignored.
- Respects every `.gitignore` in the tree (root and nested, with git's
  override semantics) and an optional exclusion list.
- Works fully offline — the AI part is optional.
- Can show how long each finding's current line has been committed using Git
  history.
- When AI is on, it sends only each comment's ID, marker, and text. No file
  names, no paths, no line numbers, no code.

## Install

Requires Python 3.12+.

```bash
pipx install todoscope        # recommended
# or
uv tool install todoscope     # if you use uv
# or
python3 -m pip install todoscope
```

## Use

```bash
todoscope src/                # scan a folder recursively (local only)
todoscope src/main.py         # scan one file
todoscope .                   # scan the whole project
todoscope src/ --ai           # add AI interpretations and priorities
todoscope src/ --blame        # add who-authored-each-finding via git blame
todoscope src/ --age          # add time since each finding was committed
todoscope src/ --age --blame  # show both age and attribution
todoscope src/ --quiet        # one numbered finding per line, nothing else
todoscope src/ --verbose      # extra details on stderr
todoscope src/ --format json  # machine-readable JSON report on stdout
todoscope src/ --format sarif # SARIF 2.1.0 report for code-scanning tools
```

That's it. Findings are sorted by folder depth, then path, then line, and
every text mode uses the same canonical line:

```text
1. src/auth/session.py:84: TODO: Handle expired refresh tokens
```

Scanning is local by default — `--ai` is opt-in and never runs when
`--quiet` is given (the combination prints a note and behaves like plain
`--quiet`). `--blame` requires a Git repository and adds one attribution line
per finding. `--age` also requires Git and shows the number of days since the
finding's current marker line was committed. Uncommitted lines are identified
as such, while unavailable history is reported without failing the scan. Both
options are rejected with `--quiet`; when combined, they share a single
`git blame --porcelain` call per file. Git history data never reaches the AI.

`--format json` prints a deterministic JSON document to stdout (scan
metadata, findings, skipped counts, optional blame and age data, and the AI
section with a machine-readable status/reason). Age entries include a status,
an exact day count, and the commit date; uncommitted or unavailable entries
use `null` for values that do not apply. Without `--ai`, the AI section is
`null`. `--format sarif` prints a deterministic SARIF 2.1.0 document with one
rule per configured marker; AI priorities map to SARIF levels (High →
`error`, Medium → `warning`, Low/Unclear → `note`), and blame and age data
are attached as result properties when requested. Verbose details and errors
always go to stderr. Neither format ever contains API keys or environment
values.

## Configuration

Everything optional lives in a `.todoscope.json` in your project root:

```json
{
  "markers": ["TODO", "FIXME"],
  "extensions": [".py", ".js", ".jsx", ".ts", ".tsx", ".rs"],
  "exclude": ["tests/fixtures/", "generated/"],
  "model": "your-ai-model-id",
  "max_ai_characters": 20000
}
```

| Key | What it does |
|---|---|
| `markers` | Replaces the default marker list (`["TODO"]`). Matching is case-sensitive and prefix-based; the longest matching marker wins. |
| `extensions` | Replaces the default scanned extensions. |
| `exclude` | Skips exact project-root-relative paths or directory prefixes. |
| `model` | Required for AI analysis. There is **no default model**. |
| `max_ai_characters` | Lower AI payload limit (hard ceiling: 100,000). |

Invalid configuration stops with a clear error (exit code 3).

## AI analysis

AI is opt-in: pass `--ai` to request it. To enable it you need both:

1. An API key — from your shell (`TODOSCOPE_API_KEY`) or a `.env` file in the
   project root:

   ```dotenv
   TODOSCOPE_API_KEY=...
   TODOSCOPE_SECONDARY_API_KEY=...
   ```

   Shell values win over `.env`. If a key comes from `.env`, that file must
   be ignored by your `.gitignore`, otherwise AI is refused for safety.

2. A `model` in `.todoscope.json`.

When enabled, TodoScope makes **one** request and then prints one complete
report: per finding you get a short interpretation and an estimated priority
(High / Medium / Low / Unclear), plus an overall summary. If the request
fails and a secondary key is configured, an interactive terminal offers one
retry with it — the secondary key is never used silently.

> Priorities are estimated from comment text only. No source code was
> provided to the AI.

Before any request, comment text is screened for likely credentials (API
keys, tokens, private-key headers, credential assignments). If any finding
looks like a secret, the AI request is refused and the suspicious findings
are listed locally — the local report is unaffected. Detection is
conservative: it flags unambiguous secret shapes, never prose.

### Using DeepSeek (or another OpenAI-compatible provider)

The OpenAI SDK reads `OPENAI_BASE_URL` from your environment. For DeepSeek:

```bash
export OPENAI_BASE_URL=https://api.deepseek.com
todoscope .
```

or as a permanent alias in `~/.zshrc`:

```zsh
alias todoscope="OPENAI_BASE_URL=https://api.deepseek.com /home/$USER/.local/bin/todoscope"
```

## Privacy

The only data from your repository that reaches the AI is each finding's ID,
marker, and extracted comment text. Everything else stays local. Before any
request, comment text is screened for likely credentials and the request is
refused if any are found. Comments are
treated as untrusted data — instructions written inside a comment can never
change TodoScope's behaviour. **Never put credentials or secrets in code
comments**, because comment text may be sent to the AI.

## Exit codes

- `0` — scan finished (including local-only results after any AI problem)
- `1` — unexpected failure
- `2` — bad path/usage, or an ignored target refused in non-interactive mode
- `3` — configuration error

## Use in CI

TodoScope is CI-friendly: finding TODOs is **not** an error, so scans never
fail a pipeline just because comments exist. Common patterns:

- Log findings: `todoscope . --quiet` (one line per finding).
- Machine-readable reports: `todoscope . --format json` and upload or parse
  the JSON in later steps.
- Code-scanning alerts: `todoscope . --format sarif > todoscope.sarif` and
  upload the file with `github/codeql-action/upload-sarif` (or another
  SARIF consumer) to surface findings as alerts in the Security tab.
- AI in CI: set `TODOSCOPE_API_KEY` as a repository secret and a `model` in
  `.todoscope.json`; non-interactive runs skip the secondary key safely.

Ready-made examples live in [`examples/ci/`](examples/ci/):

- `scan-pr.yml` — scan on pull requests, print findings, upload the JSON
  report as an artifact.
- `scan-quiet.yml` — minimal log-only variant.

## Development

```bash
uv sync                       # set up the environment
uv run pytest                 # tests
uv run ruff check .           # lint
uv run ruff format --check .  # format check
uv build                      # wheel + sdist
```

Continuous integration runs these same checks on every push and pull
request (Python 3.12 and 3.13).

## Releasing

1. Bump `version` in `pyproject.toml` (minor for features, patch for fixes).
2. Add a `CHANGELOG.md` entry for the new version.
3. Commit, then tag and push the tag:

```bash
git tag vX.Y.Z          # e.g. git tag v0.8.2
git push
git push --tags
```

The publish workflow verifies everything, uploads to PyPI using the
`PYPI_TOKEN` repository secret, and creates a GitHub release automatically.

To re-publish an older tag (for example, backfilling a version that never
made it to PyPI), run the workflow manually: Actions → Publish → Run
workflow, and set the `ref` input to the tag name (e.g. `v0.5.0`).

## Changelog

See [CHANGELOG.md](CHANGELOG.md).
