Metadata-Version: 2.5
Name: pdfredactcli
Version: 0.2.1
Summary: Redact text in a PDF (true redaction, not just visual overlay) using PyMuPDF
Project-URL: Homepage, https://github.com/alberto743/pdfredact
Project-URL: Repository, https://github.com/alberto743/pdfredact
Author: Alberto P.
License-Expression: MPL-2.0
License-File: COPYING
Classifier: Environment :: Console
Classifier: Intended Audience :: Developers
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Topic :: Security
Classifier: Topic :: Text Processing
Requires-Python: >=3.9
Requires-Dist: pymupdf>=1.24
Requires-Dist: pyyaml>=6
Provides-Extra: test
Requires-Dist: pytest>=8; extra == 'test'
Description-Content-Type: text/markdown

# pdfredact

*[Leggi questo in italiano](README.it.md)*

Redact text in a PDF (true redaction, not just a visual overlay) using
[PyMuPDF](https://pymupdf.readthedocs.io/). Finds occurrences of the specified text/pattern,
applies a redaction annotation, and "burns" it into the page content, physically removing the
underlying text (not recoverable via copy-paste or text extraction).

## Installation

Requires Python 3.9 or later. Dependencies are PyMuPDF and PyYAML (for `--config`), both of
which publish prebuilt wheels for Linux, Windows, and macOS (no compiler required).

The package is published on PyPI as [`pdfredactcli`](https://pypi.org/project/pdfredactcli/)
(the installed command is `pdfredact`).

### With pip

```sh
pip install pdfredactcli
```

or, from a local clone of the repository:

```sh
pip install .
```

or, for development (editable install with test dependencies):

```sh
pip install -e .[test]
```

### With pipx (recommended for a command-line tool)

[pipx](https://pipx.pypa.io/stable/) installs the tool in an isolated virtual environment and
exposes only the `pdfredact` command on `PATH`, without touching the system Python.

**Linux/macOS:**

```sh
python3 -m pip install --user pipx
python3 -m pipx ensurepath
pipx install pdfredactcli
```

Or, from a local clone of the repository, replace the last line with `pipx install .` (run
from the root of the repository).

**Windows:**

On Windows it's convenient to install `pipx` via [Scoop](https://scoop.sh/), which also
manages updating Python itself if needed:

```powershell
# If Scoop isn't already installed:
Set-ExecutionPolicy RemoteSigned -Scope CurrentUser
Invoke-RestMethod -Uri https://get.scoop.sh | Invoke-Expression

scoop install pipx
pipx ensurepath
```

Then, in a new terminal (so `ensurepath` takes effect):

```powershell
pipx install pdfredactcli
```

Or, from a local clone of the repository, run `pipx install .` from the repository folder
instead.

In both cases, after installation the `pdfredact` command is available directly in a new
terminal.

## Usage

```sh
pdfredact input.pdf output.pdf -t "Mario Rossi" -t "CF: ABCDEF"
pdfredact input.pdf output.pdf -r "\bMCNP-\d{4}\b"
pdfredact input.pdf output.pdf -t "Confidential" --case-sensitive
pdfredact input.pdf output.pdf -t "Mario" --no-whole-word  # also matches "Mariotti"
pdfredact input.pdf output.pdf -t "foo" --pages 1,2,5-7
pdfredact input.pdf output.pdf --box "1:56,700,300,730"
pdfredact input.pdf output.pdf -t "foo" --fill-color "#ff0000"
pdfredact input.pdf -t "Mario Rossi"              # writes input_redacted.pdf
pdfredact --config job.yaml
pdfredact input.pdf output.pdf --config rules.yaml -t "extra one-off term"
pdfredact --version
```

The output path is optional: if omitted, it defaults to `<input>_redacted.pdf` next to the
input file.

Equivalent without installing, from the root of the repository:

```sh
python -m pdfredact input.pdf output.pdf -t "Mario Rossi"
```

### Whole-word matching

By default, `-t/--text` terms only match on word boundaries: `-t "Mario"` redacts a standalone
"Mario" but leaves "Mariotti" or "mario2" untouched. This only constrains the *edges* of a term
that are themselves letters/digits, so a term ending in punctuation (e.g. `"Confidential:"`)
still correctly matches even though nothing changes on that side - it just won't match inside
"NonConfidential:" either. Use `--no-whole-word` (or `whole_word: false` in `--config`) to go
back to plain substring matching. This only affects `-t`; `-r/--regex` patterns are never
touched, since you already have full manual control there via your own `\b`.

Word boundaries are decided from where the characters actually sit on the page, not just from
the extracted text: many PDFs separate words (table columns, form fields, right-aligned values)
by moving the text cursor instead of writing a space, so a visible gap counts as a boundary
even when there is no whitespace character between the two words.

### Rectangle coordinates (`--box`)

Format: `PAGE:x0,y0,x1,y1`

- `PAGE` is 1-based (page 1 = first page)
- `x0,y0,x1,y1` in PDF points (72 pt = 1 inch), origin at the top-left (same coordinate
  system returned by `page.search_for()`)
- Corner order doesn't matter: the rectangle is normalized.
- A box that falls entirely outside the page redacts nothing: it's reported on stderr and
  not counted as a redacted occurrence, so a typo'd coordinate can't look like a success.

### Config file (`--config`)

Any option can also be set in a YAML file instead of retyped on every invocation:

```yaml
# job.yaml
input: input.pdf              # optional if given positionally on the command line
output: output.pdf            # optional; defaults to <input>_redacted.pdf if unset everywhere

text:                         # literal terms to redact (like repeated -t)
  - "Mario Rossi"
  - "CF: ABCDEF"

regex:                        # regex patterns to redact (like repeated -r)
  - '\bMCNP-\d{4}\b'

boxes:                        # explicit rectangles, same "PAGE:x0,y0,x1,y1" format as --box
  - "1:56,700,300,730"

case_sensitive: false          # like --case-sensitive
whole_word: true               # like --whole-word
pages: "1,2,5-7"               # like --pages
fill_color: "#000000"          # like --fill-color; keep the quotes, an unquoted
                               # '#' would start a YAML comment
```

All keys are optional, and `pdfredact --config job.yaml` alone is a valid invocation if `input`
is set in the file. Values from the config file and the command line are merged:

- `text`, `regex`, and `boxes` from the command line are **added** to the config file's lists.
- `input`, `output`, `pages`, `fill_color`, `case-sensitive`, and `whole-word` from the command
  line **override** the config file's value when explicitly passed. To override a config file's
  `case_sensitive: true` back to `false`, pass `--no-case-sensitive` (plain `--case-sensitive`
  can only set it to `true`); likewise `--no-whole-word` overrides a config file's
  `whole_word: true` back to `false`.

Each key is validated the same way as its CLI equivalent (same `--box`/`--pages`/`--fill-color`
formats); an unknown key or a value of the wrong type/shape fails immediately with exit code 2
rather than being silently ignored.

### Exit codes

`0` = success, `2` = input/usage error.

## Known limitations

- Only PDF input is accepted. PyMuPDF can also open TXT, EPUB, SVG, CBZ and image files, but
  redaction annotations are PDF-only, so those inputs are rejected with exit code 2; convert
  them to PDF first.
- Document metadata (Author, Title, XMP) and annotation/comment content are not handled,
  since they don't appear in `get_text()`.
- A term split across multiple lines in the PDF layout might not be found.
- Scanned PDFs (image-only, with no extractable text) require OCR upstream: the tool finds
  nothing to redact in that case.
- Whole-word matching (the default for `-t`) only checks the character immediately before/after
  a match on the same line, so it's most reliable for terms that don't themselves span a line
  break. If a PDF packs two words so tightly that the extracted text runs them together with no
  gap at all, there is no boundary left to find: `--no-whole-word` is the fallback there.

Always verify the output with `pdftotext` and `pdfinfo -meta` before distribution.

## Windows compatibility

The project is tested in CI on Linux, Windows, and macOS (see `.github/workflows/tests.yml`)
and is compatible with Windows without modifications: it only uses `os.path` (no hardcoded
separators), no POSIX-only calls, and `os.path.samefile` has worked correctly on Windows
since Python 3.2.

## Development

```sh
pip install -e .[test]
pytest
pytest tests/test_core.py::test_redact_pdf_literal_term   # single test
```

## AI-assisted development

This project's code, tests, and documentation were developed with the assistance of AI
tools (Claude Code). Every change was reviewed before being published; please report any
issues you find via the project's issue tracker.

## License

[MPL-2.0](COPYING). The repository is [REUSE](https://reuse.software/) compliant; to verify:
`pipx run reuse lint`.
