Metadata-Version: 2.4
Name: pdfblah
Version: 0.4.1
Summary: Trustworthy true edits to native PDFs: replace, redact, remove text, scrub PII, anonymize, mail-merge, and edit metadata. Fonts and layout kept perfect.
Project-URL: Homepage, https://pdfblah.com
Project-URL: Source, https://github.com/KuvopLLC/pdfblah
Project-URL: Issues, https://github.com/KuvopLLC/pdfblah/issues
Author: Kuvop LLC
License-Expression: MIT
License-File: LICENSE
Keywords: acrobat-alternative,anonymize,cli,edit,find,mail-merge,pdf,pdf-editor,pii,redact,replace,scrub,text
Classifier: Environment :: Console
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: End Users/Desktop
Classifier: Operating System :: MacOS
Classifier: Operating System :: POSIX
Classifier: Programming Language :: Python :: 3
Classifier: Topic :: Text Processing
Classifier: Topic :: Utilities
Requires-Python: >=3.9
Requires-Dist: faker>=20
Requires-Dist: pdfplumber>=0.11
Requires-Dist: pikepdf>=8
Requires-Dist: python-stdnum>=1.19
Provides-Extra: test
Requires-Dist: pytest; extra == 'test'
Requires-Dist: reportlab; extra == 'test'
Description-Content-Type: text/markdown

# pdfblah

[![PyPI](https://img.shields.io/pypi/v/pdfblah)](https://pypi.org/project/pdfblah/)
[![Python](https://img.shields.io/pypi/pyversions/pdfblah)](https://pypi.org/project/pdfblah/)
[![CI](https://github.com/KuvopLLC/pdfblah/actions/workflows/ci.yml/badge.svg)](https://github.com/KuvopLLC/pdfblah/actions/workflows/ci.yml)
[![License: MIT](https://img.shields.io/badge/license-MIT-blue.svg)](LICENSE)

Trustworthy, true edits to native PDFs, from the command line. Replace, redact, and
remove text, scrub or anonymize personal data, mail-merge a template, and read or
edit metadata. Fonts, spacing, and alignment stay perfect, and nothing you didn't
ask for is touched.

![pdfblah demo](https://pdfblah.com/demo.gif?v=3)

Most tools "edit" a PDF by painting a box over the old text and drawing new text
on top, which leaves the original underneath (copy and paste still reveals it) and
often adds a watermark. `pdfblah` rewrites the real text in the content stream, so:

- the old text is genuinely gone (`pdftotext`, Ctrl-F, and copy show only the new value)
- no overlay, no watermark
- your metadata (dates, Producer, XMP) is kept byte for byte, unless you choose to edit it
- alignment is auto-detected and kept, so right-aligned numbers stay flush
- fonts it cannot reproduce are refused instead of garbled

Pure Python. No system dependencies.

## Install

```sh
pipx install pdfblah      # recommended, isolated; or:  pip install pdfblah
```

On a Mac with Homebrew, use Homebrew's pipx:

```sh
brew install pipx && pipx install pdfblah
```

Also works with [uv](https://github.com/astral-sh/uv): `uv tool install pdfblah`.

## Use

Replace the first match:

```sh
pdfblah in.pdf out.pdf --find "Old Name" --replace "New Name"
```

Options:

```sh
--scope all         change every match           (default: first)
--scope 3           change the 3rd match
--ci                ignore case
--word              whole word only ("cat" will not match "category")
--regex             treat --find as a regex (\1 backrefs work in --replace)
--page 2            only page 2
--replace ""        delete the text
```

Many rules from a file (`FIND | REPLACE | FLAGS` per line):

```sh
pdfblah in.pdf out.pdf --rules rules.txt
```

```
# rules.txt
Old Company Name | New Company Name | all
CONFIDENTIAL DRAFT | FINAL | ci
Jane Doe | John Smith | all word
Total | Sum | 2
delete this phrase |
```

## Commands

The same engine (locate real text, then substitute something) as four presets.

**redact** removes the matched text for real (gone from `pdftotext`, Ctrl-F and copy)
and draws a bar over each spot. `--no-bar` removes the text with no mark.

```sh
pdfblah redact in.pdf out.pdf --find "Account 12345"
pdfblah redact in.pdf out.pdf --find "\d{3}-\d{2}-\d{4}" --regex   # every SSN
```

**scrub** finds structured personal data (email, IBAN, credit card, SSN, phone) and
removes it, or masks it. Cards and IBANs are checksum-validated, so ordinary numbers
are left alone.

```sh
pdfblah scrub in.pdf out.pdf
pdfblah scrub in.pdf out.pdf --types email,credit_card --mask "[redacted]"
```

**anonymize** replaces detected data with realistic, shape-preserving fakes so a
document is safe to share. The same value maps to the same fake; `--seed` makes it
reproducible. Names are swapped only when you list them.

```sh
pdfblah anonymize in.pdf out.pdf --names "Alison Cohen,Matthew Reider" --seed 7
```

**merge** fills a template once per data row: every `{{column}}` placeholder becomes
that row's value, one output PDF per row.

```sh
pdfblah merge template.pdf people.csv --out ./letters --name-col name
```

**meta** reports everything metadata-ish (DocInfo, XMP, pages, version, encryption),
which often reveals more than you expect (author, software, timestamps). With an
output file it can strip or set fields. Every other command keeps metadata intact.

```sh
pdfblah meta in.pdf                                  # report what's in there
pdfblah meta in.pdf clean.pdf --strip                # remove all metadata
pdfblah meta in.pdf out.pdf --set author="Jane Roe"  # set a field
```

Metadata edits can also ride along with any other command, or live in a rules file:

```sh
pdfblah redact in.pdf out.pdf --find "Acme Corp" --strip-metadata
pdfblah in.pdf out.pdf --find OLD --replace NEW --set-metadata author="Ops"
```

```
# rules.txt
@strip-metadata
@set-metadata author = Redacted Dept
CONFIDENTIAL | PUBLIC | all
```

## Library

```python
from pdfblah import process, redact, scrub, anonymize, merge, apply_rules

process("in.pdf", "out.pdf", "999.00", "42.00", scope="all", ci=True)
process("in.pdf", "out.pdf", r"\d{4}-\d{4}", "REDACTED", scope="all", regex=True)
redact("in.pdf", "out.pdf", "Account 12345")
scrub("in.pdf", "out.pdf", types=["email", "credit_card"])
anonymize("in.pdf", "out.pdf", names=["Alison Cohen"], seed=7)
merge("template.pdf", [{"name": "Alice"}], "./out")
```

Each call returns a report dict (`ok`, `count`, `refused`, `reason`, ...).

## What it does not do

Scanned PDFs (image only, no text layer) cannot be edited. Fonts that are not
embedded and not standard, or use a custom encoding, are refused rather than
rendered wrong. This is by design: a wrong-looking edit is worse than a clear "no".

## Hosted version

Want it without installing anything, or for a non-technical colleague? The hosted
version at **[pdfblah.com](https://pdfblah.com)** does the same edit in the browser:
upload, preview for free, download.

## License

MIT, (c) 2026 Kuvop LLC.
