Metadata-Version: 2.4
Name: humanizer-pro
Version: 4.15.0
Summary: Deterministic, zero-dependency audit for AI-writing tells, leaked chatbot artifacts, and rewrite fidelity.
Author: Humanizer Pro contributors
License: MIT
Project-URL: Homepage, https://github.com/eddyplolz/humanizer-pro
Project-URL: Changelog, https://github.com/eddyplolz/humanizer-pro/blob/main/CHANGELOG.md
Project-URL: Issues, https://github.com/eddyplolz/humanizer-pro/issues
Keywords: ai-writing,linter,prose,markdown,sarif
Classifier: Environment :: Console
Classifier: License :: OSI Approved :: MIT License
Classifier: Programming Language :: Python :: 3
Classifier: Topic :: Text Processing :: Linguistic
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Dynamic: license-file

# Humanizer Pro

[![CI](https://github.com/eddyplolz/humanizer-pro/actions/workflows/ci.yml/badge.svg)](https://github.com/eddyplolz/humanizer-pro/actions/workflows/ci.yml) [![License: MIT](https://img.shields.io/badge/license-MIT-blue.svg)](LICENSE) ![Python 3.10+](https://img.shields.io/badge/python-3.10%2B-blue.svg) ![False positives: 0.2%](https://img.shields.io/badge/false_positives-0.2%25_measured-green.svg) ![Runs locally](https://img.shields.io/badge/runs-locally-lightgrey.svg)

**Finds the habits that make text read as machine-written, removes them, and checks that every fact survived.** Runs on your computer with no account, no API key, and no network connection. Measured false-positive rate on 1,264 human-written documents: 0.2%.

| Before | After | What changed |
|---|---|---|
| In today's rapidly evolving digital landscape, effective collaboration serves as a crucial cornerstone for organizations seeking to unlock their full potential. | Good collaboration depends less on the tool than on whether people know what decisions they own, where work is tracked, and how quickly blockers get resolved. | Stock opener cut; "serves as a cornerstone" became a plain claim; the vague promise became three concrete conditions. |
| It is important to note that this approach is not just about tools, but about creating a vibrant culture of innovation. | *(folded into the sentence above)* | "It is important to note" and "not just X, but Y" are formulas, and "vibrant culture of innovation" said nothing the first sentence did not. |

Three ways to use it:

```bash
# 1. As a skill in Claude Code (then type /humanizer-pro)
git clone https://github.com/eddyplolz/humanizer-pro.git ~/.claude/skills/humanizer-pro

# 2. As a skill in Codex and other agents
git clone https://github.com/eddyplolz/humanizer-pro.git ~/.agents/skills/humanizer-pro

# 3. As a command-line checker (Python 3.10+)
pip install humanizer-pro && humanizer-audit draft.md
```

Windows paths and the GitHub Action are in [Install the skill](#install-the-skill) and [Check files automatically](#check-files-automatically).

**Who it is for:** anyone cleaning up an AI draft before it goes out; editors and docs teams who want a check in CI; Wikipedia and wiki editors, since it knows wikitext and neutral tone; and coding agents auditing their own prose.

**What makes it different from a "humanizer" website:** your text never leaves your machine, every flag comes with the reason and the line, a compare mode proves the rewrite kept your numbers, names, dates, and links, and the error rates are measured on public data you can rebuild yourself.

## Contents

1. [What it does](#what-it-does)
2. [What it does not do](#what-it-does-not-do)
3. [Your privacy](#your-privacy)
4. [Install the skill](#install-the-skill)
5. [Use the skill](#use-the-skill)
6. [Check files from the command line](#check-files-from-the-command-line)
7. [Check files automatically](#check-files-automatically)
8. [How well it works](#how-well-it-works)
9. [How it decides](#how-it-decides)
10. [Help improve it](#help-improve-it)
11. [Credits and license](#credits-and-license)
12. [Version history](#version-history)

## What it does

AI drafts repeat the same habits: stock phrases, inflated importance, and filler that sounds complete but says little. Readers notice, and then they stop trusting the text.

Humanizer Pro:

- finds those habits and explains each one
- rewrites the draft in plain language, if you ask it to
- checks that a rewrite kept the numbers, dates, full names, links, and quotes it can detect
- flags text that a chatbot left behind, such as citation codes and placeholder fields

Here is an example.

Before:

> In today's rapidly evolving digital landscape, effective collaboration serves as a crucial cornerstone for organizations seeking to unlock their full potential. It is important to note that this approach is not just about tools, but about creating a vibrant culture of innovation.

After:

> Good collaboration depends less on the tool than on whether people know what decisions they own, where work is tracked, and how quickly blockers get resolved.

The command-line checker explains the first version line by line (shortened):

```text
$ humanizer-audit eval/fixtures/ai-slop-general.md
eval/fixtures/ai-slop-general.md — risk 60
  WARNING L3:C1 family3.filler_framing — Filler framing or superficial analysis [In today's]
  WARNING L3:C37 family4.ai_vocab_cluster — AI-vocabulary cluster [cornerstone, crucial, enhance, foster, landscape, robust, unlock, vibrant]
  WARNING L6:C29 family5.syntactic_tell — Syntactic tell or hedged construction [It is important to note]
  WARNING L6:C75 family7.rhetorical_formula — Rhetorical formula or forced cadence [not just about tools, but]
  WARNING L7:C80 family3.filler_framing — Filler framing or superficial analysis [In conclusion]
```

## What it does not do

This project makes no claims about AI detectors.

It does not add a fake personality. Swapping one stock phrase for a casual one trades one habit for another, so the skill avoids both.

It does not rewrite text that is already clean. A built-in restraint check returns good writing close to how it arrived.

> [!IMPORTANT]
> A high score shows writing habits. It does not prove that a machine wrote the text, and it proves nothing about a person. Research has found that AI detectors wrongly flag more than 60% of essays by people writing in English as a second language (Liang et al., Stanford, *Patterns*, 2023). Do not use this tool as the only basis for an academic, hiring, or authorship decision.

## Your privacy

The command-line checker reads files on your computer and sends nothing anywhere. It makes no network connections and keeps no logs.

The skill runs inside your coding agent, such as Claude Code or Codex. Your text goes wherever your agent already sends it, and nowhere else.

Two features can put your text somewhere else, and only when you choose them:

- SARIF results contain short quotes from your text. If you upload them to GitHub code scanning, those quotes are stored with your repository's code scanning results.
- The GitHub Action runs in your own GitHub Actions workflow, so your files are read on GitHub's servers, as with any other check you run there.

## Install the skill

### Before you start

You need [git](https://git-scm.com/downloads). To use the command-line checker as well, you also need [Python](https://www.python.org/downloads/) 3.10 or newer.

### Install for Claude Code

On a Mac or Linux computer, run these commands in Terminal:

```bash
mkdir -p ~/.claude/skills
git clone https://github.com/eddyplolz/humanizer-pro.git ~/.claude/skills/humanizer-pro
```

On Windows, run these commands in Command Prompt:

```bat
mkdir "%USERPROFILE%\.claude\skills"
git clone https://github.com/eddyplolz/humanizer-pro.git "%USERPROFILE%\.claude\skills\humanizer-pro"
```

To check it worked, open the new `humanizer-pro` folder and look for a file called `SKILL.md`. Claude Code loads the skill the next time you start it. You can then type `/humanizer-pro`.

### Install for Codex and other agents

On a Mac or Linux computer:

```bash
mkdir -p ~/.agents/skills
git clone https://github.com/eddyplolz/humanizer-pro.git ~/.agents/skills/humanizer-pro
```

On Windows:

```bat
mkdir "%USERPROFILE%\.agents\skills"
git clone https://github.com/eddyplolz/humanizer-pro.git "%USERPROFILE%\.agents\skills\humanizer-pro"
```

If you use Codex and have set `CODEX_HOME`, install into `$CODEX_HOME/skills` instead. Otherwise `~/.codex/skills` also works.

## Use the skill

Paste your text after a request in plain words. You do not need to learn any commands.

| What you want | What to say |
|---|---|
| A cleaned-up draft | "Humanize this" or "make this less AI" |
| A score and reasons, with no changes | "AI check," "score this," or "do not rewrite" |
| A score, reasons, and a rewrite | "Full audit" |
| Tighter sentences | "Style edit" or "tighten this" |
| Neutral, encyclopedia-style writing | "Wiki mode," or mention wikitext or citations |

The full audit rates the draft on 6 qualities: directness, rhythm, trust, authenticity, density, and restraint. It shows its reasoning before the rewrite.

Wiki mode removes promotional tone, keeps citations, and flags claims that have no source.

Style edits use a short checklist based on *The Elements of Style*. The full 1918 text is included, and the skill only loads it when you ask.

## Check files from the command line

The checker reads text files and lists each problem with its line number and the words that triggered it. It never changes your files.

### Install the checker

Install it with [pipx](https://pipx.pypa.io/):

```bash
pipx install git+https://github.com/eddyplolz/humanizer-pro
```

You can also run it from a downloaded copy of this repository, without installing anything:

```bash
python3 scripts/humanizer_audit.py path/to/draft.md
```

On Windows, use `py -3` instead of `python3`.

### Check a file

```bash
humanizer-audit path/to/draft.md
```

You can:

- check several files or folders at once (a folder means every `.md` and `.txt` file inside it)
- add `--json` to get results a program can read
- add `--sarif results.sarif` to save results in SARIF, the format GitHub uses to show problems on pull requests
- add `--include-code` to check code samples as well (they are skipped by default, but leaked chatbot text inside code is always caught)
- use `--compare original.md revised.md` to check that a rewrite kept the numbers, dates, names of 2 or more words, links, citations, quotes, and code samples in the original (it does not catch every change: a one-word name or a changed fact, such as "delayed" becoming "canceled," can pass)

### Understand the result

The checker ends with an exit code that scripts can act on.

| Code | Meaning | What to do |
|---|---|---|
| 0 | Pass | Nothing. |
| 1 | Review: the risk score reached the threshold (60 unless you set `--fail-score`) | Read the findings and decide. |
| 2 | Block: leaked chatbot text, a placeholder, or hidden characters | Fix before you publish. |
| 3 | The checker could not run, for example a wrong option or a missing file | Check the command. |

## Check files automatically

### Before each commit

[pre-commit](https://pre-commit.com/) is a tool that runs checks every time you commit. To add this check, put the following in a file called `.pre-commit-config.yaml` at the top of your project:

```yaml
repos:
  - repo: https://github.com/eddyplolz/humanizer-pro
    rev: v4.15.0
    hooks:
      - id: humanizer-audit
```

The check runs on Markdown and text files. A review or block result stops the commit.

### In GitHub Actions

This repository is also a GitHub Action. To check your writing on every pull request:

1. In your project, create the file `.github/workflows/writing.yml`.
2. Paste in the workflow below.
3. Change `paths` to the files or folders you want checked.
4. Commit the file.

```yaml
name: Writing check
on: pull_request

permissions:
  contents: read
  security-events: write  # needed to show results on the pull request

jobs:
  audit:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: eddyplolz/humanizer-pro@v4.15.0
        with:
          paths: docs README.md
          fail-on: block
      - uses: github/codeql-action/upload-sarif@v3
        if: always()
        with:
          sarif_file: humanizer-audit.sarif
```

The last step shows each finding on the pull request, next to the line it refers to.

You can set these options:

| Option | What it does | Default |
|---|---|---|
| `paths` | Files or folders to check, separated by spaces or new lines | `.` (everything) |
| `fail-on` | When the step fails: `block`, `review`, or `never` | `block` |
| `fail-score` | Risk score that counts as a review | `60` |
| `sarif-file` | Where to save the SARIF results (leave empty to skip) | `humanizer-audit.sarif` |
| `include-code` | Set to `true` to check code samples too | `false` |

The action needs Python 3 on the runner. GitHub's Ubuntu and macOS runners have it.

## How well it works

This project measures its error rates instead of only describing them. It counts errors in both directions:

- false positives: how often it flags writing by people
- catch rate: how often it flags writing by AI

| What is measured | Writing used | Result |
|---|---|---|
| False positives | 1,264 test documents written before ChatGPT existed: Stack Exchange answers, Wikipedia articles, newspapers, essays, and Python Enhancement Proposals | 0.2% overall; 0.0% to 0.7% by type of writing ([full results](corpus/RESULTS.md)) |
| Catch rate | 638 test documents from 2023 AI models (RAID) and real ChatGPT replies (WildChat) | 3.1% overall; 0.0% to 5.5% by type of writing |

### False positives

Every document in the human set was written before ChatGPT was released, so any flag on one is a mistake by definition. The latest rates, at the default threshold, are:

| Type of writing | Documents | False-positive rate |
|---|---|---|
| Stack Exchange answers | 450 | 0.0% |
| Python Enhancement Proposals | 303 | 0.7%, including 1 blocked |
| Essays | 250 | 0.0% |
| Newspapers | 257 | 0.4% |
| Wikipedia articles | 4 | 0.0% |

By the period the text was written in:

| Period | Documents | False-positive rate |
|---|---|---|
| Before 1930 | 507 | 0.2% |
| 1930 to 2017 | 567 | 0.4% |
| 2018 to 2022 | 190 | 0.0% |

These rates come from the current rules, measured on the test group only (1,264 documents). A block counts as a false positive. The Wikipedia slice is small, so its rate says little yet.

Every document comes from a public source. The project publishes a fingerprint of each one (a unique code), with its date, word count, and a public pointer such as a Wikipedia revision id or a Stack Exchange answer id, so anyone can rebuild the collection and check the numbers. It never publishes the text.

### Catch rate

The AI text comes from 2 public sources, labeled by model:

- text from [RAID](https://github.com/liamdugan/raid), a research benchmark of 2023 AI models
- first replies from [WildChat](https://huggingface.co/datasets/allenai/WildChat-1M), a public collection of real ChatGPT conversations

A pool of answers from current Claude models to [60 published prompts](corpus/machine_prompts.json) can be added with `scripts/generate_machine.py`; it is not part of the published numbers.

The catch rate at the default threshold is 3.1% overall (20 of 638 test documents): 5.5% on chat replies, 0.0% on news and encyclopedia text. The full table by model is in the [results](corpus/RESULTS.md). 2 limits apply:

- it describes those AI models only
- the human essays and newspapers are about 100 years old, so part of any gap there reflects the era, not the author

### How the results stay honest

Every document goes into one of 2 groups, chosen by its fingerprint. About a quarter go into a tuning group, and the rest into a test group.

Rules are adjusted using the tuning group only. Every published rate uses the test group only, so a rule cannot be tuned to make its own results look better.

This project also checks its own documentation with the same rules, using `scripts/self_scan.py`. It publishes 2 scores: one that counts every phrase this page quotes as a bad example, and one that leaves quotes and code out.

## How it decides

The checker looks for 9 families of habits, learned from cleanup work on Wikipedia and elsewhere. For example:

- inflated importance, such as "stands as a testament" or "pivotal moment"
- stock rhetorical patterns, such as "not just X, but Y" or lists that always come in threes
- text left behind by a chatbot, such as "I hope this helps," knowledge-cutoff notes, or citation codes like `oaicite`

The full list, with examples, is in the [catalog of habits](reference/tell-catalog.md).

2 principles guide every edit:

- a single instance proves nothing (one "crucial" is a coincidence, but a cluster is a pattern)
- over-editing is a failure (if a change makes good writing worse, the original goes back)

The checker is stricter for some kinds of writing than others. A chat message, an essay, a news story, and an encyclopedia article each have their own [rules for strictness](reference/registers.md).

## Help improve it

New rules go through a review process before they are added. A proposed rule is rejected if it:

- encourages over-editing
- repeats an existing rule
- reflects one person's taste
- takes more than about 80 words to state

The full process is in the [improvement guide](reference/improvement-loop.md).

Before you open a pull request, run these 2 checks. Both must pass.

```bash
python3 -m pytest -q tests
python3 scripts/self_scan.py
```

On Windows, use `py -3` instead of `python3`.

<details>
<summary>What each file and folder is for</summary>

| Path | What it is |
|---|---|
| `SKILL.md` | The skill itself: routing, principles, checklists, and scoring |
| `reference/` | Detailed guides: the catalog of habits, worked examples, style and wiki guides, strictness rules, and the improvement process |
| `scripts/humanizer_audit.py` | The checker |
| `scripts/self_scan.py` | Checks this project's own documentation |
| `scripts/corpus.py`, `scripts/fp_measure.py` | Build the test collection and measure error rates |
| `scripts/generate_machine.py` | Creates AI-written samples for the catch rate (maintainers only; needs an Anthropic API key) |
| `corpus/` | Document fingerprints, the published prompts, and measured results |
| `eval/` | Sample texts and the expected results for each |
| `tests/` | Automated tests |
| `pyproject.toml`, `action.yml`, `.pre-commit-hooks.yaml` | Packaging, the GitHub Action, and the pre-commit check |
| `agents/openai.yaml` | Display name and default prompt for Codex |
| `CHANGELOG.md` | Every change, by version |
| `WARP.md` | Guide for maintainers |

</details>

## Credits and license

This project uses the [MIT License](LICENSE).

It is a rebuild of [blader/humanizer](https://github.com/blader/humanizer) by Siqi Chen (MIT, copyright 2025). It adds a catalog of habits, detection of leaked chatbot text, a fact-checking compare mode, wiki and style modes, and measured error rates. The `LICENSE` file includes both copyright notices.

It draws on these sources:

- [blader/humanizer](https://github.com/blader/humanizer) by Siqi Chen (MIT)
- [Stop Slop](https://github.com/hardikpandya/stop-slop) by Hardik Pandya (MIT)
- [avoid-ai-writing](https://github.com/conorbronsdon/avoid-ai-writing) by Conor Bronsdon (MIT): the vocabulary tiers, strictness rules, self-check, and error-rate measurement design
- [humanize](https://github.com/harshaneel/humanize) by Harshaneel Gokhale (MIT): rhythm counts and several patterns
- [Wikipedia: Signs of AI writing](https://en.wikipedia.org/wiki/Wikipedia:Signs_of_AI_writing) by WikiProject AI Cleanup ([CC BY-SA 4.0](https://creativecommons.org/licenses/by-sa/4.0/))
- [Project Gutenberg #37134](https://www.gutenberg.org/ebooks/37134): the public-domain text of *The Elements of Style* by William Strunk Jr.
- [US-PD-Newspapers](https://huggingface.co/datasets/PleIAs/US-PD-Newspapers) by PleIAs: public-domain newspapers in the human test set
- [RAID](https://github.com/liamdugan/raid) by Dugan and others (MIT): AI-written test set
- [WildChat-1M](https://huggingface.co/datasets/allenai/WildChat-1M) by Ai2 (ODC-BY): AI-written test set
- [Python Enhancement Proposals](https://github.com/python/peps) (public domain): technical writing in the human test set

## Version history

Current release: **v4.15.0**. It puts the package on PyPI, rebuilds the measurement corpus from public sources only, publishes error rates in both directions, adds four editing principles and several new checks to the audit, and tests on Windows and macOS. Every change is listed in the [changelog](CHANGELOG.md).
