Metadata-Version: 2.4
Name: anti-slop-kit
Version: 0.3.0
Summary: Detect and prevent slop in AI-generated code and documentation
Author: Anti-Slop Kit Contributors
License: MIT
Project-URL: Homepage, https://github.com/ameobius-ai/anti-slop-kit
Project-URL: Repository, https://github.com/ameobius-ai/anti-slop-kit
Project-URL: Documentation, https://github.com/ameobius-ai/anti-slop-kit#readme
Project-URL: Issues, https://github.com/ameobius-ai/anti-slop-kit/issues
Keywords: ai,code-quality,linting,documentation,fidelity
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: MIT License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Software Development :: Quality Assurance
Classifier: Topic :: Text Processing :: Linguistic
Requires-Python: >=3.9
Description-Content-Type: text/markdown
License-File: LICENSE
Provides-Extra: dev
Requires-Dist: pytest>=7.0.0; extra == "dev"
Requires-Dist: pytest-cov>=4.0.0; extra == "dev"
Requires-Dist: mypy<2.0.0,>=1.0.0; extra == "dev"
Requires-Dist: black>=23.0.0; extra == "dev"
Requires-Dist: ruff>=0.1.0; extra == "dev"
Requires-Dist: pre-commit>=3.0.0; extra == "dev"
Requires-Dist: bandit>=1.7.0; extra == "dev"
Requires-Dist: pip-audit>=2.6.0; extra == "dev"
Provides-Extra: hermes
Requires-Dist: pydantic>=2.0.0; extra == "hermes"
Dynamic: license-file

[![CI](https://github.com/ameobius-ai/anti-slop-kit/actions/workflows/ci.yml/badge.svg)](https://github.com/ameobius-ai/anti-slop-kit/actions/workflows/ci.yml)
[![Python 3.9+](https://img.shields.io/badge/python-3.9+-blue.svg)](https://www.python.org/downloads/)
[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT)
[![Code style: black](https://img.shields.io/badge/code%20style-black-000000.svg)](https://github.com/psf/black)
[![Dependabot](https://img.shields.io/badge/dependabot-enabled-blue.svg)](https://github.com/ameobius-ai/anti-slop-kit/network/updates)

# anti-slop-kit

Controlled-language writing skills and deterministic linters that remove AI slop
from technical prose. Five languages: English (ASD-STE100 mechanics), Russian
(GOST R 58049-2017, clause 8.2.3), Spanish, German and French.

A skill tells the model how to write. A linter proves whether the model did it.
The linter is the part most anti-slop advice leaves out.

## Installation

### Prerequisites

- Python 3.9 or higher
- pip (Python package manager)

### Quick Install

Install via pip:

    pip install anti-slop-kit

### Development Install

Clone and install in development mode:

    git clone https://github.com/ameobius-ai/anti-slop-kit.git
    cd anti-slop-kit
    pip install -e ".[dev]"
    pip install pre-commit
    pre-commit install

### Verify Installation

Check if package is installed:

    anti-slop-check --version

The command lists every rulepack with its version and calibration date.
The language linters themselves ship in the source tree, so a
non-editable `pip install` carries only the `tools` package - install
editable (`pip install -e .`) to lint with the rulepacks.

Or run tests:

    python -m pytest tests/


## Usage

See the [documentation](docs/) for detailed usage examples and API reference.

### Quick Start

1. Install the package
2. Import the main module
3. Use the analysis functions
4. Review the results

For more examples, check the [examples directory](examples/) and [API documentation](docs/api.md).

### Configuration

Create a configuration file to customize behavior. See [configuration guide](docs/configuration.md) for details.

### Command Line

Use the CLI tool for batch processing and automation. Run `anti-slop --help` for available commands.

## Who this is for

**Use this for:** API docs, runbooks, release notes, incident reports, onboarding
docs, support macros, changelogs — any text where a reader must act correctly on
the first read. Also useful as a gate on LLM-generated documentation.

**Do not use this for:** essays, marketing copy where voice is the point, fiction,
or anything where rhythm and register matter more than being parsed correctly.
The sentence-length and semicolon rules will fight you, and they should: they come
from maintenance-manual standards, not from general writing advice.

**What the score means:** violations per 100 words. A smoke alarm, not a grade.
The useful signal is the delta across revisions of the same text. An absolute
threshold means something only once a team picks one for a document class — the
CI samples gate at 2, and that is a convention, not a law.

## What is in here

```
AGENTS.md                instructions for an agent working in this repository
en/SKILL.md              ste-writing skill, English
en/ste-lint.py           English linter, 11 rule groups
en/samples/              one slop text and one clean rewrite
ru/SKILL.md              utrya-writing skill, Russian
ru/ru-ste-lint.py        Russian linter, 13 rule groups + typography
ru/samples/              one slop text and one clean rewrite
es/, de/, fr/            Spanish, German, French linters and samples
harness/SKILL.md         separate skill: how to design an agent harness
evals/                   eval harness: 14 tasks, 4 conditions, scorer, runner
examples/                five before/after pairs with measured scores
tests/                   unittest suite, standard library only
tools/                   JSON wrappers for the five linters
scripts/check.sh         the whole gate: tests, then the sample linters
hooks/pre-commit         git hook that blocks a commit above the limit
hooks/pre-push           git hook that runs the whole gate before a push
.pre-commit-config.example.yaml
RESULTS.md               measured scores and their limits
CONTRIBUTING.md          how to contribute rules, tests, and fixes
demo.sh                  one-command demo: samples, findings, tests
```

## Quick start

```sh
git clone https://github.com/ameobius-ai/anti-slop-kit
cd anti-slop-kit

./demo.sh
```

Or step by step:

```sh
python3 en/ste-lint.py en/samples/baseline.md en/samples/ste.md
python3 ru/ru-ste-lint.py ru/samples/baseline.md ru/samples/utr.md

python3 -m unittest discover -s tests
```

No dependencies. Python 3.9 or later. The linters use the standard library only,
because a skill directory is copied as a unit and must keep working after the copy.

## The gate

One entry point runs everything this project checks:

```sh
bash scripts/check.sh          # tests, then the sample linters
bash scripts/check.sh tests
bash scripts/check.sh lint
```

`.github/workflows/ci.yml` calls the same script, so a green local run and a
green CI run cannot disagree about what they checked.

GitHub Actions is disabled at the account level for the account that hosts this
repository: `POST /actions/workflows/ci.yml/dispatches` answers 422, `Actions
has been disabled for this user`. Until that is lifted the workflow never runs
here, and the local hooks are the only enforcement that exists:

```sh
ln -s ../../hooks/pre-commit .git/hooks/pre-commit   # blocks one bad file
ln -s ../../hooks/pre-push   .git/hooks/pre-push     # blocks a bad push
```

The workflow file stays in the tree because a fork with Actions enabled runs it
unchanged.

## Score

The score is violations per 100 words. Lower is cleaner.

| Text | Score | Longest sentence |
| --- | --- | --- |
| `en/samples/baseline.md` | 29.94 | 49 words |
| `en/samples/ste.md` | 0.83 | 14 words |
| `ru/samples/baseline.md` | 33.33 | 27 words |
| `ru/samples/utr.md` | 0.00 | 11 words |

Read `RESULTS.md` before you quote these numbers. Two texts per language is a
smoke test, not a benchmark. `evals/` holds the harness for a real measurement
across seven tasks per language and four prompt conditions. First live runs were
executed on 2026-08-04 (EN 23/28 cells, RU 24/28, via a local OpenAI-compatible
gateway); see `evals/README.md` for the setup and scores. No number on this page
comes from it yet. A separate lane scores the detection side:
`evals/detection_benchmark.py` runs every kit signal over a labeled AI/human
corpus (public eras via the fetchers in `scripts/`, the current generation via
`scripts/corpus_from_evals_run.py` on top of a `run.py` run); the measured
tables are in `RESULTS.md`, under the era ladder.

## Russian is a first-class citizen

The RU side is not a translation of the EN side. English plain-language tooling
is crowded; a deterministic Russian linter is rare. It targets канцелярит,
отглагольные существительные, цепочки родительного падежа and причастные
обороты, plus typography (ёлочки, тире), against ГОСТ Р 58049-2017 §8.2.3 (УТР).
It carries its own lexicon, its own morphology handling (ё-folding, a participle
stoplist), and its own samples and scores.

## Use it in a pipeline

The linters return exit code 1 when a file scores above the limit, so they can
gate a build:

```sh
python3 en/ste-lint.py --max 5 docs/*.md
python3 ru/ru-ste-lint.py --max 5 --json README.ru.md
```

Exit codes:

- `0`: every file is at or below the limit, or no limit was given
- `1`: at least one file is above the limit
- `2`: bad option or unreadable file

Git hook:

```sh
ln -s ../../hooks/pre-commit .git/hooks/pre-commit
chmod +x hooks/pre-commit
ANTI_SLOP_MAX=3 git commit          # change the limit for one commit
git commit --no-verify              # skip the hook
```

For [pre-commit](https://pre-commit.com), copy `.pre-commit-config.example.yaml`
and adjust the two paths.

## Explain a score

A score says where the problems are, not only how many. `--explain` prints one
line per finding: line number, rule, matched text, and the suggested fix.

```sh
python3 en/ste-lint.py --explain docs/draft.md
```

```text
draft.md                     words=  412 total=   9 per100w=  2.18 maxsent= 24
  L14    passive_voice         'is handled'                           Name the actor. Use active voice.
  L22    banned_word           'utilize'                              Use 'use' instead.
```

## Split the score

One total hides two different problems. `--breakdown` prints them apart: `slop`
counts banned words, marketing adjectives, AI filler and hedges; `cl` counts the
controlled-language mechanics, which are sentence length, passive voice,
nominalizations and participle chains.

```sh
python3 ru/ru-ste-lint.py --breakdown ru/samples/baseline.md
python3 en/ste-lint.py --only slop docs/draft.md
```

```text
baseline.md            words=  117 total=   39 per100w= 33.33 maxsent= 27 slop=   15 cl=   24
```

The split changes what you do next. In `ru/samples/baseline.md`, 24 of the 39
findings are structural, so a search for banned words finds 15 and misses the
larger half. `--only slop` and `--only cl` gate on one component alone, which
helps when a document class tolerates long sentences but not marketing language.

## GitHub Actions annotations

`--format github` emits workflow commands, so findings appear inline on pull
request diffs when the linter runs in GitHub Actions:

```yaml
- name: Lint prose
  run: python3 en/ste-lint.py --format github --max 5 docs/*.md
```

Each finding becomes a `::warning` annotation with file, line, rule name and
suggested fix. Combine with `--max` to fail the job and annotate at once.

## Exclude a region

The linters skip frontmatter, code blocks, inline code, link targets, bare URLs
and HTML comments. To exclude prose as well:

```markdown
<!-- anti-slop: off -->
A quoted paragraph that you must not rewrite.
<!-- anti-slop: on -->
```

## What the score does not tell you

The linters match patterns. They do not read.

- A score of 0 says nothing about whether the text is correct or complete.
- Every rule can produce a false positive. Passive voice is right when the actor
  is unknown. Some long sentences are clear.
- Use the score to find candidates for a rewrite, not to grade a writer.

## Examples

See the [examples directory](examples/) for practical usage examples including:

- **Basic Analysis**: Analyze single text files
- **Batch Processing**: Process multiple files efficiently  
- **Custom Patterns**: Add your own detection rules
- **Configuration**: Customize behavior with YAML
- **CI/CD Integration**: Use with GitHub Actions and pre-commit hooks

### Quick Start

1. Install the package
2. Create a configuration file (optional)
3. Run analysis on your files
4. Review results and adjust threshold

For detailed examples with code, visit the [examples directory](examples/) or check the [documentation](docs/).

### Common Use Cases

- **Code Review**: Check PR descriptions for AI patterns
- **Documentation**: Ensure technical writing quality
- **Content Creation**: Review blog posts and articles
- **CI/CD Pipeline**: Automatically check content quality
- **Pre-commit Hook**: Catch issues before committing

For more examples, see [USAGE.md](docs/USAGE.md).
## Performance
Benchmarks and optimization information for anti-slop-kit.

### Benchmarks
Performance varies by file size:
- Small files (<10KB): <0.1s, ~50MB memory
- Medium files (10-100KB): 0.1-1s, ~100MB memory
- Large files (100KB-1MB): 1-10s, ~200MB memory
- Very large files (>1MB): 10s+, ~500MB memory

### Optimization Tips
1. Exclude large directories (node_modules, .git, dist)
2. Enable parallel processing for multiple files
3. Use incremental analysis with git diff
4. Filter by file types to skip irrelevant files
5. Process in batches for better memory management

### Resource Usage
- CPU: 1 core for single file, multiple cores for batch
- Memory: 50MB base + 10-50MB per file
- Disk: Read-only analysis, minimal writes

### Performance Tuning
Configure in .anti-slop.yaml: parallel, workers, chunk_size, cache

### Known Limitations
- Large files (>10MB) may cause high memory usage
- Complex regex patterns slow down analysis
- Network features add latency

For more details, see performance benchmarks in the test suite.
## Security

### Security Features

anti-slop-kit includes several security-focused features:

- **Dependency Scanning**: Automated checks for known vulnerabilities
- **Code Analysis**: Detection of potentially unsafe patterns
- **Input Validation**: Sanitization of user-provided content
- **Secure Defaults**: Conservative security settings out of the box

### Best Practices

Follow these security best practices when using anti-slop-kit:

1. **Keep Dependencies Updated**: Regularly update anti-slop-kit and its dependencies
2. **Review Configuration**: Audit your `.anti-slop.yaml` for sensitive data
3. **Use Virtual Environments**: Isolate dependencies to prevent conflicts
4. **Monitor Logs**: Check analysis logs for suspicious patterns
5. **Limit Permissions**: Run with minimal required permissions
6. **Secure Configuration Files**: Don't commit secrets to version control

### Reporting Vulnerabilities

If you discover a security vulnerability, please report it responsibly:

1. **Do NOT open a public issue**
2. **GitHub Security Advisory**: create a private advisory via the [Security tab](https://github.com/ameobius-ai/anti-slop-kit/security/advisories)
3. Include detailed description and reproduction steps
4. Allow reasonable time for response and fix

We will acknowledge receipt within 48 hours and provide a timeline for fixing the issue.

### Security Considerations

- **Data Privacy**: anti-slop-kit processes text locally, no data sent to external servers
- **File Access**: Only reads files you explicitly specify
- **Network Access**: Minimal network requests (only for dependency updates if enabled)
- **Code Execution**: Does not execute analyzed code, only parses and analyzes

### Dependency Security

anti-slop-kit uses Dependabot for automated dependency updates:

- Weekly security scans
- Automatic PRs for security updates
- Manual review required before merging

For production use, consider:
- Pinning dependency versions
- Running `pip-audit` regularly
- Using lock files (requirements.txt or Pipfile.lock)

### Security Updates

Security updates are released as soon as possible after vulnerability discovery.
Subscribe to GitHub releases or watch the repository for security announcements.

For more information, see [SECURITY.md](SECURITY.md).
## FAQ

### General Questions

**Q: What does anti-slop-kit do?**
A: It analyzes text to detect AI-generated patterns and provides a quality score.

**Q: What languages are supported?**
A: English, Russian, Spanish, German and French — each linter lives in its `<lang>/` directory, and every clean sample is pinned in the gate.

**Q: Is it free to use?**
A: Yes, MIT license. Free for personal and commercial use.

**Q: Where can I get help?**
A: Check [TROUBLESHOOTING.md](docs/TROUBLESHOOTING.md) or open a [GitHub issue](https://github.com/ameobius-ai/anti-slop-kit/issues).

### Usage Questions

**Q: How do I install it?**
A: `pip install anti-slop-kit` (see [Installation](#installation) section above)

**Q: How do I configure it?**
A: Create `.anti-slop.yaml` in your project root (see [Configuration](#usage) section)

**Q: Can I use it with pre-commit?**
A: Yes, see [Examples](#examples) section for pre-commit hook setup

**Q: Does it work in CI/CD?**
A: Yes, see [Examples](#examples) for GitHub Actions integration

### Technical Questions

**Q: What's the scoring system?**
A: 0-100 scale: 90-100 Excellent, 70-89 Good, 50-69 Fair, 0-49 Poor

**Q: Can I customize detection patterns?**
A: Yes, add custom patterns in configuration file (see [Examples](#examples))

**Q: How fast is it?**
A: See [Performance](#performance) section for benchmarks

**Q: Does it send data anywhere?**
A: No, all processing is local. No external API calls.

### Troubleshooting

**Q: Installation fails?**
A: See [TROUBLESHOOTING.md](docs/TROUBLESHOOTING.md) for common solutions

**Q: Too many false positives?**
A: Adjust strictness level or add exclusions in configuration

**Q: Running too slow?**
A: See [Performance](#performance) for optimization tips

For more questions, see [TROUBLESHOOTING.md](docs/TROUBLESHOOTING.md) or open a [GitHub issue](https://github.com/ameobius-ai/anti-slop-kit/issues).
## Contributing

See `CONTRIBUTING.md` for the ground rules (standard library only, no shared
modules between linters, a test in the same commit as a rule change) and how to
add a banned word or report a false positive.

## Sources

- ASD-STE100 Simplified Technical English, Issue 9 (15 January 2025), ASD and the
  STEMG: https://asd-ste100.org. The specification is copyrighted. This repository
  reproduces the mechanics and no part of the text. Request a free copy from ASD.
- GOST R 58049-2017, clause 8.2.3, controlled Russian technical language.
- The English skill follows the approach shown in
  https://github.com/woosal1337/blog/tree/main/videos/ep01-the-cure-for-ai-slop.
  The skill and the linter here are written from scratch and share no code with it.
- https://github.com/talkstream/ru-text is a larger Russian rule set (about 1044
  rules) and works well next to this kit.
- `harness/SKILL.md` is built from
  https://github.com/ai-boost/awesome-harness-engineering (CC0) and the sources it
  lists.

## License

MIT. See `LICENSE`.


## Acknowledgments

### Contributors

Thanks to all contributors who have helped improve anti-slop-kit:

- [@ameobius-ai](https://github.com/ameobius-ai) - Project creator and maintainer
- Community contributors - Thank you for your PRs, issues, and feedback!

### Dependencies

anti-slop-kit is built on the shoulders of giants:

- [Python](https://www.python.org/) - Programming language
- [pytest](https://pytest.org/) - Testing framework
- [PyYAML](https://pyyaml.org/) - YAML parsing
- [GitHub Actions](https://github.com/features/actions) - CI/CD platform
- [mypy](http://mypy-lang.org/) - Static type checker
- [black](https://black.readthedocs.io/) - Code formatter
- [isort](https://pycqa.github.io/isort/) - Import sorter
- [flake8](https://flake8.pycqa.org/) - Linting

### Inspiration

This project was inspired by:

- [Write Good](https://github.com/btford/write-good) - Naive linter for English prose
- [Hemingway Editor](https://hemingwayapp.com/) - Writing clarity tool
- [Grammarly](https://www.grammarly.com/) - Writing assistant
- [Proselint](https://github.com/amperser/proselint/) - Linter for prose

### Special Thanks

- The open source community for amazing tools and libraries
- Everyone who has reported issues and suggested improvements
- Technical writers and editors who provided feedback on pattern definitions


## Citation

If you use anti-slop-kit in your research, please cite:

```
ameobius-ai. (2026). anti-slop-kit: Text Quality Analysis Toolkit.
https://github.com/ameobius-ai/anti-slop-kit
```


## Related Projects

Other tools for improving writing quality:

- **[Write Good](https://github.com/btford/write-good)** - Naive linter for English prose
- **[Proselint](https://github.com/amperser/proselint)** - Linter for prose
- **[Hemingway Editor](https://hemingwayapp.com/)** - Writing clarity tool
- **[Grammarly](https://www.grammarly.com/)** - AI writing assistant
- **[LanguageTool](https://languagetool.org/)** - Grammar and style checker


## Roadmap

### Shipped (ahead of the Q4 plan)

- Custom rule engine for user-defined patterns (`tools/aslint/custom_rules.py`,
  `--rules rules.yaml`)

### Current Focus (Q3 2026)

- Improve test coverage to 90%+ (branch coverage tracked via `.coveragerc`)
- Create web interface for interactive analysis

### Planned Features (Q4 2026)

- IntelliJ IDEA plugin
- REST API for remote analysis
- Batch processing optimizations

### Future Vision (2027)

- Real-time analysis mode
- Integration with popular writing tools
- Enterprise features (team dashboards, reporting)
