Metadata-Version: 2.4
Name: fastnltk
Version: 0.5.5
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: Implementation :: CPython
Classifier: Programming Language :: Python :: Implementation :: PyPy
Classifier: Programming Language :: Rust
Classifier: License :: OSI Approved :: Apache Software License
Classifier: Topic :: Text Processing :: Linguistic
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Requires-Dist: nltk>=3.10
Requires-Dist: numpy>=2.2.6,<3.0
Requires-Dist: fastnltk[dev,test,lint] ; extra == 'all'
Requires-Dist: maturin>=1.14 ; extra == 'dev'
Requires-Dist: pytest>=9.1 ; extra == 'dev'
Requires-Dist: pytest-benchmark>=5.2 ; extra == 'dev'
Requires-Dist: ruff>=0.16 ; extra == 'dev'
Requires-Dist: mypy>=2.3 ; extra == 'dev'
Requires-Dist: ruff>=0.16 ; extra == 'lint'
Requires-Dist: mypy>=2.3 ; extra == 'lint'
Requires-Dist: pre-commit>=4.6 ; extra == 'lint'
Requires-Dist: pytest>=9.1 ; extra == 'test'
Requires-Dist: pytest-benchmark>=5.2 ; extra == 'test'
Requires-Dist: hypothesis>=6.161 ; extra == 'test'
Requires-Dist: nltk>=3.10 ; extra == 'test'
Provides-Extra: all
Provides-Extra: dev
Provides-Extra: lint
Provides-Extra: test
License-File: LICENSE
Summary: Drop-in Rust-accelerated replacement for NLTK. Same API, 5-50x faster.
Keywords: nlp,nltk,tokenization,tagging,stemming,parsing,text-processing,rust
Home-Page: https://github.com/wyattferguson/fastnltk
Author: Wyatt Ferguson
License-Expression: Apache-2.0
Requires-Python: >=3.10
Description-Content-Type: text/markdown; charset=UTF-8; variant=GFM
Project-URL: Changelog, https://github.com/wyattferguson/fastnltk/blob/main/CHANGELOG.md
Project-URL: Homepage, https://github.com/wyattferguson/fastnltk
Project-URL: Repository, https://github.com/wyattferguson/fastnltk

<div align="center">
  <h1>fastNLTK</h1>
  <p><strong>NLTK with a Rust engine.</strong><br>
  Drop-in replacement. Same API, Same data, 10× faster.</p>

  <p>
    <a href="https://pypi.org/project/fastnltk/"><img src="https://img.shields.io/pypi/v/fastnltk.svg" alt="PyPI"></a>
    <a href="https://pypi.org/project/fastnltk/"><img src="https://img.shields.io/pypi/pyversions/fastnltk.svg" alt="Python"></a>
    <a href="https://github.com/wyattferguson/fastnltk/actions/workflows/quality.yml"><img src="https://github.com/wyattferguson/fastnltk/actions/workflows/quality.yml/badge.svg" alt="CI"></a>
    <a href="https://www.rust-lang.org/"><img src="https://img.shields.io/badge/rust-1.97%2B-blue" alt="Rust"></a>
    <a href="LICENSE"><img src="https://img.shields.io/badge/license-Apache--2.0-blue.svg" alt="License"></a>
  </p>
</div>

---

[**NLTK**](https://www.nltk.org) is fantastic. The API is clean, and it does just about everything you'd want from an NLP library. The only drawback being it's pure Python, and on large inputs that starts to show. A single
call is fine. A million calls in a data pipeline is a different story.

**fastNLTK** is NLTK with the hot path rewritten in Rust. Same API, same
data, same results. Just faster — 5× to 455× depending on what you're doing.
No new dependencies, no YAML files, no config. Simply change your import and watch your code fly.

```python
# Before
import nltk
tokens = nltk.word_tokenize("The quick brown fox.")

# After
import fastnltk as nltk
tokens = nltk.word_tokenize("The quick brown fox.")  # same call, 5–50× faster
```

Your NLTK data (corpora, models, pickles, all of it) still works. Nothing to
re-download, nothing to migrate.

## Benchmarks

**368 Python drop-in compatibility tests against NLTK. 6 skipped (chat stdin). 1 expected failure.**

Benchmarked on release builds against NLTK 3.10. [Full results →](BENCHMARKS.md)

| Operation                  | NLTK      | fastNLTK | Speedup  |
| -------------------------- | --------- | -------- | -------- |
| TextTiling tokenizer       | 22043 ms  | 32 ms    | **698×** |
| Maxent train               | 33 ms     | 0.08 ms  | **431×** |
| windowdiff                 | 2.38 ms   | 0.01 ms  | **174×** |
| edit_distance              | 2.44 ms   | 0.02 ms  | **144×** |
| HMM tagger                 | 8.58 ms   | 0.10 ms  | **88×**  |
| pk                         | 2.23 ms   | 0.03 ms  | **83×**  |
| Treebank detokenizer       | 6.69 ms   | 0.12 ms  | **54×**  |
| Sentiment (VADER)          | 67.54 ms  | 1.79 ms  | **38×**  |
| Punkt sentence tokenizer   | 14.28 ms  | 0.43 ms  | **33×**  |
| Expression.fromstring      | 16.06 ms  | 0.54 ms  | **30×**  |
| Tweet tokenizer            | 83.00 ms  | 3.26 ms  | **25×**  |
| CFG grammar parser         | 0.05 ms   | 0.00 ms  | **23×**  |
| Quadgram collocations      | 98.74 ms  | 5.11 ms  | **19×**  |
| Lancaster stemmer          | 31.50 ms  | 1.42 ms  | **22×**  |
| Snowball stemmer           | 21.60 ms  | 1.77 ms  | **12×**  |

Geometric mean across 51 benchmarks: **9.5×**. Module-level breakdown:

| Module                        | Geo Mean | Top single |
| ----------------------------- | -------- | ---------- |
| [metrics](BENCHMARKS.md)      | **128×** | 174×       |
| [sentiment](BENCHMARKS.md)    | **38×**  | 38×        |
| [sem](BENCHMARKS.md)          | **30×**  | 30×        |
| [classify](BENCHMARKS.md)     | **25×**  | 431×       |
| [collocations](BENCHMARKS.md) | **14×**  | 19×        |
| [tree](BENCHMARKS.md)         | **11×**  | 11×        |
| [translate](BENCHMARKS.md)    | **10×**  | 10×        |
| [stem](BENCHMARKS.md)         | **9×**   | 22×        |
| [chunk](BENCHMARKS.md)        | **9×**   | 9×         |
| [tokenize](BENCHMARKS.md)     | **8×**   | 698×       |
| [cluster](BENCHMARKS.md)      | **6×**   | 6×         |
| [tag](BENCHMARKS.md)          | **5×**   | 88×        |
| [parse](BENCHMARKS.md)        | **4×**   | 23×        |
| [probability](BENCHMARKS.md)  | **4×**   | 4×         |
| [chat](BENCHMARKS.md)         | **3×**   | 3×         |
| [ccg](BENCHMARKS.md)          | **3×**   | 3×         |

## What's accelerated

Every module that has a Rust-backed engine:

| Module         | What's in Rust                                                                                    |
| -------------- | ------------------------------------------------------------------------------------------------- |
| `tokenize`     | Treebank, Toktok, Tweet, Regexp, Space, MWE, TextTiling, Punkt, SExpr, Logos DFA                  |
| `stem`         | Porter, Lancaster (full 124 rules), Snowball, Regexp, WordNet, ARLSTem, Cistem, ISRI, RSLP        |
| `tag`          | PerceptronTagger, HMM (integer Viterbi), TnT, Default/Unigram/Bigram/Trigram/Regexp/Affix |
| `classify`     | NaiveBayes, Maxent (GIS), TextCat                                                                 |
| `corpus`       | PlaintextCorpusReader, TaggedCorpusReader, CategorizedPlaintextCorpusReader                        |
| `probability`  | FreqDist, ConditionalFreqDist (shared references), MLE/Laplace/Lidstone prob dists                |
| `lm`           | MLE, Lidstone, Laplace, Kneser-Ney interpolated, Witten-Bell, StupidBackoff                       |
| `collocations` | Bigram/Trigram/Quadgram finders                                                                   |
| `metrics`      | edit_distance, jaccard, windowdiff, pk, BLEU, association, agreement, Spearman                    |
| `parse`        | CFG, Earley chart parser                                                                          |
| `tree`         | Tree (bracket parse, subtrees, productions, leaves)                                               |
| `chunk`        | RegexpParser (NP/VP IOB extraction)                                                               |
| `sentiment`    | VADER                                                                                             |
| `sem`          | FOL expression parser, model evaluation                                                           |
| `inference`    | Tableau prover, Resolution prover, Discourse                                                      |
| `cluster`      | K-means                                                                                           |
| `chat`         | Eliza-style chatbot                                                                               |
| `translate`    | BLEU score                                                                                        |

Not in Rust yet? Those calls fall through to NLTK automatically. Your code still works.

## Install

```bash
pip install fastnltk
```

Pre-built wheels for Linux (x86_64, aarch64), macOS (x86_64, arm64), Windows (x64).
Python 3.10–3.13.

Make sure you have the NLTK data you need:

```bash
python -m nltk.downloader punkt averaged_perceptron_tagger wordnet
```

## Usage

Everything lives under `fastnltk` with the same names and signatures as `nltk`.

```python
from fastnltk import word_tokenize, pos_tag, sent_tokenize

# Sentence segmentation (Punkt, Rust)
sents = sent_tokenize("Dr. Smith left at 5 p.m. He went home.")
# → ['Dr. Smith left at 5 p.m.', 'He went home.']

# Word tokenization (Treebank, Rust)
tokens = word_tokenize("The quick brown fox jumps over the lazy dog.")
# → ['The', 'quick', 'brown', 'fox', 'jumps', 'over', 'the', 'lazy', 'dog', '.']

# POS tagging (Perceptron, Rust)
tagged = pos_tag(tokens)
# → [('The', 'DT'), ('quick', 'JJ'), ...]
```

Drop it in as a direct NLTK replacement:

```python
import fastnltk as nltk
# All your existing nltk.* calls now run through Rust
nltk.word_tokenize("Hello, world!")
nltk.pos_tag(["Hello", "world"])
nltk.ne_chunk(nltk.pos_tag(["John", "lives", "in", "Boston"]))
```

For module-level imports:

```python
from fastnltk.stem import PorterStemmer, LancasterStemmer
from fastnltk.tag import PerceptronTagger
from fastnltk.parse import CFG, EarleyChartParser
from fastnltk.probability import FreqDist, ConditionalFreqDist
from fastnltk.lm import MLE, KneserNeyInterpolated
from fastnltk.metrics import edit_distance, jaccard_distance
from fastnltk.collocations import BigramCollocationFinder
from fastnltk.tree import Tree

# Same API as NLTK everywhere
stemmer = LancasterStemmer()
stemmer.stem("maximum")          # → 'maxim'
stemmer.stem("presumably")       # → 'presum'

fd = FreqDist("hello world")
fd["l"]                           # → 3
fd.max()                          # → 'l'

tagger = PerceptronTagger()
tagger.tag(["I", "love", "NLP"])  # → [('I', 'PRP'), ('love', 'VBP'), ('NLP', 'NNP')]

tree = Tree.from_string("(S (NP I/PRP) (VP love/VBP NLP/NNP))")
tree.leaves()                     # → ['I/PRP', 'love/VBP', 'NLP/NNP']
tree.productions()                # → ['S -> NP VP', 'NP -> I/PRP', 'VP -> love/VBP NLP/NNP']
```

## From source

```bash
git clone https://github.com/wyattferguson/fastnltk
cd fastnltk
pip install maturin
maturin develop --release
```

## Development

```bash
pip install -e ".[dev]"
maturin develop --release

cargo test          # Rust unit tests (309 pass)
pytest tests/       # 375 Python tests (368 pass, 6 skip, 1 xfail)

cargo fmt --all -- --check
cargo clippy --lib
ruff check fastnltk/ tests/
```

See [`CONTRIBUTING.md`](CONTRIBUTING.md) for full setup and PR workflow.

## Compatibility

The goal is 100% drop-in. Right now **368 of 375 tests pass** (6 skipped — chat bots
read stdin), with only **1 expected failure**:

- **CCG `fromstring`** — NLTK 3.10's `ccg.chart.fromstring` is broken (upstream bug)

Every critical-path API (tokenize, tag, stem, metrics, prob, parse, chunk, sentiment,
classify, collocations, tree, cluster, translate, chat) is verified byte-identical
to NLTK across all tested inputs.

## Platform

| Platform | Arch            | Wheel |
| -------- | --------------- | ----- |
| Linux    | x86_64, aarch64 | ✅    |
| macOS    | x86_64, arm64   | ✅    |
| Windows  | x64             | ✅    |

## License

[Apache 2.0](LICENSE). Not affiliated with NLTK or its maintainers.

## Contact + Support

Created by [Wyatt Ferguson](https://github.com/wyattferguson)

For any questions or comments heres how you can reach me:

**:octopus: Follow me on [Github @wyattferguson](https://github.com/wyattferguson)**

**:mailbox_with_mail: Email me at [wyattxdev@duck.com](wyattxdev@duck.com)**

**:tropical_drink: Follow on [BlueSky @wyattf](https://wyattf.bsky.social)**

