⚡ Why Chunkr?
🏎️ 1,000+ MB/s Throughput
Built in memory-safe Rust with SIMD separator scanning and zero-copy string borrowing directly into Python. Operates 2x to 20x faster than LangChain and LlamaIndex.
🧠 18+ Chunking Strategies
From Recursive and Token BPE to Late Chunking, Tree-sitter AST code chunking, Table preservation, and Hierarchical Parent-Child trees.
📄 Native PDF Extraction
Built-in PDF parsing delivering 600+ pages/sec — 17x faster than pure-Python pypdf and 3x faster than PyMuPDF without external system dependencies.
🌉 Zero-Copy Ecosystem Bridges
Export directly to langchain.Document or llama_index.TextNode with zero runtime conversion penalty.
🥊 Head-to-Head Comparison
| Feature / Metric | Chunkr (chunkr-rs) | LangChain Splitters | LlamaIndex Node Parsers | Chonkie |
|---|---|---|---|---|
| Core Implementation | Rust + PyO3 (Native) | Pure Python | Pure Python | Python / Rust partial |
| Recursive Split Speed | 500 – 1,000+ MB/s | 200 – 350 MB/s | 150 – 300 MB/s | ~400 MB/s |
| Memory Model | Zero-Copy Slices | String Duplication | Object Churn | String Duplication |
| Multithreading | Rayon (True Multi-core) | ThreadPool (GIL bound) | Async / GIL bound | None |
| Late Chunking | Built-in | ❌ None | Experimental | ❌ None |
| AST Code Chunking | Tree-sitter (Rust & Python) | Regex-based | ❌ None | ❌ None |
| Native PDF Parsing | Built-in (600+ pgs/s) | External (pypdf) | External (pypdf) | ❌ None |
| Direct Adapters | LangChain, LlamaIndex, Pandas | Native | Native | Basic helper |
🐍 Python Quickstart
import chunkr
# 1. High-speed Recursive Character Chunking
chunker = chunkr.RecursiveChunker(chunk_size=500, overlap=50)
docs = chunker.chunk("Your long text or document...")
# 2. Hierarchical Parent-Child (Small-to-Big) Chunking
hier = chunkr.HierarchicalChunker(parent_size=1200, child_size=300)
pairs = hier.chunk_hierarchical("Your document...")
# 3. Late Chunking with Token Boundary Snapping
late = chunkr.LateChunker(chunk_size=300, overlap=30)
chunks = late.chunk("Your document...")
# 4. Zero-Copy Export to LangChain and LlamaIndex
langchain_docs = chunkr.to_langchain(docs)
llamaindex_nodes = chunkr.to_llamaindex(docs)
❓ Frequently Asked Questions
What is Chunkr and how does it differ from other text splitters?
Chunkr is an ultra-fast document chunking library written in Rust with PyO3 native bindings for Python. It replaces pure-Python splitters with a high-throughput, zero-copy engine offering 18+ chunking strategies, including Late Chunking, Tree-sitter AST code parsing, and native PDF loading.
Why is it installed as chunkr-rs but imported as chunkr?
On PyPI, the name
chunkr-rs is registered to prevent collisions with third-party OCR tools. When installed via pip install chunkr-rs, you import it directly in Python using import chunkr.Do I need Rust installed to use Chunkr in Python?
No. Pre-compiled native binary wheels are published to PyPI for Linux, Windows, and macOS (Intel and Apple Silicon).
pip install chunkr-rs installs pre-built binaries instantly.How does Late Chunking work in Chunkr?
Late Chunking preserves global document context across embedding vectors. The full text is passed through the transformer model first, and Chunkr's
LateChunker calculates token span boundaries to mean-pool chunk embeddings directly from full-document representations.