PyPI Python Versions Crates.io MIT License

chunkr⚡

Blazingly fast document & text chunking engine written in Rust with native Python bindings for LLMs, Vector DBs, and RAG pipelines.

$ pip install chunkr-rs
GitHub Repository PyPI Package

⚡ Why Chunkr?

🏎️ 1,000+ MB/s Throughput

Built in memory-safe Rust with SIMD separator scanning and zero-copy string borrowing directly into Python. Operates 2x to 20x faster than LangChain and LlamaIndex.

🧠 18+ Chunking Strategies

From Recursive and Token BPE to Late Chunking, Tree-sitter AST code chunking, Table preservation, and Hierarchical Parent-Child trees.

📄 Native PDF Extraction

Built-in PDF parsing delivering 600+ pages/sec — 17x faster than pure-Python pypdf and 3x faster than PyMuPDF without external system dependencies.

🌉 Zero-Copy Ecosystem Bridges

Export directly to langchain.Document or llama_index.TextNode with zero runtime conversion penalty.

🥊 Head-to-Head Comparison

Feature / Metric Chunkr (chunkr-rs) LangChain Splitters LlamaIndex Node Parsers Chonkie
Core Implementation Rust + PyO3 (Native) Pure Python Pure Python Python / Rust partial
Recursive Split Speed 500 – 1,000+ MB/s 200 – 350 MB/s 150 – 300 MB/s ~400 MB/s
Memory Model Zero-Copy Slices String Duplication Object Churn String Duplication
Multithreading Rayon (True Multi-core) ThreadPool (GIL bound) Async / GIL bound None
Late Chunking Built-in ❌ None Experimental ❌ None
AST Code Chunking Tree-sitter (Rust & Python) Regex-based ❌ None ❌ None
Native PDF Parsing Built-in (600+ pgs/s) External (pypdf) External (pypdf) ❌ None
Direct Adapters LangChain, LlamaIndex, Pandas Native Native Basic helper

🐍 Python Quickstart

import chunkr

# 1. High-speed Recursive Character Chunking
chunker = chunkr.RecursiveChunker(chunk_size=500, overlap=50)
docs = chunker.chunk("Your long text or document...")

# 2. Hierarchical Parent-Child (Small-to-Big) Chunking
hier = chunkr.HierarchicalChunker(parent_size=1200, child_size=300)
pairs = hier.chunk_hierarchical("Your document...")

# 3. Late Chunking with Token Boundary Snapping
late = chunkr.LateChunker(chunk_size=300, overlap=30)
chunks = late.chunk("Your document...")

# 4. Zero-Copy Export to LangChain and LlamaIndex
langchain_docs = chunkr.to_langchain(docs)
llamaindex_nodes = chunkr.to_llamaindex(docs)

❓ Frequently Asked Questions

What is Chunkr and how does it differ from other text splitters?
Chunkr is an ultra-fast document chunking library written in Rust with PyO3 native bindings for Python. It replaces pure-Python splitters with a high-throughput, zero-copy engine offering 18+ chunking strategies, including Late Chunking, Tree-sitter AST code parsing, and native PDF loading.
Why is it installed as chunkr-rs but imported as chunkr?
On PyPI, the name chunkr-rs is registered to prevent collisions with third-party OCR tools. When installed via pip install chunkr-rs, you import it directly in Python using import chunkr.
Do I need Rust installed to use Chunkr in Python?
No. Pre-compiled native binary wheels are published to PyPI for Linux, Windows, and macOS (Intel and Apple Silicon). pip install chunkr-rs installs pre-built binaries instantly.
How does Late Chunking work in Chunkr?
Late Chunking preserves global document context across embedding vectors. The full text is passed through the transformer model first, and Chunkr's LateChunker calculates token span boundaries to mean-pool chunk embeddings directly from full-document representations.