Metadata-Version: 2.3
Name: langchain-jtoken
Version: 0.1.0
Summary: Compress LangChain Document page_content with jtoken for fewer LLM prompt tokens
Project-URL: Homepage, https://github.com/HermannSamimi/jtoken
Project-URL: Repository, https://github.com/HermannSamimi/jtoken
Project-URL: Issues, https://github.com/HermannSamimi/jtoken/issues
Author-email: Hermann Samimi <hermannsamimi@gmail.com>
License: MIT
Keywords: json,langchain,llm,rag,token-optimization
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: MIT License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Software Development :: Libraries :: Python Modules
Requires-Python: >=3.10
Requires-Dist: jtoken>=0.3.5
Requires-Dist: langchain-core>=0.3
Description-Content-Type: text/markdown

# langchain-jtoken

A [LangChain](https://python.langchain.com) document transformer that compresses
JSON-shaped `Document.page_content` with [jtoken](https://github.com/HermannSamimi/jtoken)
so retrieved documents (Elasticsearch hits, MongoDB JSON, API responses) occupy
fewer tokens in the LLM context window — **losslessly**, so any agent can decode
them back with `jtoken.decode_document` when needed.

Measurable wins (tiktoken, cl100k_base): **−11%** Elasticsearch hits, **−19%**
MongoDB extended JSON, **−13%** nested API events on the
[reproducible benchmark](../../benchmarks/benchmark.py).

## Install

```bash
pip install langchain-jtoken
```

## Usage

```python
from langchain_jtoken import JSONTokenDocumentTransformer

transformer = JSONTokenDocumentTransformer()
compressed_docs = transformer.transform_documents(docs)
```

Non-JSON content passes through unchanged (`strict=True` raises instead).

## Test

```bash
pytest
```

MIT — © Hermann Samimi