Metadata-Version: 2.5
Name: hypertok
Version: 0.1.0
Summary: Atom-level token optimization for LLM context windows: structure-aware atomization, budgeted selection, and auditable compression.
Project-URL: Homepage, https://github.com/meet2147/hypertoken
Project-URL: Repository, https://github.com/meet2147/hypertoken
Project-URL: Issues, https://github.com/meet2147/hypertoken/issues
Author-email: Meet Jethwa <meetjethwa3@gmail.com>
License: Apache-2.0
License-File: LICENSE
Keywords: context-window,document-intelligence,evidence-graph,llm,nlp,prompt-compression,rag,retrieval,token-optimization,tokens
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: License :: OSI Approved :: Apache Software License
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Topic :: Text Processing :: Linguistic
Classifier: Typing :: Typed
Requires-Python: >=3.9
Provides-Extra: all
Requires-Dist: beautifulsoup4>=4.12; extra == 'all'
Requires-Dist: pypdf>=4.0; extra == 'all'
Requires-Dist: tiktoken>=0.5; extra == 'all'
Requires-Dist: transformers>=4.30; extra == 'all'
Provides-Extra: dev
Requires-Dist: mypy>=1.8; extra == 'dev'
Requires-Dist: pypdf>=4.0; extra == 'dev'
Requires-Dist: pytest-cov>=4.1; extra == 'dev'
Requires-Dist: pytest>=7.4; extra == 'dev'
Requires-Dist: reportlab>=4.0; extra == 'dev'
Requires-Dist: ruff>=0.5; extra == 'dev'
Provides-Extra: hf
Requires-Dist: transformers>=4.30; extra == 'hf'
Provides-Extra: html
Requires-Dist: beautifulsoup4>=4.12; extra == 'html'
Provides-Extra: pdf
Requires-Dist: pypdf>=4.0; extra == 'pdf'
Provides-Extra: tiktoken
Requires-Dist: tiktoken>=0.5; extra == 'tiktoken'
Description-Content-Type: text/markdown

# HyperToken

**Atom-level token optimization for LLM context windows.**

[![PyPI](https://img.shields.io/pypi/v/hypertok.svg)](https://pypi.org/project/hypertok/)
[![Python](https://img.shields.io/pypi/pyversions/hypertok.svg)](https://pypi.org/project/hypertok/)
[![License](https://img.shields.io/badge/license-Apache--2.0-blue.svg)](LICENSE)

```bash
pip install hypertok
```

```python
from hypertok import HyperToken
```

Zero required dependencies. Pure Python. Works offline.

> The project is **HyperToken**; the package is `hypertok`. The shorter name on
> PyPI belongs to an unrelated project that also installs a top-level
> `hypertoken/` module, so sharing it would let the two overwrite each other.

---

## The problem

Every context-window compressor works on **windows**: slide a fixed span across
the document, score it, keep the best ones. That is convenient and it throws
away the thing that makes a document a document. A 512-token window cuts a
function in half, separates a table row from its caption, and drops the heading
that told you what the paragraph was about.

HyperToken works one level down, on **atoms** — the smallest structural element
the source actually has:

| Source | An atom is |
|---|---|
| PDF / plain text | one heading, paragraph, table row, footnote, citation |
| Markdown | one heading, list item, fenced block, table row, quote |
| Source code | one function or method, with its scope recorded |
| JSON / API payload | one leaf value, at its full path |
| HTML | one block-level element |

…and one level down again, on **units** — the sentences inside an atom. The
optimizer selects atoms, then trims units inside the atoms it keeps. Nothing is
cut mid-structure, and every emitted token is traceable to where it came from.

## Results

`python benchmarks/bench.py` — 16 queries over 3 documents, measuring the
fraction of gold answer spans that survive compression:

| Budget | truncate (head) | truncate (tail) | fixed chunks + BM25 | **HyperToken** |
|---:|---:|---:|---:|---:|
| 80 tokens | 12.5% | 25.0% | 50.0% | **93.8%** |
| 150 tokens | 37.5% | 37.5% | 81.2% | **93.8%** |
| 300 tokens | 68.8% | 75.0% | 87.5% | **100.0%** |
| 600 tokens | 100.0% | 100.0% | 87.5% | **100.0%** |

At 300 tokens HyperToken retains every answer while emitting *fewer* tokens
(242 avg) than the chunk baseline (259 avg) that retains 87.5%. The budget is a
hard ceiling: across all presets and budgets, zero runs overshot.

## Quick start

```python
from hypertok import HyperToken

ht = HyperToken(preset="rag")

result = ht.optimize(
    open("contract.txt").read(),
    budget=2000,
    query="termination notice period",
)

print(result.text)     # the optimized context, ≤ 2000 tokens
print(result.stats)    # 18420 -> 1987 tokens (16433 saved, 89.2% reduction)
```

One-liner:

```python
from hypertok import optimize

text = optimize(document, budget=1000, query="what changed in v4.2?").text
```

No budget, no dropping — just squeeze out the waste:

```python
ht.compress(scraped_html).text     # every atom survives, each one smaller
```

Several documents into one shared budget:

```python
ht.optimize_many([doc_a, doc_b, doc_c], budget=4000, query=q, weights=[3, 1, 1])
```

## Command line

```bash
hypertok optimize report.pdf --budget 2000 --query "revenue by region" --stats
hypertok compress notes.md --preset aggressive -o small.md
hypertok atoms contract.txt -n 20        # see how a document decomposes
hypertok count report.pdf --tokenizer tiktoken
hypertok presets
```

```
hypertok · preset=rag · tokenizer=heuristic
  3,309 → 194 tokens   █·······················  94.1% saved
  budget 200 · within budget · headroom 6
  atoms: 8 kept, 1 compressed, 81 dropped
  by structure label:
    body                 2,562 →    111  █··············· (32 atoms)
    citation               332 →      0  ················ (15 atoms)
    heading                180 →     36  ███············· (27 atoms)
```

## How it works

Five stages. Each one is a replaceable module.

```
 source ──▶ atomize ──▶ graph ──▶ score ──▶ select ──▶ compress ──▶ render
             │            │         │          │           │
     structure-aware   reading   relevance   budgeted    sentence
        atoms with      order    density     knapsack    trimming
        type labels   headings   structure   + support   inside kept
                     references  redundancy   pull-in       atoms
```

**1 · Atomize.** Structure-aware decomposition. For markup the structure is
read off the syntax; for PDF-extracted text it is inferred with heuristics that
recover headings, captions, footnotes, citations, table rows and list items —
the label vocabulary from the SAGER corpus study. Hard-wrapped lines are
rejoined and words broken across a line break are healed, which is free tokens
before anything else runs.

**2 · Graph.** Atoms are linked by reading-order adjacency, heading
containment, and cross-references (`"see Table III"` → the Table III caption).

**3 · Score.** Four signals per atom: BM25 query relevance with conservative
stemming, information density, a structural prior per label, and a SimHash
near-duplicate penalty. Every component is exposed for audit.

**4 · Select.** A value-density greedy pass with three corrections that matter
more than exact knapsack packing:

- **Support pull-in** — admitting an atom admits its heading ancestors and
  referenced captions, so kept content stays legible.
- **Redundancy rejection** — an atom restating admitted content is skipped.
- **Marginal trimming** — the first atom that does not fit is compressed into
  the remaining space rather than dropped.

The plan is then rendered and re-measured; if the real string overshoots, the
lowest-value content is dropped and it is rendered again. `budget` is a
guarantee, not a hint.

**5 · Compress.** Lossless normalization always (Unicode folding, whitespace,
decoration). Then, by level, stock-phrase simplification, repeated-phrase
abbreviation, sentence trimming, and function-word pruning. Numbers, currency,
dates, percentages, URLs, emails, identifiers and code spans are protected at
every level, as are negations and modals (`not`, `must`, `shall`, `except`) —
deleting a "not" from a contract clause is not a token saving.

## Presets

| Preset | What it does |
|---|---|
| `safe` | Lossless only. Never rewords, never trims a sentence. For load-bearing text. |
| `balanced` | **Default.** Normalization, phrase simplification, sentence trimming. |
| `rag` | Relevance-heavy, strong cross-chunk dedup, page annotations for citations. |
| `summary` | No query attached; density and structure carry selection. |
| `code` | Never reformats code. Drops whole definitions, not lines. |
| `aggressive` | Maximum reduction that still reads as English. |
| `extreme` | Function-word pruning. Telegraphic; models read it, people mostly cannot. |

Every field is overridable:

```python
from hypertok import HyperToken, ScoreWeights, SelectorConfig

ht = HyperToken(
    preset="rag",
    weights=ScoreWeights(relevance=0.8, informativeness=0.1, coverage=0.1),
    selector=SelectorConfig(support=False, redundancy_threshold=0.75),
)
```

## The audit trail

Every run explains itself. This is the part that makes atom-level optimization
usable in production rather than just smaller.

```python
result = ht.optimize(doc, budget=1000, query="refund policy")

for record in result.dropped[:3]:
    print(record.type, record.tokens_before, record.reason)
# citation  41  not selected
# footnote  22  redundant with 8f2a1c04 (0.94)
# body     118  no budget remaining

result.manifest()   # full JSON: per-atom decision, cost, score, reason
```

```python
from hypertok.report import format_report, summarize
print(format_report(result))
summarize(result)["by_type"]   # tokens kept vs dropped, per structure label
```

## Tokenizers

The default is a dependency-free estimator calibrated against `cl100k_base` and
biased slightly high, so a plan that fits the estimate fits the real tokenizer.
For exact counts:

```bash
pip install 'hypertok[tiktoken]'   # or [hf], [pdf], [html], [all]
```

```python
HyperToken(tokenizer="tiktoken")            # cl100k_base
HyperToken(tokenizer="gpt-4o")              # model name → its encoding
HyperToken(tokenizer="hf:bert-base-uncased")
HyperToken(tokenizer=my_tokenizer)          # anything with .count(str) -> int
```

## Budgets that account for the whole prompt

A context window is shared. Budgeting only the evidence is why prompts overflow.

```python
from hypertok import Budget

budget = Budget(
    total=8000,
    reserve_output=1500,    # room for the model to answer
    reserve_system=400,
    reserve_history=1200,
)
ht.optimize(document, budget, query)   # gets 4900 tokens
```

## Working with the pieces

The stages are usable on their own.

```python
from hypertok import HyperToken, AtomType

ht = HyperToken()
doc = ht.atomize(open("paper.pdf"))          # or a path, string, or list of chunks

for atom in doc.of_type(AtomType.TABLE_ROW):
    print(atom.page, atom.text)

scored = ht.score(doc, query="revenue")
print(scored[0].explain())
# body#12 value=0.8421 (coverage=1.000, density=0.774, redundancy=0.000, ...)
```

```python
from hypertok import AtomGraph
graph = AtomGraph(doc)
graph.support(atom.id)      # what this atom needs to be legible
graph.descendants(head.id)  # everything under a heading
```

## Optional extras

| Extra | Adds |
|---|---|
| `hypertok[tiktoken]` | Exact OpenAI-family token counts |
| `hypertok[hf]` | Hugging Face tokenizer counts |
| `hypertok[pdf]` | PDF ingestion via `pypdf` |
| `hypertok[html]` | Robust HTML parsing via BeautifulSoup |
| `hypertok[all]` | All of the above |

PDF extraction is fault-tolerant by design: pages that fail are recorded in
`Document.meta["failed_pages"]` and the batch continues, the same posture that
let the SAGER corpus study finish 1,059 of 1,076 PDFs without aborting.

## Background

HyperToken's data model comes from **SAGER — Structure-Aware Graph Evidence
Retrieval for Document Intelligence** (Meet Jethwa), which processed 1,076 PDFs
into 492,530 evidence atoms and showed that an evidence-graph substrate beats
flat chunk retrieval on messy real-world corpora (MRR@10 0.167 → 0.812).

That work applied the substrate to *retrieval*. HyperToken applies it to
*budgeting* — the problem you hit next, once retrieval works and the results no
longer fit.

## Development

```bash
git clone https://github.com/meet2147/hypertoken
cd hypertoken
pip install -e '.[dev,all]'

pytest                          # test suite
ruff check src tests            # lint
python benchmarks/bench.py      # reproduce the results table
```

Point the corpus harness at a real directory to check ingestion robustness,
structure recovery, throughput and budget compliance across many documents:

```bash
python benchmarks/corpus_run.py ~/corpus --budgets 500 2000 --json report.json
```

It reports failures with reasons, flags documents whose atomization looks
wrong, and exits non-zero on any budget violation — so it works as a CI gate
over a fixture corpus.

## License

Apache-2.0. See [LICENSE](LICENSE).
