Metadata-Version: 2.5
Name: edgenote
Version: 1.0.0
Summary: Place important notes at the edges of your LLM prompts to fight the lost-in-the-middle problem.
Project-URL: LinkedIn, https://www.linkedin.com/in/md-tareq-shah-alam/
Author-email: Md Tareq Shah Alam <tareqshah.027@gmail.com>
License: Proprietary
License-File: LICENSE
Keywords: bm25,chunking,context-window,llm,lost-in-the-middle,prompt-engineering,prompt-optimization,rag,reranking,token-counting
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: License :: Other/Proprietary License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Python: >=3.10
Requires-Dist: regex>=2023.0
Requires-Dist: sentence-transformers>=3.0
Requires-Dist: tiktoken>=0.7
Requires-Dist: transformers>=4.40
Provides-Extra: all
Provides-Extra: chunker
Provides-Extra: dev
Requires-Dist: pytest>=8; extra == 'dev'
Provides-Extra: regex
Provides-Extra: rerank
Description-Content-Type: text/markdown

# edgenote

[![PyPI](https://img.shields.io/pypi/v/edgenote.svg)](https://pypi.org/project/edgenote/)
[![License](https://img.shields.io/badge/License-Proprietary-red.svg)]()
[![Tests](https://img.shields.io/badge/tests-220%20passed-brightgreen.svg)]()
[![LinkedIn](https://img.shields.io/badge/LinkedIn-Md%20Tareq%20Shah%20Alam-0A66C2?logo=linkedin&logoColor=white)](https://www.linkedin.com/in/md-tareq-shah-alam/)

> Prompt edge-pinning framework for LLMs. Mitigates the "Lost in the Middle" attention degradation problem by anchoring critical constraints and high-scoring context at prompt boundaries.

---

## The Problem

Transformer architectures exhibit U-shaped attention distributions ([Liu et al., 2023](https://arxiv.org/abs/2307.03172)). Information positioned at the **beginning (primacy)** and **end (recency)** of the context window is retrieved reliably, while critical data in the **middle** suffers severe recall degradation.

`edgenote` structures prompts defensively:
1. **Edge Pinning:** Anchors high-priority instructions, constraints, and facts to both the head and tail.
2. **U-Shaped Interleaving:** Organises ranked documents so that top-scoring chunks occupy the high-attention edges while lower-ranked data remains in the middle.
3. **Exact Token Budgeting:** Measures real token lengths and evicts or compresses low-priority middle context when limits are exceeded.

---

## Empirical Benchmark

Multi-document needle-in-a-haystack retrieval evaluation across context positions:

| Position in Context | Baseline Prompt | `edgenote` | Overhead |
| :--- | :---: | :---: | :---: |
| **Start** (0.0) | CORRECT | CORRECT | +44 tokens |
| **25%** | WRONG | CORRECT | +44 tokens |
| **Middle** (0.5) | WRONG | CORRECT | +44 tokens |
| **75%** | WRONG | CORRECT | +44 tokens |
| **End** (1.0) | CORRECT | CORRECT | +44 tokens |

---

## Installation

```bash
pip install edgenote
```

`edgenote` ships fully equipped with native tokenisation and neural ranking runtimes:

| Subsystem | Engine | Supported Models |
| :--- | :--- | :--- |
| **OpenAI Tokeniser** | `tiktoken` | `gpt-4o`, `o1`, `o3-mini`, `cl100k_base`, `o200k_base` |
| **Open-Weights Tokeniser** | `transformers` | Llama 3, Mistral, Gemma, Qwen, DeepSeek |
| **Neural Re-ranking** | `sentence-transformers` | Cross-encoder relevance scoring & bi-encoder embeddings |
| **Pattern Engine** | `regex` | Unicode-aware BPE tokenisation |

---

## Quick Start

```python
from edgenote import Session

s = Session()
s.pin("Hardware budget is strictly capped at $185,000.", label="Constraint")
s.add("Proposal A: Liquid cooling loop overhaul ($240,000)", relevance=0.7)
s.add("Proposal B: High-density compute cluster ($180,000)", relevance=0.92)

result = s.render("Which proposal satisfies our constraints?")
print(result.text)

messages = result.to_messages()
```

---

## Core Capabilities

### 1. Model-Accurate Token Counting

Pass any standard model identifier to bind the exact tokenizer backend:

```python
from edgenote import Session

s = Session(token_counter="gpt-4o")
s = Session(token_counter="meta-llama/Meta-Llama-3-8B-Instruct")
```

Custom counting callables are also accepted:

```python
import tiktoken
from edgenote import Session

enc = tiktoken.encoding_for_model("gpt-4o")
s = Session(token_counter=enc.encode)
```

### 2. Automated Re-ranking

Automatically score and reorder retrieved documents at render time without manual relevance labels:

```python
from edgenote import Session

s = Session(reranker="cross-encoder")
s.add("Quarterly financial filings")
s.add("Personnel roster and team structure")
s.add("Enterprise procurement guidelines")

result = s.render("What were the fourth-quarter operating expenditures?")
```

Pure-Python zero-overhead rankers are also available:

```python
from edgenote import Session
from edgenote.reranker import BM25Reranker, TFIDFReranker

s = Session(reranker=BM25Reranker())
s = Session(reranker=TFIDFReranker())
```

### 3. Context Compression

Summarize low-priority context when token limits are reached instead of dropping text:

```python
from edgenote import Session
from groq import Groq

client = Groq()

def summarize(text: str, max_tokens: int) -> str:
    response = client.chat.completions.create(
        model="llama3-8b-8192",
        messages=[
            {"role": "system", "content": f"Summarize in {max_tokens} tokens or fewer."},
            {"role": "user", "content": text},
        ],
        max_tokens=max_tokens,
    )
    return response.choices[0].message.content

s = Session(compressor=summarize)
s.add("Long technical specification...")
result = s.render("Summarize system latency limits", budget=2048)
```

### 4. Text Chunking

Split large documents before ingestion:

```python
from edgenote.chunker import StructuralChunker, SemanticChunker

chunker = StructuralChunker(max_tokens=512, overlap_tokens=64)
chunks = chunker.chunk(document_text)

semantic_chunker = SemanticChunker(threshold=0.5)
semantic_chunks = semantic_chunker.chunk(document_text)
```

### 5. Framework Integrations

#### LangChain

```python
from edgenote.integrations import from_langchain

result = from_langchain(
    docs,
    query="What is the operating budget?",
    pins=["Strictly cite sources using document headers."],
    budget=4000,
)
messages = result.to_messages()
```

#### LlamaIndex

```python
from edgenote.integrations import from_llamaindex

result = from_llamaindex(
    nodes,
    query="Synthesize quarterly performance metrics.",
    pins=["Output format: Markdown table."],
)
```

#### Generic Dictionaries

```python
from edgenote.integrations import from_dicts

docs = [
    {"text": "Annual recurring revenue reached $12M", "relevance": 0.95, "source": "Finance"},
    {"text": "Total headcount expanded to 120", "relevance": 0.40, "source": "HR"},
]
result = from_dicts(docs, query="Provide financial summary")
```

---

## CLI & Model Cache

`edgenote` provides a command-line interface for verification and pre-caching neural weights:

```bash
# Verify installation and active backends
edgenote

# Run live prompt assembly demonstration
edgenote --demo

# Pre-cache weights for air-gapped environments
edgenote download bge-small
```

---

## API Reference

### `Session` Parameters

| Parameter | Type | Default | Description |
| :--- | :--- | :--- | :--- |
| `style` | `Style` | `None` | Layout and section formatting configuration |
| `token_counter` | `Union[str, TokenCounter, Callable]` | `None` | Model name or tokenizer callable for token measurement |
| `orderer` | `Union[str, Callable]` | `"edges"` | Context layout strategy (`"edges"`, `"none"`, or custom) |
| `reranker` | `Union[str, Reranker, Callable]` | `None` | Document scoring engine (`"cross-encoder"`, `"embedding"`, etc.) |
| `compressor` | `Union[str, Compressor, Callable]` | `None` | Eviction or summarization strategy for budget fitting |
| `unranked_relevance`| `float` | `0.0` | Default score assigned to unranked chunks |

### `Rendered` Attributes

| Attribute | Type | Description |
| :--- | :--- | :--- |
| `text` / `user_text` | `str` | Fully assembled user prompt |
| `system` | `str` | System prompt text |
| `tokens` | `int` | Total measured token consumption |
| `pin_overhead` | `int` | Token cost incurred by tail-edge pin repetition |
| `dropped` | `List[str]` | Chunk texts evicted or summarized to fit budget |
| `to_messages()` | `List[Dict]` | Returns OpenAI-compatible payload `[{"role": ..., "content": ...}]` |

---

## Author

[**Md Tareq Shah Alam**](https://tareqshahalam.is-a.dev/)
- LinkedIn: [Md Tareq Shah Alam](https://www.linkedin.com/in/md-tareq-shah-alam/)
- Email: [tareqshah.027@gmail.com](mailto:tareqshah.027@gmail.com)
- License: Proprietary (All Rights Reserved)