Metadata-Version: 2.5
Name: pdfschema
Version: 0.2.0
Summary: Extract named columns from PDF tables as structured rows — built for feeding LLMs data instead of whole documents.
Project-URL: Homepage, https://github.com/nullbite-coder/pdfschema
Project-URL: Issues, https://github.com/nullbite-coder/pdfschema/issues
Author-email: Kush Rawal <kushrawal00@gmail.com>
License: MIT License
        
        Copyright (c) 2026 Kush Rawal
        
        Permission is hereby granted, free of charge, to any person obtaining a copy
        of this software and associated documentation files (the "Software"), to deal
        in the Software without restriction, including without limitation the rights
        to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
        copies of the Software, and to permit persons to whom the Software is
        furnished to do so, subject to the following conditions:
        
        The above copyright notice and this permission notice shall be included in all
        copies or substantial portions of the Software.
        
        THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
        IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
        FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
        AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
        LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
        OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
        SOFTWARE.
License-File: LICENSE
Keywords: ai,context,csv,document-parsing,extraction,json,llm,pdf,pdfplumber,prompt,rag,schema,table,tokens
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: MIT License
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Topic :: Software Development :: Libraries :: Python Modules
Classifier: Topic :: Text Processing
Classifier: Topic :: Text Processing :: Markup
Classifier: Topic :: Utilities
Requires-Python: >=3.11
Requires-Dist: pdfplumber>=0.11
Provides-Extra: dev
Requires-Dist: build; extra == 'dev'
Requires-Dist: pytest>=8.0; extra == 'dev'
Requires-Dist: reportlab>=4.0; extra == 'dev'
Requires-Dist: twine; extra == 'dev'
Provides-Extra: pandas
Requires-Dist: pandas>=2.0; extra == 'pandas'
Description-Content-Type: text/markdown

# pdfschema

**Pull just the columns you need out of a PDF, so your LLM prompt carries data instead of a document.**

```console
pip install pdfschema
```

```python
import pdfschema

rows = pdfschema.extract_rows(
    "statement.pdf",
    schema=["Date", "Particulars", "Debit"],
)
```

```json
[
  {"Date": "01-04-2025", "Particulars": "PAYMENT MADE VIA UPI TO VPA MERCHANT7X2K@EXAMPLEPAY AGAINST REF 100000000001", "Debit": "390.00"}
]
```

<sub>Every example here uses **fictional** transaction data. The measurements
below are real, taken from a 125-page statement that is not published.</sub>

## The problem this solves

Feeding a whole PDF to a model is the expensive way to answer a narrow question.
A 125-page bank statement is ~44,000 tokens of prompt for every single call —
and most of it is page furniture, letterhead, disclaimers and columns the
question never touches. At scale that is the bill.

pdfschema lets you name the columns you actually need and get structured rows
back. You extract once, then filter, aggregate or sample in Python before
anything reaches the model.

### Measured on a real 125-page statement

All six columns, typed schema:

| What you send | Tokens (approx.) | vs. raw page text |
| --- | ---: | ---: |
| Raw extracted page text | 43,951 | — |
| `to_csv()` | 28,540 | **−35%** |
| `to_tson()` | 29,974 | **−32%** |
| `to_toon()` | 30,005 | **−32%** |
| `to_tson_columnar()` | 30,246 | **−31%** |
| `to_json()` | 50,577 | **+15%** ⚠️ |

Narrow the columns and it compounds: three columns as CSV is 21,565 tokens
(−51%), two columns 4,687 (−89%).

**Don't put JSON in a prompt.** An array of objects repeats every key on every
row, which on this document costs *more* than the raw text it replaced. Keep
`to_json()` for the code that consumes the rows afterwards — that is what it's
good at — and send one of the formats that names the fields once:

```python
table = pdfschema.extract("statement.pdf", schema={
    "Date": "date", "Particulars": "string", "Debit": "number",
})[0]

prompt = f"Transactions:\n{table.to_toon(name='transactions')}\n\nWhich merchant took the most?"
```

```
transactions[970]{Date,Particulars,Debit}:
  2025-04-01,PAYMENT MADE VIA UPI TO VPA MERCHANT7X2K@EXAMPLEPAY AGAINST REF …,390
  2025-04-02,PAYMENT RECEIVED VIA UPI FROM VPA J.DOE-1@EXAMPLEBANK FROM …,null
```

<sub>(`…` marks text elided for width.)</sub>

**The real win is narrowing.** Most of the saving comes from dropping columns
and boilerplate the question never needed — and from filtering rows in Python,
which no amount of prompt engineering over raw PDF text does reliably:

```python
april = [r for r in rows if r["Date"].month == 4]   # with schema={"Date": "date"}
```

## Why a schema, and not a grid detector

Most extractors look for the lines a page draws and infer columns from them.
Bank statements, invoices and most generated business PDFs draw **row bands but
no vertical rules** — so there is nothing to find, and those tools hand back one
column with every row collapsed into a single string.

pdfschema starts from the header. Your labels are located on the page and their
x positions *become* the column boundaries, so an unruled table reads exactly
like a ruled one. The same labels identify the table's continuation on the next
page, which is how a 125-page statement comes back as one table rather than 125
fragments.

## Modes

Six independent switches that compose freely. Pick what you care about, leave
the rest at their defaults.

| Mode | Values | Default | Set with | Controls |
| --- | --- | --- | --- | --- |
| **Extraction** | schema / discovery | discovery | `schema=` | Which tables come back, and what the keys are called |
| **Value** | `string`, `number`, `date` | `string` | mapping `schema=` | The Python type of each cell |
| **Output** | `json`, `csv`, `toon`, `tson`, `tson-columnar`, lists, DataFrame | — | `to_*()` method | The text or object you hand on |
| **Cell join** | `smart`, `space`, `newline` | `smart` | `join=` | How a cell wrapped over several lines is rebuilt |
| **Merge** | on / off | on | `merge=` | Whether a table split across pages returns as one |
| **Date reading** | day-first / month-first | day-first | `dayfirst=` | How `01/04/2025` is read |
| **Failure** | lenient / strict | lenient | `strict=` | Whether a cell that won't convert is `None` or an error |

`table.strategy` additionally *reports* how rows were found — `"ruled"` (the page
drew them), `"clustered"` (inferred from word positions) or `"mixed"`.

### Extraction mode 1 — schema

Name the columns; get exactly those keys back, in your order and your spelling.

```python
rows = pdfschema.extract_rows("statement.pdf", schema=["Date", "Balance"])
```

Matching is **normalised subset**: a table matches when its header supplies every
label you asked for, compared ignoring case and whitespace. Columns you didn't
ask for are dropped. `"transaction id"`, `"Transaction  ID"`, `"TRANSACTION ID"`
and `"TransactionID"` all find the same column.

**Headers that wrap across lines** — common in narrow ruled columns, where
`Application No` prints as `Application` / `No` — are read cell by cell from the
drawn grid, so you ask for the label as it reads:
`schema=["Application No", "Hypothecation Type"]`.

If no table carries your columns you get `SchemaNotFoundError` — naming the
unmatched labels and listing the headers that *do* exist — not an empty list your
pipeline silently passes along.

### Extraction mode 2 — discovery

Omit the schema and every table found comes back, keyed by the labels printed on
the page:

```python
for table in pdfschema.extract("statement.pdf"):
    print(table.header, table.n_rows, table.page_span)
# ['Date', 'Transaction ID', 'Particulars', 'Debit', 'Credit', 'Balance'] 970 1-125
```

Those labels are exactly what `schema=` accepts, so this is the first step when
onboarding an unfamiliar document.

### Value modes

| Value mode | Python type | `None` when |
| --- | --- | --- |
| `string` | `str` | never — an empty cell is `""` |
| `number` | `float` | the cell is empty or a lone dash |
| `date` | `datetime.date` | the cell is empty or a lone dash |

```python
rows = pdfschema.extract_rows("statement.pdf", schema={
    "Date": "date", "Debit": "number", "Particulars": "string",
})
rows[0]["Date"]   # datetime.date(2025, 4, 1)
rows[0]["Debit"]  # 390.0
```

`number` handles currency symbols, thousands separators, percentages and
accounting negatives like `(250.50)`. `date` tries a fixed list of formats.

### Other modes

```python
pdfschema.extract(path, schema, pages="1,3,5-7")   # page selection
pdfschema.extract(path, schema, merge=False)       # one table per page
pdfschema.extract(path, schema, join="newline")    # keep the PDF's line breaks
pdfschema.extract(path, schema, dayfirst=False)    # 01/04/2025 is 4 January
pdfschema.extract(path, schema, strict=True)       # raise instead of storing None
```

`join` decides how a cell wrapped over several lines is rebuilt. PDFs often chop
text every N characters regardless of word boundaries, so a naive join gives you
either `@EXAMPLEPAYAGAINST` or `VP A MERCHANT`; `"smart"` (the default) works the wrap
style out per column and rebuilds the original string.

## Output modes and types

Every extraction returns a `list[Table]`. The output mode is whichever method you
then call — nothing about the extraction changes, only the shape you hand on.

| Method | Returns | Use for |
| --- | --- | --- |
| `.rows` / `to_dicts()` | `list[dict]` | Python code |
| `to_json(indent=)` | `str` | APIs, files, another program. **Not prompts.** |
| `to_lists()` | `list[list]` | Spreadsheet writers, the `csv` module |
| `to_dataframe()` | `pandas.DataFrame` | Analysis, aggregation, joins |
| `to_csv()` | `str` | **Prompts (cheapest)**, spreadsheets, DB load |
| `to_toon(name=)` | `str` | **Prompts (most reliable)** — [TOON](https://github.com/toon-format/spec), adds a row count and field list for ~5% over CSV |
| `to_tson()` | `str` | [zenoaihq TSON](https://github.com/zenoaihq/tson) — single line, delimiter-based |
| `to_tson_columnar(name=)` | `str` | [tsonformat.com TSON](https://tsonformat.com/) — indentation-based, transposed |
| `to_prompt(fmt, **opts)` | `str` | The format comes from config or a flag |

`to_dataframe()` needs `pip install pdfschema[pandas]`; the rest are always
available. The two TSONs are unrelated formats that share a name — one
delimiter-based, one indentation-based — and pdfschema implements both rather
than picking for you.

**Picking a combination:**

| You want | Extraction | Value | Output |
| --- | --- | --- | --- |
| Rows in an LLM prompt, cheapest | schema | typed | `to_csv()` |
| Rows in a prompt, model must not miscount | schema | typed | `to_toon(name=…)` |
| Rows for your own code | schema | typed | `to_json()` / `.rows` |
| Analysis and aggregation | schema | typed | `to_dataframe()` |
| Onboarding an unfamiliar document | discovery | string | `.header` |
| Faithful transcription | schema | string | `to_csv()` |

One thing worth knowing: every prompt format must quote a string that would
otherwise read as a number, so an untyped `Debit` column emits `"390.00"` where
`schema={"Debit": "number"}` emits `390`. Typing the numeric columns makes the
prompt smaller *and* less ambiguous.

These formats are newer than CSV and how reliably a given model reads one varies
by model — benchmark against yours before switching a pipeline over.

## Command line

```console
$ pdfschema discover statement.pdf
[1] pages 1-125  rows 970  (ruled)
    schema: ["Date", "Transaction ID", "Particulars", "Debit", "Credit", "Balance"]

$ pdfschema extract statement.pdf --schema "Date:date,Debit:number" -o rows.json
$ pdfschema extract statement.pdf --schema "Date,Balance" --pages 1-5 --format csv
$ pdfschema extract statement.pdf --schema "Date:date" --format toon --name transactions
```

`--format` takes `json` (default), `csv`, `toon`, `tson`, `tson-columnar`. Also
`--pages`, `--name`, `--join`, `--no-merge`, `--monthfirst`, `--strict`,
`-o/--output`, `--indent`. `extract` exits 2 if the schema is not found.

## What you get

- **Deterministic and auditable.** No model in the loop, no temperature, no
  hallucinated cells. The same PDF gives the same rows every time — which is what
  you want in the layer *feeding* an LLM.
- **Wrapped cells rejoined correctly**, multi-page tables merged, multi-table
  pages handled.
- **Errors you can branch on.** `SchemaNotFoundError` carries `.unmatched` and
  `.headers_found`; everything deliberate derives from `PdfSchemaError`.

## Typical uses

- **RAG ingestion** — turn statements, invoices and reports into rows worth
  embedding, instead of embedding page images or raw text dumps.
- **Agent tools** — back a `get_transactions(start, end)` tool with a real
  extraction rather than a whole-document prompt.
- **Batch pipelines** — thousands of documents a day, where a 50% prompt
  reduction is the difference in the invoice.
- **Pre-flight validation** — check a document has the columns you expect before
  spending a model call on it.
- **Plain ETL** — no LLM anywhere; PDF to CSV, database or DataFrame.

## Requirements

Python 3.11+ and [pdfplumber](https://github.com/jsvine/pdfplumber). Optional:
`pdfschema[pandas]` for `to_dataframe()`.

## Documentation

Full documentation, limitations and examples:
**https://github.com/nullbite-coder/pdfschema**

MIT licensed.
