Metadata-Version: 2.4
Name: foliodoc
Version: 0.1.0
Summary: Fast document conversion for Python: PDFs, scans, images and DOCX to Markdown, JSON or text, with tables, headings and reading order.
Author-email: Meet Jethwa <meetjethwa3@gmail.com>
License: MIT
Project-URL: Homepage, https://meet2147.github.io/foliodoc/
Project-URL: Source, https://github.com/Meet2147/foliodoc
Project-URL: Issues, https://github.com/Meet2147/foliodoc/issues
Keywords: pdf,ocr,document-conversion,markdown,tables,layout,docling,rag
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: MIT License
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Text Processing :: Markup
Classifier: Topic :: Scientific/Engineering :: Image Recognition
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: pypdfium2<6,>=4.20
Requires-Dist: pillow>=10
Requires-Dist: numpy>=1.24
Requires-Dist: opencv-python-headless>=4.8
Provides-Extra: ocr
Requires-Dist: rapidocr>=3.4; extra == "ocr"
Requires-Dist: onnxruntime>=1.17; extra == "ocr"
Requires-Dist: pyobjc-framework-Vision>=10; sys_platform == "darwin" and extra == "ocr"
Provides-Extra: apple
Requires-Dist: pyobjc-framework-Vision>=10; sys_platform == "darwin" and extra == "apple"
Provides-Extra: rapidocr
Requires-Dist: rapidocr>=3.4; extra == "rapidocr"
Requires-Dist: onnxruntime>=1.17; extra == "rapidocr"
Provides-Extra: tesseract
Requires-Dist: pytesseract>=0.3; extra == "tesseract"
Provides-Extra: bench
Requires-Dist: reportlab>=4; extra == "bench"
Requires-Dist: rapidfuzz>=3; extra == "bench"
Requires-Dist: apted>=1.0; extra == "bench"
Requires-Dist: lxml>=5; extra == "bench"
Requires-Dist: pyarrow>=14; extra == "bench"
Provides-Extra: dev
Requires-Dist: pytest; extra == "dev"
Requires-Dist: reportlab>=4; extra == "dev"
Dynamic: license-file

# foliodoc

Fast document conversion for Python (PDF / scans / images / DOCX → Markdown, JSON, text) on Windows, Linux and macOS. It solves the same problem as [docling](https://github.com/docling-project/docling) with no layout neural network and no GPU: exact text from the PDF text layer, OCR only where a page has none, and geometric reconstruction of reading order, headings, lists and tables. See [Benchmarks](#benchmarks) for where that wins and where it doesn't.

```python
from foliodoc import convert

doc = convert("report.pdf")        # or .png/.jpg/.tiff scan, or .docx
print(doc.to_markdown())
doc.tables                         # [[["Item", "Qty"], ["Apple", "3"]], ...]
doc.to_json()                      # blocks with type, text, bbox, page, heading level, cells
```

```bash
foliodoc report.pdf scan.png -o out/ --to md --timings
```

## Install

```bash
pip install foliodoc              # digital PDFs and Word files
pip install "foliodoc[ocr]"       # + OCR for scans/images: Apple Vision on macOS, RapidOCR on every OS
```

Python 3.10+ on Windows, Linux and macOS. Individual engines: `foliodoc[apple]`, `foliodoc[rapidocr]`, `foliodoc[tesseract]` (the last also needs the `tesseract` binary). Documentation: https://meet2147.github.io/foliodoc/

## How it gets its accuracy

| Problem | What foliodoc does |
|---|---|
| Most PDFs already contain the exact text | Reads the text layer through pdfium: every character comes with its box, font size and weight. **No OCR, 100% exact, about 10 ms per page.** OCR is used only for pages with no text layer or a broken font encoding, and for large embedded images that have no text drawn over them. |
| Scanned pages | `ocr="auto"` picks the best engine installed: **Apple Vision** on macOS (Neural Engine, on-device), **RapidOCR** (PaddleOCR v5 on ONNX Runtime, CPU) on Windows, Linux and macOS, then Tesseract. You can plug in your own engine with `register_ocr_engine`. Pages stream through a bounded pool of engines. |
| Skewed scans | Projection-profile deskew (coarse-to-fine search, about 30 ms), applied only when it clearly helps. |
| OCR skips lines or returns half a line | Every ink line on the page is checked against the OCR boxes. Lines that were skipped or only partly read are re-read from a tight crop, and the result replaces the old reading only if it is clearly more complete. It never adds duplicates. |
| OCR merges two table cells into one "line" | Lines are split wherever the *pixels* show a real gap that is wide compared with the line's own word spacing. Justification stretches all spaces in a line equally, so prose isn't split. Boxes are then snapped to the ink. |
| OCR gives no font information | Stroke thickness (average run length of ink pixels) separates bold and large headings from body text far more reliably than box height. |
| Reading order in multi-column layouts | Recursive XY-cut that refuses horizontal cuts through a multi-column zone. Tables and figures are treated as single units. |
| Tables | Ruled grids come from vector lines (or morphological line detection on scans). Rules-only "booktabs" tables are supported. Tables without rules are accepted only if the rows share column gutters and every column is aligned on its left edge, right edge or center, which prose never is. |
| Tables broken across pages or columns | Stitched back together, including a lone header row, repeated headers, and continuation rows. |
| Page furniture | Repeated headers and footers, matched fuzzily to tolerate OCR noise, plus page numbers are dropped from the body. |
| Hyphenation | "experi-\nment" → "experiment". |

## Benchmarks

Two benchmarks, both reproducible from `bench/`. Every number below comes from the final code. Both systems ran one after the other on an Apple M2, warm (model loading excluded), with default settings unless a row says otherwise. Docling is version 2.63.

### Summary

| | Where foliodoc wins | Where docling wins |
|---|---|---|
| **Speed** | Digital PDFs: **~20× faster** (0.08 vs 1.66 s/page). Scans with Apple Vision on both: **5–10× faster**. | Cross-platform OCR on small forms (FUNSD): docling 4.2 s vs foliodoc 6.7 s per form. |
| **Text accuracy** | Digital PDFs (slightly), scanned forms and receipts with Apple Vision, receipt key fields (total found 97% vs 77–93%). | Real page images from OmniDocBench (books, papers, slides), and forms with the cross-platform OCR. |
| **Structure** | Synthetic tables split across pages or columns. | **Real-world layout**: tables, headings, list items, header/footer removal (DocLayNet), and table structure (OmniDocBench). Docling's neural layout and table models are much better here. |
| **Robustness** | 1 failure in 7,822 documents. | 214 failures in 7,822: docling rejects large images (over Pillow's 179-megapixel "decompression bomb" limit after its internal upscaling). |

foliodoc is not more accurate than docling across the board. It is much faster, it reads text as well or better on digital PDFs, forms and receipts, and it loses clearly on layout and table structure in real documents.

### Real-world: 7,822 public documents

| Dataset | Docs | What it is | Ground truth |
|---|---|---|---|
| [DocLayNet v1.2](https://huggingface.co/datasets/docling-project/DocLayNet-v1.2) test | 4,999 | Digital PDF pages: financial reports, manuals, papers, laws, tenders, patents | Human layout labels + the page's text |
| [OmniDocBench](https://huggingface.co/datasets/opendatalab/OmniDocBench) | 1,651 | Page images: books, papers, slides, notes, newspapers (English and Chinese) | Reading-order text + table HTML |
| [SROIE 2019](https://huggingface.co/datasets/rth/sroie-2019-v2) | 973 | Scanned receipts | Text lines + company/date/address/total |
| [FUNSD](https://huggingface.co/datasets/nielsr/funsd) | 199 | Noisy scanned forms | Words |

**Default configuration** (on macOS both systems use Apple Vision for OCR; docling through `ocrmac`):

| Dataset | Metric | foliodoc | docling |
|---|---|---|---|
| DocLayNet (4,999 digital PDFs) | Text token F1 | **0.939** | 0.931 |
| | Table / heading / list-item detection F1 (IoU ≥ 0.5) | 0.35 / 0.41 / 0.36 | **0.86 / 0.86 / 0.87** |
| | Page headers/footers leaking into text ↓ | 32% | **7%** |
| | Seconds per page (mean / median) | **0.08 / 0.03** | 1.66 / 0.82 |
| OmniDocBench, English pages | Text edit distance ↓ (all 755 pages) | 0.288 | **0.282** |
| | Text edit distance ↓ (703 pages both processed) | 0.298 | **0.229** |
| | Table TEDS (703 pages both processed) | 0.343 | **0.493** |
| OmniDocBench, all 1,651 pages | Seconds per page | **0.99** | 7.94 |
| SROIE (973 receipts) | Token F1 | **0.691** | 0.595 |
| | Token F1 (840 receipts both processed) | **0.702** | 0.640 |
| | Total / date found (840 both processed) | **97.1% / 95.8%** | 89.5% / 89.3% |
| | Seconds per receipt | **0.41** | 4.01 |
| FUNSD (199 forms) | Token F1 | **0.899** | 0.856 |
| | Seconds per form | **0.46** | 2.26 |
| All | Failed documents | **1** | 214 |

**Cross-platform configuration.** Both systems use RapidOCR (PaddleOCR models on ONNX Runtime, CPU), the OCR they use on Windows and Linux. This ran on a subset: all FUNSD forms, the SROIE test split, and the English and mixed-language OmniDocBench pages that both finished (493). DocLayNet needs no OCR, so its results above apply on every OS.

| Dataset | Metric | foliodoc + RapidOCR | docling + RapidOCR |
|---|---|---|---|
| FUNSD (199) | Token F1 | 0.810 | **0.844** |
| | Seconds per form | 6.71 | **4.18** |
| SROIE test (347) | Token F1 | 0.601 | **0.620** |
| | Total / date found | **97.4%** / 91.6% | 93.4% / **91.9%** |
| | Seconds per receipt | **4.66** | 5.71 |
| OmniDocBench English + mixed (493) | Text edit distance ↓ | 0.324 | **0.296** |
| | Table TEDS | 0.461 | **0.697** |
| | Seconds per page | **8.56** | 11.81 |

Notes on the real-world numbers:

- 102 DocLayNet pages are excluded from the text metric for both systems because DocLayNet's own ground-truth text is garbled on them (PDFs with broken font encodings, e.g. `Ó Ç ä ä Ê`).
- Text inside figures is excluded on both sides. Docling's layout model was trained on DocLayNet, so DocLayNet is home ground for it.
- Chinese pages: both systems ran with default (Latin-script) OCR settings and score poorly on them. foliodoc accepts `languages=["zh-Hans", "en-US"]` for Apple Vision; that setting was not benchmarked.
- A failed document counts as empty output. Every document that crashed or hung was recorded as a failure; none were silently skipped.
- The cross-platform run was cut short to save time: OmniDocBench covers 493 of the planned 871 pages. Both systems are scored on exactly the same pages.
- Bugs found in foliodoc while running this benchmark were fixed before the final numbers: effective font size, word spacing in PDFs without space glyphs, letter-spaced and rotated text, page-border frames, content outside the visible page, a pdfium threading crash, a Vision stall, and an XY-cut infinite loop on mirrored glyphs.

Reproduce:

```bash
python bench/real_prep.py funsd sroie doclaynet omnidocbench    # after downloading the datasets to bench/real/
python bench/real_run.py foliodoc doclaynet                        # resumable; one JSON line per document
bench/.venv-docling/bin/python bench/real_run.py docling doclaynet
python bench/real_run.py folio-rapidocr funsd                   # cross-platform OCR configuration
python bench/real_score.py > bench/real/results.json
```

### Synthetic: exact ground truth

`bench/corpus.py` generates documents where every character, heading and table cell is known: reports, two-column papers, invoices and dense pages, each as a digital PDF, a clean 300 dpi scan, and a noisy 200 dpi scan (skew, blur, noise, JPEG). The table shows the held-out set (seed 1000, generated after tuning; 36 files, 78 pages).

| Input | System | Character error ↓ | Word F1 | Table structure | Table cells | Headings F1 | Sec/page ↓ |
|---|---|---|---|---|---|---|---|
| Digital PDF | **foliodoc** | **0.03%** | **1.000** | **0.917** | **0.972** | 1.000 | **0.011** |
| | docling | 9.70% | 0.987 | 0.528 | 0.858 | 1.000 | 1.65 |
| Clean scan (Apple Vision) | **foliodoc** | **0.38%** | **0.998** | **0.917** | **0.965** | 0.989 | **0.49** |
| | docling | 1.61% | 0.989 | 0.528 | 0.847 | 0.996 | 1.55 |
| Noisy scan (Apple Vision) | **foliodoc** | **0.17%** | **0.996** | **0.917** | **0.969** | 0.996 | **0.42** |
| | docling | 1.75% | 0.993 | 0.528 | 0.854 | 1.000 | 2.11 |
| Clean scan (RapidOCR) | **foliodoc** | **0.99%** | **0.989** | **0.917** | **0.958** | 0.945 | **4.7** |
| | docling | 7.70% | 0.928 | 0.528 | 0.827 | 0.986 | 5.9 |
| Noisy scan (RapidOCR) | **foliodoc** | **0.25%** | **0.994** | **0.917** | **0.939** | 0.985 | 5.8 |
| | docling | 8.00% | 0.946 | 0.528 | 0.815 | 1.000 | 5.8 |

On these synthetic documents docling's character error is high for two reasons confirmed by diffing: it reorders blocks in two-column layouts, and on justified text it cuts off the last word of some lines (`$94,118.11 → $94,11`). The synthetic set is narrower than real documents, which is why the real-world results above are the ones to trust.

## Layout

```
foliodoc/pdf.py       pdfium text layer (chars → segments with words), vector rulings, images, scan extraction
foliodoc/ocr.py       OCR engines behind one interface (apple / rapidocr / tesseract)
foliodoc/raster.py    deskew, rule detection, weak-line re-read, ink-gap splitting, stroke-based styling
foliodoc/layout.py    rows, ruled/unruled tables, XY-cut reading order, paragraphs, headings, lists, furniture, stitching
foliodoc/docx.py      DOCX straight from XML
foliodoc/convert.py   orchestration, streaming OCR pool
bench/             corpus generator, scorer, runner, per-document inspector
```
