Metadata-Version: 2.4
Name: pagelens-ocr
Version: 0.1.0
Summary: CPU-only OCR for scanned documents (PDF, JPG, JPEG, PNG) that keeps headings, tables and layout
Author: Maruti L Sankannanavar
License-Expression: Apache-2.0
Keywords: ocr,pdf,scanned documents,layout,tables,cpu
Classifier: Programming Language :: Python :: 3
Classifier: Operating System :: OS Independent
Classifier: Topic :: Scientific/Engineering :: Image Recognition
Classifier: Topic :: Text Processing
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: rapidocr>=3.9
Requires-Dist: rapid-layout>=1.2
Requires-Dist: rapid-table>=3.0
Requires-Dist: onnxruntime>=1.20
Requires-Dist: openvino>=2025.0
Requires-Dist: opencv-python-headless>=4.10
Requires-Dist: numpy>=1.26
Requires-Dist: pypdfium2>=4.30
Dynamic: license-file

# pagelens-ocr

OCR for scanned documents that runs on a normal CPU - no GPU, no cloud. Give it a **PDF, JPG, JPEG or PNG**
(the type is detected automatically) and it prints the text of every page, keeping headings, side-by-side
fields and tables in place.

## Install

Python 3.10 or newer, then:

```
pip install pagelens-ocr
```

The first run downloads the OCR models once (~250 MB, needs internet); after that it works offline.

## Use from the command line

```
pagelens-ocr scan.pdf
pagelens-ocr photo.jpg
pagelens-ocr scan.pdf > scan.txt        # save the text to a file
```

Output - the file name and page number, then that page's text:

```
scan.pdf - Page 1
# INVOICE
Invoice No   : 1234                    Date : 01/01/2025
| Item     | Qty | Amount |
|----------|-----|--------|
| Paper A4 | 2   | 500.00 |

scan.pdf - Page 2
...
```

## Use from Python

```python
from pagelens_ocr import extract, extract_text

for page in extract("scan.pdf"):          # one result per page
    print(page["file"], page["page"], page["seconds"])
    print(page["text"])

text = extract_text("photo.jpg")          # all pages as one string
```

## What it does

Per page: crop scanner borders, straighten tilted scans, remove punch holes -> detect and read every text line
(PP-OCRv6; a fast model reads every line, a stronger one re-reads uncertain lines) -> find headings, text blocks,
tables and logos (PP-DocLayout v2) -> rebuild table rows and columns (SLANet+) -> print in reading order.
Typically 4-9 seconds per page on a laptop CPU.

Handwriting is not supported. Always check important values (numbers, dates, IDs) against the original.

## License

Apache-2.0
