Metadata-Version: 2.5
Name: xyslice
Version: 0.1.0
Summary: Extract text, tables, and images from PDF and Office documents using XY cut layout analysis
Author-email: DanielViglione <dviglione02@gmail.com>
License-Expression: AGPL-3.0-or-later
License-File: LICENSE
Keywords: document-extraction,layout-analysis,pdf,xy-cut
Classifier: Development Status :: 3 - Alpha
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Text Processing
Requires-Python: >=3.12
Requires-Dist: openpyxl>=3.1.5
Requires-Dist: pymupdf>=1.26.0
Requires-Dist: python-docx>=1.2.0
Requires-Dist: python-pptx>=1.0.2
Provides-Extra: training
Requires-Dist: joblib>=1.5.1; extra == 'training'
Requires-Dist: pandas>=2.2.3; extra == 'training'
Requires-Dist: pyarrow>=19.0.1; extra == 'training'
Requires-Dist: scikit-learn>=1.7.0; extra == 'training'
Description-Content-Type: text/markdown

# Xyslice

Xyslice extracts text, tables, and images from PDF and Office documents. It uses recursive XY cut analysis to split document regions along whitespace gaps and assemble text in reading order.

The PyPI distribution is named `xyslice`. The Python package is named `peppermint`.

## Installation

Python 3.12 or newer is required.

```bash
pip install xyslice
```

## Extract text from a PDF

```python
from peppermint.extraction.pdfxycut import extract_blocks

blocks = list(extract_blocks("example.pdf"))
for block in blocks:
    if block["type"] == "text":
        print(block["page"], block["text"])
```

PDF blocks include their type, bounding box, and page number. Text blocks contain extracted text, table blocks contain cell rows, and image blocks contain Base64 image data. PDF page numbers start at 1.

## Supported formats

| Format | Extraction function | Return value |
| --- | --- | --- |
| PDF | `peppermint.extraction.pdfxycut.extract_blocks` | Iterator of blocks |
| Word DOCX | `peppermint.extraction.docxycut.extract_blocks` | List of page groups containing `blocks` |
| PowerPoint PPTX | `peppermint.extraction.pptxycut.extract_blocks` | Iterator of blocks with slide numbers in `page` |
| Excel XLSX and CSV | `peppermint.extraction.xlsxcut.extract_blocks` | Iterator of table and image blocks with sheet information |

Each function accepts a file path. For example, extract spreadsheet tables with:

```python
from peppermint.extraction.xlsxcut import extract_blocks

for block in extract_blocks("example.xlsx"):
    if block["type"] == "table":
        for row in block["rows"]:
            print(row)
```

## Layout features and training

The optional training dependencies provide pandas, a Parquet engine, scikit-learn, and model serialization:

```bash
pip install "xyslice[training]"
python -m peppermint.features.build example.pdf --out features.parquet
```

Feature extraction produces geometry, font, spacing, alignment, and border features. The training function in `peppermint.models.train_layout_classifier` expects a Parquet dataset containing a `label` column. No pretrained classifier is included.

## Limitations

Xyslice is an early release. Layout extraction uses heuristics, so results depend on document structure and formatting. Scanned PDF text requires OCR before extraction. DOCX and PPTX layout positions are estimated from document properties rather than rendered by Microsoft Office. Legacy binary Office formats such as DOC, PPT, and XLS are not supported.

## License

Xyslice is licensed under AGPL-3.0-or-later. The release includes the full license text. Its PDF extraction dependency, PyMuPDF, is offered under AGPL or commercial licensing; see the [PyMuPDF licensing documentation](https://pymupdf.readthedocs.io/en/latest/about.html#license-and-copyright).
