Metadata-Version: 2.4
Name: text2rel
Version: 0.5
Summary: An HTML cleaner customized for kinship ties extraction from large text corpra.
Requires-Python: >=3.10,<3.13
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Requires-Dist: beautifulsoup4 (>=4.12,<5.0)
Requires-Dist: numpy (==1.26.4)
Requires-Dist: openai
Requires-Dist: pandas (==2.2.2)
Requires-Dist: pymupdf (>=1.24,<2)
Requires-Dist: regex
Requires-Dist: tiktoken
Description-Content-Type: text/markdown

![Alt text](images/Kinship_Ties_Extraction.png)
![PyPI version](https://img.shields.io/pypi/v/html-cleaner-kinship)



This is the package implementation of the code for the kinship ties and relational information from large text corpora, such as genealogies, biographies, and historical dictionaries. 

### Important information: 
- For example package usage, see ```test_workflow.ipynb```
- For information about individual functions before the package implementation and release, see ```examples``` folder. 

### Expected input files

Each document should normally have these three files together in the same
folder under `documents_root`:

```text
0_GenelogiesAndBiographies/
├── Example genealogy/
│   ├── Example genealogy_mod.htm       # source HTML used for the inventory
│   ├── Example genealogy_mod.pdf       # PDF with a readable OCR text layer
│   └── Example genealogy_Original.pdf  # original PDF without OCR
└── Biographies/
    └── Example biography/
        ├── Example biography_mod.htm
        ├── Example biography_mod.pdf
        └── Example biography_Original.pdf
```

The filenames must share the same base name. The package uses the HTML for
inventory creation, prefers `*_mod.pdf` for page matching, and falls back to
`*_Original.pdf`, which can be OCRed with Tesseract when necessary.

### Current release: 
Current release includes inventory creation, setting text bounds, assigning font usage and selecting appropriate text chunks from the files (including restriction to "MainText" fonts and JumPJumP insertion). 

### Files structure: 

1. #### cleaner.py 
Contains the HTMLCleaner class and thus the main logic. 

2. #### utils.py
Contains helper functions (like JumPJumP insertion)

3. #### io.py
Contains the file processing logic, such as loading and saving the JSON inventory. 

4. #### page_matching.py
Adds one-based PDF `start_page` and `end_page` values to chunk metadata using
exact and fuzzy text matching. Missing PDFs are skipped with null page values
and explicit filename guidance.

### Add PDF page numbers

If you already created an `HTMLCleaner`, use `cleaner.add_pdf_pages()`. If you
only have an inventory JSON file, use `add_pages_to_inventory_file()` directly.

```python
summary = cleaner.add_pdf_pages(
    documents_root="../0_GenelogiesAndBiographies",
    output_path="exact_html_inventory_new_ids_cleaned_pages.json",
)
```

The default fuzzy threshold is `0.80`. The matcher prefers `*_mod.pdf`, falls
back to `_Original.pdf` when needed, and can use Tesseract for image-only
originals when Tesseract is installed.

