Metadata-Version: 2.4
Name: text2rel
Version: 1.0.1
Summary: An HTML cleaner customized for kinship ties extraction from large text corpra.
Requires-Python: >=3.10,<3.13
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Provides-Extra: plots
Requires-Dist: beautifulsoup4 (>=4.12,<5.0)
Requires-Dist: jellyfish
Requires-Dist: matplotlib ; extra == "plots"
Requires-Dist: numpy (==1.26.4)
Requires-Dist: openai
Requires-Dist: pandas (==2.2.2)
Requires-Dist: pymupdf (>=1.24,<2)
Requires-Dist: rapidfuzz
Requires-Dist: regex
Requires-Dist: scipy ; extra == "plots"
Requires-Dist: tiktoken
Description-Content-Type: text/markdown

# text2rel

[![PyPI version](https://img.shields.io/pypi/v/text2rel.svg)](https://pypi.org/project/text2rel/)
[![Python versions](https://img.shields.io/pypi/pyversions/text2rel.svg)](https://pypi.org/project/text2rel/)
[![PyPI status](https://img.shields.io/pypi/status/text2rel.svg)](https://pypi.org/project/text2rel/)
[![PyPI downloads](https://img.shields.io/pypi/dm/text2rel.svg)](https://pypistats.org/packages/text2rel)
[![Source](https://img.shields.io/badge/source-GitHub-181717.svg?logo=github)](https://github.com/EliteKinshipNetworks/Code)

```bash
pip install text2rel
```

[![Kinship ties extraction](https://raw.githubusercontent.com/EliteKinshipNetworks/Code/Package_setup/Package/images/Kinship_Ties_Extraction.png)](https://github.com/EliteKinshipNetworks/Code)



This is the package implementation of the code for the kinship ties and relational information from large text corpora, such as genealogies, biographies, and historical dictionaries. 

### Important information: 
- For example package usage, see ```test_workflow.ipynb```
- For information about individual functions before the package implementation and release, see ```examples``` folder. 

### Expected input files

Each document should normally have these three files together in the same
folder under `documents_root`:

```text
0_GenelogiesAndBiographies/
├── Example genealogy/
│   ├── Example genealogy_mod.htm       # source HTML used for the inventory
│   ├── Example genealogy_mod.pdf       # PDF with a readable OCR text layer
│   └── Example genealogy_Original.pdf  # original PDF without OCR
└── Biographies/
    └── Example biography/
        ├── Example biography_mod.htm
        ├── Example biography_mod.pdf
        └── Example biography_Original.pdf
```

The filenames must share the same base name. The package uses the HTML for
inventory creation, prefers `*_mod.pdf` for page matching, and falls back to
`*_Original.pdf`, which can be OCRed with Tesseract when necessary.

### Current release: 
Current release includes inventory creation, setting text bounds, assigning font usage and selecting appropriate text chunks from the files (including restriction to "MainText" fonts and JumPJumP insertion). 

### Files structure: 

1. #### cleaner.py 
Contains the HTMLCleaner class and thus the main logic. 

2. #### utils.py
Contains helper functions (like JumPJumP insertion)

3. #### io.py
Contains the file processing logic, such as loading and saving the JSON inventory. 

4. #### page_matching.py
Adds one-based PDF `start_page` and `end_page` values to chunk metadata using
exact and fuzzy text matching. Missing PDFs are skipped with null page values
and explicit filename guidance.

### Add PDF page numbers

If you already created an `HTMLCleaner`, use `cleaner.add_pdf_pages()`. If you
only have an inventory JSON file, use `add_pages_to_inventory_file()` directly.

```python
summary = cleaner.add_pdf_pages(
    documents_root="../0_GenelogiesAndBiographies",
    output_path="exact_html_inventory_new_ids_cleaned_pages.json",
)
```

The default fuzzy threshold is `0.80`. The matcher prefers `*_mod.pdf`, falls
back to `_Original.pdf` when needed, and can use Tesseract for image-only
originals when Tesseract is installed.

### Human Labeler Allocation

When text chunks are allocated to labelers, each generated TXT filename
indicates its task type:

- `Task_0` contains chunks shared across all labelers.
- `Task_1` through `Task_9` contain labeler-specific, unshared chunks.

### New section 4 assessment workflow

The reusable assessment follows sections 4.2.2.1 and 4.2.2.2 of
`240426 Information extraction system 260801.ipynb`. Install the optional plot
dependencies, load the processed human and AI tables, and then create the
assessment tables before drawing anything:

```bash
pip install -e ".[plots]"
```

```python
from text2rel import (
    ReliabilityAssessment,
    build_sweep_report,
    plot_sweep_report,
    plot_threshold_curves,
)

assessment = ReliabilityAssessment("inventory.json")
assessment.load_relations_data(relations_csv="human_relations.csv")
assessment.load_chatgpt_relations("machine_relations.csv")

# one_to_one is the package's stricter default for headline metrics.
relation_results = assessment.assess_llm_against_humans(data="relations")
plot_threshold_curves(relation_results["threshold_metrics"], direction="one_to_one")

# directional_best reproduces the notebook's two reusable-candidate views.
diagnostics = assessment.assess_llm_against_humans(
    data="relations", matching_mode="directional_best"
)
fn_report = build_sweep_report(
    diagnostics["human_to_llm_best_pairs"], data="relations"
)
plot_sweep_report(fn_report)
```

`assess_llm_against_humans()` returns the selected pairs, threshold metrics,
false-negative and false-positive review tables, structural exclusions, and
configuration metadata. Plot functions only render these returned tables; they
do not change matching or scores. See the separate **New section 4 assessment
plots and tables** section at the end of `test_workflow.ipynb` for relations
and events, notebook-style titles, confidence/length diagnostics, and review
summaries.

### Convert downloaded OpenAI batches to CSV

Use the submitted cleaned inventory (or the exact submitted subset) to restore
chunk IDs. The function accepts a JSONL path or `download_batch_results()` output:

CSV saving is automatic: omitted or `None` `output_csv` writes `llm_relations.csv`
or `llm_events.csv` in the current working directory. Pass a path to choose another
destination. Reruns overwrite the destination; both DataFrames are still returned.

```python
from text2rel import process_batch_results

relations, relation_audit = process_batch_results(
    "batch_output_relations.jsonl", submitted_inventory,
    extraction_type="relations", output_csv="llm_relations.csv",
)
events, event_audit = process_batch_results(
    "batch_output_events.jsonl", submitted_inventory,
    extraction_type="events", output_csv="llm_events.csv",
)
assessment.load_llm_relations("llm_relations.csv")
assessment.load_llm_events("llm_events.csv")
```

Taxonomy can be applied to the returned tables before assessment. These are the
ungrouped `load_llm_*` schemas, not historical `load_chatgpt_*` grouped CSVs.
Errors or missing responses raise by default; pass `strict=False` to inspect
partial results and their audit. Keep the audit, which also records successful
empty extractions. Custom schemas must retain `s/r/t` or `S/V`; extra fields
are preserved. For historical `req_N` IDs, pass an explicit mapping from each
request ID to its original chunk metadata (including `chunk_id`).

### Download Labelbox annotations to human CSVs

Use `download_labelbox_results(project_ids, allocation, tasks_dir, output_dir)`
with `project_ids={"relations": ["project_id"], "events": ["project_id"]}`.
It reuses the Done-only exporter and saves `human_relations.csv`,
`human_events.csv`, and coverage CSVs. Each returned dictionary value is a
`(facts, coverage)` pair. Supply `api_key` and optionally `author_map`.

Keep the original allocation JSON and uploaded TXT files unchanged: offsets in
the TXT locate each annotation; allocation row IDs recover original chunk IDs.
For manual JSON/JSONL exports use
`process_labelbox_results(exports, allocation, tasks_dir, "relations", output_csv)`
(or `"events"`). Ambiguous mappings raise before saving that CSV.
Coverage records empty completed annotations separately; the current assessment
does not automatically score these empty chunks. Downloads include only Done
tasks and do not establish that every assigned task has been completed.

### Labelbox event construction modes

`labelbox_events_to_dataframe(..., construction_mode="graph")` is the default.
It follows section 2.4.3.2 of the 260801 notebook: connected annotations are
rebuilt around one Verb, safe reversed arrows are corrected, and incomplete or
ambiguous structures remain auditable. Use `construction_mode="legacy_simple"`
only when explicitly reproducing the package's earlier permissive conversion.
Relation conversion uses its audited endpoint joins separately.

