Metadata-Version: 2.4
Name: legere
Version: 0.11.0
Summary: Compact CRNN OCR for printed text (Portuguese charset) — pure PyTorch, CPU-friendly, ~360k parameters
Author-email: Rafael Cezar <imrafaelcezar@gmail.com>
License-Expression: MIT
Project-URL: Homepage, https://github.com/imrafaelcezar/legere
Project-URL: Repository, https://github.com/imrafaelcezar/legere
Project-URL: Issues, https://github.com/imrafaelcezar/legere/issues
Keywords: ocr,crnn,ctc,pytorch,computer-vision,portuguese,text-recognition
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Scientific/Engineering :: Image Recognition
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: torch>=2.0
Requires-Dist: numpy>=1.24
Requires-Dist: opencv-python-headless>=4.8
Requires-Dist: pillow>=10.0
Provides-Extra: pdf
Requires-Dist: pypdfium2>=4.28; extra == "pdf"
Provides-Extra: onnx
Requires-Dist: onnx>=1.15; extra == "onnx"
Requires-Dist: onnxruntime>=1.17; extra == "onnx"
Provides-Extra: train
Requires-Dist: rich>=13.0; extra == "train"
Requires-Dist: psutil>=5.9; extra == "train"
Dynamic: license-file

# Legere

**A compact CRNN OCR engine for printed text, built from scratch in pure PyTorch and trained entirely on synthetic Portuguese-language pages. Optimized for CPU inference, with no pretrained models and no external OCR libraries.**

Legere reads full document pages — images or PDFs — and returns text in
reading order with per-line bounding boxes. It's built from scratch around a
single small recognition network (CompactCRNN, ~360k parameters, hard-capped
at 1M) paired with a classic, non-neural preprocessing pipeline: deskew,
column detection and line segmentation all run on plain image projections,
no learned layout model involved. A second tiny network (LineClassifier,
~8k parameters) filters out non-text ink — barcodes, QR blocks, stamps,
rules — before it ever reaches the recognizer.

Nothing in the training data is real: every page, line, word, stain, camera
blur and printer artifact is generated on the fly by a synthetic page
generator, so there is no dataset to license, no PII to scrub and no
copyright to worry about. The trained weights for both networks ship
inside the package (about 1.4 MB total) — `pip install legere` gets you a
working OCR engine with no downloads, no external services and no GPU
requirement. Inference runs on PyTorch eager, TorchScript (FP32/INT8) or
ONNX Runtime, whichever fits your deployment.

The project also ships the full machinery that produced those weights: the
synthetic data generator (multiple page styles, dozens of degradations,
~90 fonts), a training loop with a live dashboard, EMA and curriculum
support, export tooling for every backend, and a stage-by-stage profiler —
so the model is not just usable but reproducible and extensible for anyone
who wants to retrain it on their own document style.

- [Installation](#installation)
- [Quickstart](#quickstart)
- [Architecture](#architecture)
  - [Recognition model — CompactCRNN](#recognition-model--compactcrnn)
  - [Line filter model — LineClassifier](#line-filter-model--lineclassifier)
  - [Character set](#character-set)
  - [Inference pipeline](#inference-pipeline)
  - [Inference backends](#inference-backends)
  - [Synthetic data generator](#synthetic-data-generator)
  - [Training system](#training-system)
  - [File formats](#file-formats)
- [CLI reference](#cli-reference)
  - [`legere read`](#legere-read)
  - [`legere pdf`](#legere-pdf)
  - [`legere train`](#legere-train)
  - [`legere export`](#legere-export)
  - [`legere benchmark`](#legere-benchmark)
  - [`legere detector`](#legere-detector)
- [Python API](#python-api)
- [Metrics](#metrics)
- [Project layout](#project-layout)
- [Development](#development)

---

## Installation

```bash
pip install legere              # inference only (torch, numpy, opencv, pillow)
pip install legere[pdf]         # + PDF input (pypdfium2)
pip install legere[onnx]        # + ONNX Runtime backend
pip install legere[train]       # + training extras (rich, psutil)
```

Python ≥ 3.10. CPU is the primary target; CUDA is used for training when
available. Training currently expects Windows (fonts are read from
`C:\Windows\Fonts`); inference runs anywhere.

## Quickstart

CLI:

```bash
legere document.pdf                     # OCR every page, text on stdout
legere read page.png --json out.json    # + per-line text and bboxes
```

Python:

```python
from legere import Legere

ocr = Legere()                      # bundled weights, CPU
result = ocr.read("page.png")       # or a numpy array (grayscale or BGR)
print(result.text)                  # full text in reading order
for line in result.lines:
    print(line.text, line.bbox)     # (x0, y0, x1, y1) on the deskewed page
```

---

## Architecture

```
image / PDF page
      │  pdf.py (pypdfium2 render, optional)
      ▼
fit to 1920×1080 ─▶ deskew ─▶ column detection ─▶ line segmentation
      (resize)   (projection   (vertical          (horizontal projection
                  variance)     whitespace         + rule removal,
                                gutters)           per column)
      ▼
LineClassifier filter ─▶ height-32 normalization ─▶ width-bucketed batches
 (text / non-text gate)                                    │
      ▼                                                    ▼
text in reading order ◀─ CTC decode ◀─ CompactCRNN (CNN + BiGRU + linear)
```

### Recognition model — CompactCRNN

`legere/model.py` · **359,866 parameters** (hard limit of 1,000,000 enforced
at construction). Input: grayscale line crops `(B, 1, 32, W)` normalized to
`[-1, 1]`, W variable. Output: logits `(B, W/4, 105)`.

| Stage | Layer | Output shape (H×W×C) |
|---|---|---|
| Stem | Conv3×3 s2 + BN + ReLU | 16 × W/2 × 24 |
| Block 1 | DWConv3×3 s(2,2) + PWConv1×1 | 8 × W/4 × 48 |
| Block 2 | DWConv3×3 s(1,1) + PWConv1×1 | 8 × W/4 × 64 |
| Block 3 | DWConv3×3 s(2,1) + PWConv1×1 | 4 × W/4 × 96 |
| Block 4 | DWConv3×3 s(1,1) + PWConv1×1 | 4 × W/4 × 128 |
| Block 5 | DWConv3×3 s(2,1) + PWConv1×1 | 2 × W/4 × 160 |
| Block 6 | DWConv3×3 s(2,1) + PWConv1×1 | 1 × W/4 × 192 |
| Sequence | squeeze H → BiGRU(hidden 128, 1 layer) | W/4 × 256 |
| Head | Linear 256 → 105 | W/4 × 105 |

Every depthwise-separable block is `DWConv + BN + ReLU + PWConv + BN + ReLU`.
Height is fully collapsed 32 → 1; width is downsampled ×4, so the CTC gets
**one timestep per 4 input pixels**. Decoding: greedy best-path (default) or
prefix beam search (pure Python, `beam_width` ≥ 1, per-frame log-prob pruning
at −9.0).

### Line filter model — LineClassifier

`legere/model.py` · ~8k parameters. A binary text/non-text classifier over
the same normalized crops: stem Conv3×3 s2 (16ch) → 3 depthwise-separable
blocks (16→32→48→64, all stride 2×2) → global average pool → Linear(64→1).
It rejects non-text segments (barcodes, QR blocks, rules, doodles, stamps)
before recognition. Bundled at `legere/data/lineclf.pt`; decision threshold
**0.1** (deliberately conservative: keeps ≈98.5% of real text while
rejecting ≈75% of distractors). Retrain with [`legere detector`](#legere-detector).

### Character set

`legere/charset.py` · 104 characters + CTC blank (index 0) = **105 classes**:

```
a-z  A-Z  0-9  áàâãçéêíóôõúü  ÁÀÂÃÇÉÊÍÓÔÕÚÜ
espaço . , ; : ! ? ' " ( ) [ ] { } % $ # @ & * + - = / \ _ < > | ° º ª
```

`encode_text` / `decode_indices` map text ↔ class indices; `sanitize_text`
drops anything outside the charset (used by the data generator so ground
truth is always encodable).

### Inference pipeline

`legere/inference.py` + `legere/segment.py`. All classic stages operate on
the grayscale page; **no neural network is involved until recognition**.

1. **Fit to Full HD** — pages larger than 1920×1080 are downscaled
   (`INTER_AREA`), aspect preserved. Smaller pages pass through.
2. **Deskew** — searches the rotation angle (±4°, coarse 0.5° then fine
   0.1° steps) that maximizes the variance of the horizontal ink-projection
   profile, computed on a ≤800px-wide binarized copy. Rotation is applied
   only when |angle| ≥ 0.05°.
3. **Column detection** — on the binarized, rule-free page, a *gutter* is a
   run of near-empty pixel columns (≤0.2% of height inked) at least
   `max(24 px, 1.5% of width)` wide, strictly inside the inked region. Text
   crossing the middle (centered titles, section bars) breaks the run, so
   single-column forms are never split. Each region must hold ≥8% of the
   page ink; otherwise the page stays single-column. Reading order:
   left-to-right columns, top-to-bottom lines.
4. **Line segmentation** (per column) — adaptive Gaussian threshold
   (block 31, C 15, ink=white); long horizontal/vertical rules (table
   borders, underlines) removed by morphological opening (kernels
   `max(25, dim/20)`) and painted out of the grayscale crop source; rows
   with ≥2 ink pixels form bands; bands separated by ≤2 px merge (accents);
   bands <6 px tall are dropped; x-extent refined by the band's vertical
   projection; 3 px margin around each crop.
5. **Line filter** — LineClassifier scores every candidate crop; segments
   below threshold are dropped (`--no-line-filter` disables).
6. **Normalization & batching** — crops resized to height 32 (width capped
   at 1280 px ≈ 320 timesteps), padded with white to the batch max width
   (multiple of 4), scaled to [-1, 1]. Batches are width-bucketed under a
   budget of 48×384 padded pixels (≤64 lines per batch) so tensor size stays
   bounded.
7. **Recognition & decode** — CRNN forward under `torch.inference_mode()`,
   CTC greedy (or beam) decode. With `min_confidence > 0`, lines whose mean
   per-frame max-softmax falls below the threshold are dropped (second,
   recognizer-side distractor gate).

### Inference backends

`load_model(path, torchscript=...)` resolves the backend:

| Backend | How | Notes |
|---|---|---|
| PyTorch eager FP32 | default; `path=None` uses bundled weights | reference implementation |
| TorchScript FP32 | `--torchscript` + `exports/model_fp32.ts.pt` | scripted + frozen |
| TorchScript INT8 | `--torchscript` + `exports/model_int8.ts.pt` | dynamic quantization of GRU + Linear |
| ONNX Runtime | any `*.onnx` path (auto-detected) | dynamic batch/width axes, opset 17; typically the fastest CPU backend |

### Synthetic data generator

`legere/training/pagegen.py`. Infinite, deterministic-per-seed page stream;
every rendered line returns its exact text + ink bbox, so training crops are
extracted directly (no annotation, no leakage of real data).

**Page styles** (`style=` / `--style`): `document` (headers, paragraphs,
key/value lines, tables, optionally two-column), `report` (bordered
label/value form grids with gray/inverted section bars, measurement values,
gray bold footer — modeled after real technical reports), `mixed` (50/50
per page).

**Content**: ~65% real Portuguese words, ~35% syllable-built pseudo-words
(prevents the BiGRU from memorizing a lexicon), plus structured formats:
dates, money (R$), CPF/CNPJ, coordinates, percentages, measurements
(`NN.NN mm`, `N.NN º`), serials, dotted UIDs, timestamps, phone/CEP/email.
~18% of sentences are punctuation-dense (`08:45:12`, `1.234,56`,
`a, b; c: d.`) to teach the `.,;:` distinctions.

**Photometric degradations** (page-level, scaled by `severity` ∈ [0,1] —
ink never moves, so GT bboxes stay exact): paper texture + fiber grain,
irregular illumination (linear gradient + radial vignette), bleed-through
(blurred mirrored page ghost), stains and coffee rings, printer toner
banding / low-toner fade / toner dropout speckle, scanner streaks, Gaussian
noise, blur, contrast/brightness jitter, JPEG artifacts, and a phone-photo
preset combining the harsh variants.

**Geometric augmentations** (line-crop level, training only): rotation
±3°, perspective quad jitter, vertical sine warp, SpecAugment-style
occlusion bars.

**Structural variety**: inverted text (light on dark bars), underline /
strikethrough / highlight decorations, variable letter tracking
(0.88–1.30×), two-column layouts, barcode/QR distractors (no GT), and
overlapping stamps / signature squiggles / scribbles. ~90 Windows fonts.

### Training system

`legere/training/train.py` · `legere/training/dataset.py`.

- **Loss**: CTC (`blank=0`, `zero_infinity`); grad-norm clip 5.0.
- **Optimizer**: AdamW — weight decay 0.01 on ≥2-D params only; peak LR
  2e-3 (from scratch) or 5e-4 (`--finetune`); linear warmup (500 steps) then
  cosine decay to a 2% floor.
- **Batching**: width-bucketed under `--width-budget` (default 24×384 = 9216
  padded px per batch; max batch `max(64, budget/144)` lines); a 384-sample
  buffer is sorted by width and re-chunked so padding waste stays low.
- **EMA** (default on): a shadow copy of the model is updated every step
  (`ema = 0.999·ema + 0.001·raw`; BN buffers copied); validation and the
  shipped checkpoint use the EMA weights, the raw weights are kept under
  `train_model` for `--resume`.
- **Curriculum** (`--curriculum`): severity 0.4 → 0.7 → 1.0 at 0%/15%/35%
  of `--max-steps` (the data stream is rebuilt at each phase change).
- **Workers** (`--workers N`): page generation moves to N DataLoader
  processes (per-worker seed offsets keep streams disjoint), overlapping
  data creation with forward/backward.
- **Validation**: fixed line sets at severity 0.3 — with `--style mixed`,
  1000 document lines (seed 1234) + 1000 report lines (seed 1235). Reports
  CER (overall + per style), WER and punctuation CER every `--eval-every`
  steps; metrics append to `runs/metrics.jsonl`.
- **Stopping**: hard stop at `--max-steps`; early stop once CER <
  `--target-cer` holds for `--patience-after-target` steps without
  improvement, or on a `--plateau-steps` plateau. Ctrl+C saves `last.pt`.
- **Dashboard**: Rich live UI — run mode, loss/CER sparklines, per-style
  and punctuation CER, throughput, eval history, sample predictions with
  match coloring, RAM/CPU gauges. `--no-dashboard` logs plain lines.

### File formats

**Training checkpoint** (`checkpoints/best.pt`, `last.pt`): dict with
`model` (shipped/EMA state dict), `train_model` (raw weights, EMA runs
only), `optimizer`, `scheduler`, `step`, `cer`, `best_cer`, `best_step`,
`charset`, `parameters`, `style`.

**Bundled weights** (`legere/data/model.pt`): `model` (state dict),
`charset`, `step`, `cer` — loadable with `weights_only=True`. Refresh via
`legere export --bundle`.

**Line filter** (`legere/data/lineclf.pt`): `model` (state dict),
`accuracy`.

**Metrics log** (`runs/metrics.jsonl`, one JSON per eval): `step`,
`loss_ema`, `lr`, `cer`, `wer`, `punct_cer`, `cer_document`, `cer_report`,
`best_cer`, `samples_per_sec`, `ram_mb`, `elapsed_sec`.

---

## CLI reference

The entry point is `legere` (also `python -m legere`). `legere --version`
prints the version. `legere <file>` is a shortcut for `legere read <file>`.

### `legere read`

OCR a page image or PDF; text on stdout, stats on stderr.

```bash
legere read INPUT [options]
```

| Flag | Default | Description |
|---|---|---|
| `INPUT` | — | Image (any size) or PDF path |
| `--model PATH` | bundled | Checkpoint `.pt`, TorchScript export, or `.onnx` (auto-detects ONNX Runtime) |
| `--torchscript` | off | Load `--model` with `torch.jit.load` (TorchScript/INT8 exports) |
| `--json PATH` | — | Also write per-line `{text, bbox[, page]}` JSON |
| `--beam N` | 0 | CTC beam width (0 = greedy, fastest) |
| `--threads N` | 4 | `torch.set_num_threads` |
| `--dpi N` | 200 | PDF rendering resolution |
| `--no-line-filter` | off | Disable the learned non-text segment rejector |
| `--min-confidence F` | 0.0 | Drop lines below this mean CTC confidence (try 0.5 on noisy scans) |

Multi-page PDFs print `--- page N ---` separators and add a `page` field to
the JSON entries.

### `legere pdf`

Convert PDF pages to PNG files (no OCR). Requires `legere[pdf]`.

```bash
legere pdf INPUT.pdf [--out-dir DIR] [--dpi N]
```

| Flag | Default | Description |
|---|---|---|
| `--out-dir DIR` | alongside the PDF | Output directory (`<stem>_page001.png`, ...) |
| `--dpi N` | 200 | Rendering resolution |

### `legere train`

Train the CompactCRNN on the synthetic stream. Requires `legere[train]`.

```bash
legere train [options]        # defaults, live dashboard
legere train --setup          # interactive setup (prompts every option)
```

| Flag | Default | Description |
|---|---|---|
| `--setup` | off | Interactive wizard (needs a terminal); any value entered overrides the flags below |
| `--style {document,report,mixed}` | mixed | Synthetic page style for training and validation |
| `--finetune CKPT` | — | Start a NEW run from existing weights only (fresh optimizer/schedule/step/best-CER); accepts a `.pt` path or `bundled` |
| `--resume PATH` | — | Continue an interrupted run (restores weights, optimizer, schedule, step, best CER); mutually exclusive with `--finetune` |
| `--max-steps N` | 30000 | Hard step limit |
| `--lr F` | auto | Peak LR (2e-3 scratch / 5e-4 finetune) |
| `--weight-decay F` | 0.01 | AdamW weight decay (matrix params only) |
| `--warmup-steps N` | 500 | Linear LR warmup |
| `--width-budget N` | 9216 | Sum of padded widths per batch — the batch-size knob |
| `--accum N` | 1 | Gradient accumulation steps |
| `--workers N` | 0 | Background data-generation processes (0 = in-process) |
| `--threads N` | physical cores | Torch CPU threads |
| `--ram-limit MB` | 1024 | RAM gauge limit shown in the dashboard (display only) |
| `--curriculum` | off | Ramp severity 0.4 → 0.7 → 1.0 across the run |
| `--no-ema` | off | Disable EMA weight averaging |
| `--ema-decay F` | 0.999 | EMA decay per step |
| `--eval-every N` | 500 | Validation interval (also checkpoints) |
| `--val-size N` | 2000 | Validation lines (split per style when mixed) |
| `--val-seed N` | 1234 | Validation generator seed |
| `--seed N` | 42 | Training stream / torch seed |
| `--target-cer F` | 0.01 | Early-stop target |
| `--patience-after-target N` | 2500 | Extra steps without improvement once target met |
| `--plateau-steps N` | 8000 | Stop after this many steps without CER improvement |
| `--checkpoint-dir DIR` | checkpoints | Where `best.pt` / `last.pt` go (relative to the CWD) |
| `--metrics-file PATH` | runs/metrics.jsonl | JSONL metrics log |
| `--analysis-file PATH` | runs/analysis.jsonl | Full per-eval error breakdown (char classes, confusions, per-tag CER) |
| `--profile` | off | Print per-phase step timings (data/forward/CTC/backward/optimizer) at the end |
| `--no-dashboard` | off | Plain log lines instead of the Rich dashboard |

### `legere export`

Convert a training checkpoint into deployment formats.

```bash
legere export [--checkpoint PATH] [--out-dir DIR] [--onnx] [--bundle OUT]
```

| Flag | Default | Description |
|---|---|---|
| `--checkpoint PATH` | checkpoints/best.pt | Source checkpoint (or a weights-only `.pt`) |
| `--out-dir DIR` | exports | Writes `model_fp32.ts.pt` (scripted+frozen) and `model_int8.ts.pt` (dynamic INT8 of GRU+Linear) |
| `--onnx` | off | Also write `model.onnx` (opset 17, dynamic batch/width; parity-checked vs eager when onnxruntime is installed) |
| `--bundle OUT_PT` | — | Instead of exports, write a weights-only `.pt` (state dict + charset + metadata) — use `--bundle src/legere/data/model.pt` to refresh the packaged weights |

### `legere benchmark`

Two modes. **Synthetic benchmark** (no input file): line/page latency and
full-page CER for every available backend. **Profile mode** (input file):
stage-by-stage performance report for a real image/PDF.

```bash
legere benchmark [options]                 # synthetic, all backends
legere benchmark INPUT [--json out.json]   # profile a real file
```

| Flag | Default | Description |
|---|---|---|
| `INPUT` | — | Image or PDF to profile (enables profile mode) |
| `--profile` | off | Force profile mode (implied by INPUT) |
| `--analyze` | off | Page-level error autopsy on synthetic pages (see below) |
| `--seed N` | 2024 | Page generator seed (analyze mode) |
| `--json PATH` | — | Write the profile/analyze report as JSON (stable `schema_version`) |
| `--checkpoint PATH` | bundled | Model for eager / profile mode (`.pt` or `.onnx`) |
| `--torchscript` | off | Load `--checkpoint` with torch.jit (profile mode) |
| `--beam N` | 0 | CTC beam width (profile mode) |
| `--dpi N` | 200 | PDF rendering resolution (profile mode) |
| `--fp32-ts PATH` | exports/model_fp32.ts.pt | TorchScript FP32 variant (skipped if missing) |
| `--int8-ts PATH` | exports/model_int8.ts.pt | TorchScript INT8 variant (skipped if missing) |
| `--onnx PATH` | exports/model.onnx | ONNX variant (skipped if missing) |
| `--pages N` | 5 | Synthetic Full-HD pages for the page benchmark |
| `--line-runs N` | 30 | Timed runs for single-line latency |
| `--threads N` | 4 | Torch CPU threads |
| `--style {document,report,mixed}` | document | Page style for the synthetic page benchmark |

The profile report covers: system/CPU/backend, model load time and size,
per-stage pipeline times (PDF render, resize, deskew, segmentation,
normalization, tensor prep, CRNN forward, decode), per-line width and CRNN
latency stats (avg/median/P95/max), per-batch padding waste, and lines/s +
chars/s throughput. `legere train --profile` produces the training-side
equivalent.

`legere benchmark --analyze` runs a **page-level error autopsy**: per-page
CER (with two-column flag), line-detection precision/recall, edit-distance
decomposition (recognition vs missed vs spurious lines), and — over the
aligned lines — line accuracy, per-character-class error rates, the top
character confusions (`3→5`, `:→.`, ...) and CER per tag: font, font size,
block type, decoration, degradation effect, content type and line length.
Tags are attached to every synthetic line by the page generator.

### `legere detector`

Train the LineClassifier (text/non-text gate). Positives: GT line crops
from both page styles at random severities (including inverted text).
Negatives: synthetic barcodes/QR/rules/squiggles/blank paper + random
non-text page regions.

```bash
legere detector [--steps N] [--out PATH]
```

| Flag | Default | Description |
|---|---|---|
| `--steps N` | 600 | Training steps (BCE loss, Adam) |
| `--batch-size N` | 32 | Crops per step |
| `--lr F` | 1e-3 | Learning rate |
| `--seed N` | 7 | Data/torch seed |
| `--out PATH` | lineclf.pt | Output; use `src/legere/data/lineclf.pt` to refresh the bundled filter |

---

## Python API

```python
from legere import (
    Legere,            # engine
    PageResult,        # .lines, .skew_angle, .text
    LineResult,        # .text, .bbox
    load_model,        # low-level model loader (all backends)
    run_page_ocr,      # low-level pipeline on a grayscale ndarray
    CompactCRNN, ModelConfig, CHARSET,
    LineSegment, preprocess_and_segment,
)
from legere.inference import load_line_filter, LineFilter, OnnxModel
from legere.metrics import (
    character_error_rate, word_error_rate,
    punctuation_error_rate, levenshtein,
)
from legere.pdf import render_pdf_pages, pdf_to_pngs, pdf_page_count
```

**`Legere(model=None, torchscript=False, beam_width=0, threads=None,
line_filter=True, min_confidence=0.0)`** — loads once, reuse for many pages.
`model` accepts a checkpoint path, a TorchScript export (with
`torchscript=True`) or an `.onnx` path.

- `.read(image)` → `PageResult` — `image` is a path or numpy array
  (grayscale `(H,W)` or BGR `(H,W,3)`); PDFs are rejected with a pointer to:
- `.read_pdf(path, dpi=200)` → `list[PageResult]`, one per page.

**`run_page_ocr(gray, model, beam_width=0, line_filter=None,
min_confidence=0.0)`** → `(texts, segments, skew_angle)` — the low-level
pipeline for custom flows.

**`load_line_filter(path=None, threshold=0.1)`** → `LineFilter | None` —
the bundled text/non-text gate; `keep_mask(crops)` returns booleans.

## Metrics

`character_error_rate` and `word_error_rate` are edit-distance rates over
characters/words. `punctuation_error_rate(preds, refs, marks=".,;:")`
filters both strings down to the tracked marks before the edit distance, so
`,`↔`.` and `;`↔`:` confusions register even when everything else is right.

`legere.analysis` provides the fine-grained tooling used by training and
`--analyze`: `levenshtein_ops` (edit alignment with backtrace),
`confusion_counts`, `char_class_rates` (letters/digits/accented/punct/
symbols/space), `classify_content` (money/date/time/uid/serial/... tags),
`line_accuracy`, `align_lines` (predicted↔GT line alignment) and
`analyze_lines`, which bundles everything into one JSON-serializable
report. During training the full report is appended per eval to
`runs/analysis.jsonl`, and the dashboard surfaces line accuracy, the top
confusions and the worst tag hotspots live.

## Project layout

```
src/legere/
├── __init__.py        # public API + __version__
├── charset.py         # charset, encode/decode, sanitize
├── model.py           # CompactCRNN, LineClassifier, CTC decoders
├── segment.py         # deskew, column detection, line segmentation
├── lines.py           # crop normalization + tensor batching
├── inference.py       # Legere engine, backends, filters, pipeline
├── metrics.py         # CER / WER / punctuation CER / levenshtein
├── pdf.py             # PDF rendering (pypdfium2)
├── profiling.py       # Profiler + stage-by-stage inference profiler
├── cli.py             # `legere` command
├── data/
│   ├── model.pt       # bundled CRNN weights (state dict + charset)
│   └── lineclf.pt     # bundled line-filter weights
└── training/
    ├── pagegen.py     # synthetic page generator (styles, degradations)
    ├── dataset.py     # training stream, crop extraction, validation sets
    ├── train.py       # training loop, dashboard, EMA, curriculum
    ├── lineclf.py     # LineClassifier training (`legere detector`)
    ├── export.py      # TorchScript / INT8 / ONNX / bundle exports
    └── benchmark.py   # synthetic benchmark + real-file profiler entry
```

## Development

```bash
pip install -e .[pdf,onnx,train] pytest
pytest tests/ -q
```

Tests cover: charset round-trip, model shapes and the 1M-parameter budget,
bundled-weight loading, blank/sample page reads, GT validity of augmented
pages, punctuation metric, PDF rendering/conversion/reading, profiler
sections and report, and ONNX↔eager output parity.

## License

MIT — see [LICENSE](LICENSE).
