Metadata-Version: 2.4
Name: signal-digitizer
Version: 0.1.2
Summary: Digitize scanned strip-chart / grid-plot PDFs into calibrated 1D signals, built on PyMuPDF
Author: Manoj Kumar C S
License: MIT
Project-URL: Homepage, https://github.com/yourusername/signal-digitizer
Keywords: pdf,chart digitization,signal processing,ecg,pymupdf,adaptive filtering
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Science/Research
Classifier: License :: OSI Approved :: MIT License
Classifier: Programming Language :: Python :: 3
Classifier: Topic :: Scientific/Engineering :: Image Processing
Requires-Python: >=3.9
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: pymupdf>=1.24
Requires-Dist: numpy>=1.22
Requires-Dist: opencv-python-headless>=4.6
Requires-Dist: scipy>=1.8
Provides-Extra: dev
Requires-Dist: pytest; extra == "dev"
Dynamic: license-file

# signal-digitizer

Turn a scanned strip-chart / grid-plot PDF (e.g. an ECG trace, an old lab
recorder printout) into a calibrated 1D signal.

Pipeline: rasterize page → estimate & correct skew → detect grid lines
(Canny + Hough) → calibrate axes from grid spacing → extract ink trace →
optional adaptive noise cancellation (LMS/NLMS) → calibrated (x, y) signal.

This is a plain library that depends on the official `pymupdf` package from
PyPI — it does not fork, patch, or vendor PyMuPDF in any way, so it installs
cleanly alongside any other project using `pymupdf`.

## Install

```bash
pip install signal-digitizer
```

(For local development, from this directory: `pip install -e .`)

## Usage

```python
import signal_digitizer as sd

x, y = sd.digitize(
    "chart.pdf",
    page_number=0,
    unit_per_vgap=1.0,
    unit_per_hgap=1.0,
)

# With adaptive noise cancellation (self-referencing Adaptive Line Enhancer)
x, y = sd.digitize("chart.pdf", use_anc=True, anc_mu=0.05, anc_filter_order=8)

# Validate against a ground-truth signal
result = sd.validate_signal(y, reference_signal)
print(result)  # {"pearson_r": ..., "rmse": ..., "meets_target": ...}
```

### Working from an already-open pymupdf document

```python
import pymupdf
import signal_digitizer as sd

doc = pymupdf.open("chart.pdf")
x, y = sd.digitize_page(doc[0], dpi=300, unit_per_vgap=1.0, unit_per_hgap=1.0)
```

### From a numpy image you've already rasterized

```python
x, y = sd.digitize_image(image_array, unit_per_vgap=1.0, unit_per_hgap=1.0)
```

## CLI

```bash
signal-digitizer chart.pdf --page 0 -o signal.csv

# with adaptive noise cancellation
signal-digitizer chart.pdf --page 0 --use-anc --anc-mu 0.05 --anc-filter-order 8 -o signal.csv
```

## API

- `digitize(pdf_path, ...)` — full pipeline from a PDF file path
- `digitize_page(page, ...)` — full pipeline from an open `pymupdf.Page`
- `digitize_image(image, ...)` — full pipeline from a numpy RGB image
- `render_page`, `render_pymupdf_page` — rasterization only
- `estimate_skew_deg`, `deskew` — skew correction
- `detect_grid`, `GridLines` — grid-line detection
- `calibrate_from_grid`, `AxisCalibration` — pixel → data-unit calibration
- `extract_trace_pixels`, `to_signal` — ink trace extraction
- `AdaptiveNoiseCanceller`, `denoise_adaptive`, `build_self_reference` — LMS/NLMS denoising
- `validate_signal` — Pearson-r / RMSE comparison against a reference signal

### Key `digitize()` parameters

| Parameter | Default | Meaning |
|---|---|---|
| `dpi` | 300 | Rasterization resolution |
| `unit_per_vgap` / `unit_per_hgap` | 1.0 | Data units per grid cell (x / y) |
| `ink_thresh` | 128 | Grayscale threshold below which a pixel counts as trace ink |
| `smooth_window` | `None` | Moving-average window applied after extraction |
| `correct_skew` | `True` | Estimate & correct page rotation before grid detection |
| `use_anc` | `False` | Apply adaptive LMS/NLMS denoising to the trace |
| `anc_algorithm` | `"nlms"` | `"nlms"` (recommended) or `"lms"` |
| `anc_filter_order` | 8 | Number of adaptive filter taps |
| `anc_mu` | 0.05 | Adaptation step size |

## License

MIT. See `LICENSE`.
