Metadata-Version: 2.4
Name: signal-digitizer
Version: 0.1.3
Summary: Digitize scanned strip-chart and grid-plot PDFs into calibrated one-dimensional signals with automated grid detection, skew correction, trace extraction, and adaptive noise cancellation.
Author: Manoj Kumar C S, V N Manjunath Aradhya, Nikhil D Bharadwaj
License: MIT
Project-URL: Homepage, https://gitlab.com/manojkumarcs/signal-digitizer
Keywords: pdf,chart digitization,signal processing,ecg,pymupdf,adaptive filtering
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Science/Research
Classifier: License :: OSI Approved :: MIT License
Classifier: Programming Language :: Python :: 3
Classifier: Topic :: Scientific/Engineering :: Image Processing
Requires-Python: >=3.9
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: pymupdf>=1.24
Requires-Dist: numpy>=1.22
Requires-Dist: opencv-python-headless>=4.6
Requires-Dist: scipy>=1.8
Provides-Extra: dev
Requires-Dist: pytest; extra == "dev"
Dynamic: license-file

# signal-digitizer

Turn a scanned strip-chart / grid-plot PDF (e.g. an ECG trace, an old lab
recorder printout) into a calibrated 1D signal.

Pipeline: rasterize page → estimate & correct skew → detect grid lines
(Canny + Hough) → calibrate axes from grid spacing → extract ink trace →
optional adaptive noise cancellation (LMS/NLMS) → calibrated (x, y) signal.

This is a plain library that depends on the official `pymupdf` package from
PyPI — it does not fork, patch, or vendor PyMuPDF in any way, so it installs
cleanly alongside any other project using `pymupdf`.

## Contents

- [Install](#install)
- [Usage](#usage)
- [CLI](#cli)
- [API](#api)
- [Development & Testing](#development--testing)
- [Publishing to PyPI](#publishing-to-pypi)
- [License](#license)

## Install

```bash
pip install signal-digitizer
```

(For local development, from this directory: `pip install -e .`)

## Usage

```python
import signal_digitizer as sd

x, y = sd.digitize(
    "chart.pdf",
    page_number=0,
    unit_per_vgap=1.0,
    unit_per_hgap=1.0,
)

# With adaptive noise cancellation (self-referencing Adaptive Line Enhancer)
x, y = sd.digitize("chart.pdf", use_anc=True, anc_mu=0.05, anc_filter_order=8)

# Validate against a ground-truth signal
result = sd.validate_signal(y, reference_signal)
print(result)  # {"pearson_r": ..., "rmse": ..., "meets_target": ...}
```

### Working from an already-open pymupdf document

```python
import pymupdf
import signal_digitizer as sd

doc = pymupdf.open("chart.pdf")
x, y = sd.digitize_page(doc[0], dpi=300, unit_per_vgap=1.0, unit_per_hgap=1.0)
```

### From a numpy image you've already rasterized

```python
x, y = sd.digitize_image(image_array, unit_per_vgap=1.0, unit_per_hgap=1.0)
```

## CLI

```bash
signal-digitizer chart.pdf --page 0 -o signal.csv

# with adaptive noise cancellation
signal-digitizer chart.pdf --page 0 --use-anc --anc-mu 0.05 --anc-filter-order 8 -o signal.csv
```

## API

- `digitize(pdf_path, ...)` — full pipeline from a PDF file path
- `digitize_page(page, ...)` — full pipeline from an open `pymupdf.Page`
- `digitize_image(image, ...)` — full pipeline from a numpy RGB image
- `render_page`, `render_pymupdf_page` — rasterization only
- `estimate_skew_deg`, `deskew` — skew correction
- `detect_grid`, `GridLines` — grid-line detection
- `calibrate_from_grid`, `AxisCalibration` — pixel → data-unit calibration
- `extract_trace_pixels`, `to_signal` — ink trace extraction
- `AdaptiveNoiseCanceller`, `denoise_adaptive`, `build_self_reference` — LMS/NLMS denoising
- `validate_signal` — Pearson-r / RMSE comparison against a reference signal

### Key `digitize()` parameters

| Parameter | Default | Meaning |
|---|---|---|
| `dpi` | 300 | Rasterization resolution |
| `unit_per_vgap` / `unit_per_hgap` | 1.0 | Data units per grid cell (x / y) |
| `ink_thresh` | 128 | Grayscale threshold below which a pixel counts as trace ink |
| `smooth_window` | `None` | Moving-average window applied after extraction |
| `correct_skew` | `True` | Estimate & correct page rotation before grid detection |
| `use_anc` | `False` | Apply adaptive LMS/NLMS denoising to the trace |
| `anc_algorithm` | `"nlms"` | `"nlms"` (recommended) or `"lms"` |
| `anc_filter_order` | 8 | Number of adaptive filter taps |
| `anc_mu` | 0.05 | Adaptation step size |

## Development & Testing

Set up a local dev environment and run the test suite before every release:

```bash
python -m pip install -e ".[dev]"
python -m pytest tests/ -v
```

All tests should pass before you build or publish a new version.

## Publishing to PyPI

This section walks through releasing `signal-digitizer` to PyPI end to end,
from account setup through the production upload. Run all commands from this
project's root directory (the one containing `pyproject.toml`) unless noted
otherwise.

### Step 1 — Create PyPI and TestPyPI accounts

PyPI (production) and TestPyPI (a separate sandbox index for dry runs) use
independent accounts and databases.

1. Sign up at [pypi.org](https://pypi.org/account/register/).
2. Sign up separately at [test.pypi.org](https://test.pypi.org/account/register/).
3. Verify your email address on **both** sites.
4. Enable two-factor authentication (2FA) on PyPI with an authenticator app —
   this is mandatory for new accounts before you can upload anything.

### Step 2 — Confirm the package name is available

Visit `https://pypi.org/project/signal-digitizer/` directly in your browser.

- **404 Not Found** → the name is free, proceed.
- **Project page loads** → the name is taken. Change the `name` field in
  `pyproject.toml` to something else (e.g. `chart-digitizer`,
  `ecg-chart-digitizer`). Note that the *import* name (`signal_digitizer`,
  used in `import signal_digitizer as sd`) does not have to match the PyPI
  distribution name, so you don't need to rename the package internally.

PyPI names are first-come-first-served and cannot be reassigned later, so
confirm this before you upload anything, including a TestPyPI dry run.

### Step 3 — Generate an API token

Uploads authenticate with an API token, not your account password.

1. On PyPI: **Account Settings → API tokens → Add API token**.
2. For your *first* upload, scope the token to your entire account (a
   project-scoped token isn't possible until the project exists on PyPI).
3. Copy the token immediately — it is shown only once. It starts with
   `pypi-`.
4. Repeat the same process separately on **TestPyPI** to get a second token
   for the sandbox index.
5. After your first successful production upload, come back to PyPI, create
   a new token scoped only to the `signal-digitizer` project, and delete the
   old account-wide token.

### Step 4 — Store credentials for twine

Create (or edit) `~/.pypirc` with both indexes configured:

```ini
[distutils]
index-servers =
    pypi
    testpypi

[pypi]
username = __token__
password = pypi-AgEIcHlwaS5vcmc...

[testpypi]
repository = https://test.pypi.org/legacy/
username = __token__
password = pypi-AgENdGVzdC5weXBpLm9yZw...
```

- `username` is literally the string `__token__` in both sections — not
  your actual PyPI username.
- `password` is the full token string, including the `pypi-` prefix.
- Lock the file down so it isn't world-readable:

```bash
chmod 600 ~/.pypirc
```

### Step 5 — Install the build and upload tools

```bash
python -m pip install --upgrade build twine
```

### Step 6 — Build fresh distribution artifacts

Always rebuild from a clean state right before publishing, so no stale
artifacts from earlier local testing get uploaded by accident:

```bash
rm -rf dist/ build/ src/*.egg-info
python -m build
```

This produces two files in `dist/`:

- `signal_digitizer-<version>.tar.gz` — the source distribution (sdist)
- `signal_digitizer-<version>-py3-none-any.whl` — the built wheel

### Step 7 — Validate the artifacts

```bash
python -m twine check dist/*
```

This catches metadata problems and README rendering issues (PyPI renders
`README.md` as the project description) before you commit to an upload.
Fix any errors it reports and rebuild before continuing — a given version
number can be uploaded to PyPI **exactly once** and can only be *yanked*
afterward, never overwritten.

### Step 8 — Upload to TestPyPI first

```bash
python -m twine upload --repository testpypi dist/*
```

Then, in a **clean virtual environment** (not the one you developed in),
install from TestPyPI:

```bash
python -m venv /tmp/sd-test-env
source /tmp/sd-test-env/bin/activate
pip install --index-url https://test.pypi.org/simple/ \
            --extra-index-url https://pypi.org/simple/ \
            signal-digitizer
```

The `--extra-index-url` flag is required because real dependencies
(`pymupdf`, `numpy`, `opencv-python-headless`, `scipy`) generally aren't
mirrored on TestPyPI, so pip needs to fall back to the real index for them.

Verify the installed package works correctly (not your local source tree):

```bash
python -c "import signal_digitizer as sd; print(sd.__version__)"
signal-digitizer --help
```

Deactivate and remove the test environment once you're done:

```bash
deactivate
rm -rf /tmp/sd-test-env
```

### Step 9 — Upload to production PyPI

Once the TestPyPI install checks out cleanly:

```bash
python -m twine upload dist/*
```

Within a few minutes, `pip install signal-digitizer` will work for anyone.

### Step 10 — Releasing future versions

For every subsequent release:

1. Bump the `version` field in `pyproject.toml` (PyPI rejects re-uploading
   an existing version number — even if you delete and recreate the file).
2. Update the changelog / README as needed.
3. Repeat **Step 6 → Step 9** (rebuild, validate, TestPyPI, then production).

### Optional — Trusted Publishing (recommended for CI)

If you later publish from GitHub Actions instead of your local machine,
prefer PyPI's **Trusted Publishing** (OIDC-based, no stored token at all)
over long-lived API tokens:

1. On your PyPI project page: **Publishing → Add a new publisher**.
2. Link it to your GitHub repository, workflow filename, and environment.
3. Your CI workflow can then run `twine upload` (or the official
   `pypa/gh-action-pypi-publish` action) with no secrets configured at all.

This also requires pushing the project to a real GitHub repository first —
update the placeholder `Homepage` URL in `pyproject.toml` once you do.

## License

`signal-digitizer` itself is released under the **MIT License**. In short,
that means anyone can use, copy, modify, merge, publish, distribute,
sublicense, and/or sell copies of this code — including in closed-source or
commercial projects — as long as the original copyright notice and license
text are kept somewhere in the distribution. There is no copyleft
obligation: you do not have to release your own code under MIT or any
particular license just because it uses this library.

### Copyright notice

```
MIT License

Copyright (c) 2026 Manoj Kumar C S

Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:

The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.

THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE.
```

The full text also lives in the [`LICENSE`](LICENSE) file at the project
root — that file is the authoritative copy; keep this section in sync with
it if either changes.

### Third-party dependency licenses

`signal-digitizer` only *depends on* the packages below via their public
APIs — none of their source is forked, vendored, or modified in this
repository. Each dependency keeps its own license, which applies to that
dependency's code, not to `signal-digitizer`'s own MIT-licensed code:

| Dependency | License | Notes |
|---|---|---|
| [`pymupdf`](https://pypi.org/project/PyMuPDF/) | AGPL-3.0-or-later (or Artifex commercial license) | Used only through its public API (`pymupdf.open`, `Matrix`, `get_pixmap`) for PDF rasterization. See caveat below. |
| [`numpy`](https://numpy.org/) | BSD-3-Clause | Array operations |
| [`opencv-python-headless`](https://pypi.org/project/opencv-python-headless/) | Apache-2.0 (OpenCV core) | Edge detection, Hough transform, deskewing |
| [`scipy`](https://scipy.org/) | BSD-3-Clause | Pearson correlation in `validate_signal` |

**A note on PyMuPDF's license:** PyMuPDF is dual-licensed under AGPL-3.0 (for
open-source use) or a paid commercial license from Artifex for closed-source
distribution. Because `signal-digitizer` depends on the official `pymupdf`
package as an ordinary import rather than embedding or modifying its source,
`signal-digitizer`'s own code stays MIT — but **your** obligations as an
end user of the combined stack depend on how you use PyMuPDF, not on this
library. If you distribute a closed-source application that uses PyMuPDF
(directly or via `signal-digitizer`), check whether that triggers AGPL's
copyleft/source-availability terms for your application, or whether you need
Artifex's commercial license. This is exactly the licensing tangle this
library was designed to avoid *creating* — by depending on the real
`pymupdf` package instead of forking it — but it doesn't remove your
obligations as a downstream user of PyMuPDF itself.

This is general information, not legal advice — consult a lawyer if your
use case has real licensing stakes (e.g. a proprietary commercial product
built on this).
