Metadata-Version: 2.4
Name: medh5
Version: 1.4.0
Summary: Self-describing HDF5 container for one medical imaging sample and all of its ground truth: multi-timepoint, multi-modal images with segmentation, detection, classification and registration annotations, provenance and integrity.
Author-email: Puyang Wang <pauliwang411@gmail.com>
License-Expression: MIT
Project-URL: Homepage, https://github.com/XwK-P/medh5
Project-URL: Repository, https://github.com/XwK-P/medh5
Project-URL: Issues, https://github.com/XwK-P/medh5/issues
Project-URL: Changelog, https://github.com/XwK-P/medh5/blob/main/CHANGELOG.md
Project-URL: Documentation, https://medh5.readthedocs.io/
Project-URL: Specification, https://github.com/XwK-P/medh5/blob/main/docs/spec/medh5-1.0.md
Keywords: medical imaging,hdf5,blosc2,machine learning,segmentation,detection,registration,longitudinal,nifti,dicom,rtstruct,nnunet,pytorch,monai
Classifier: Development Status :: 5 - Production/Stable
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Scientific/Engineering
Classifier: Topic :: Scientific/Engineering :: Medical Science Apps.
Classifier: Typing :: Typed
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: h5py>=3.13
Requires-Dist: hdf5plugin>=4.1
Requires-Dist: numpy>=1.24
Provides-Extra: dev
Requires-Dist: pytest>=7.0; extra == "dev"
Requires-Dist: pytest-cov>=4.0; extra == "dev"
Requires-Dist: ruff==0.16.3; extra == "dev"
Requires-Dist: mypy==1.18.2; extra == "dev"
Requires-Dist: build>=1.0; extra == "dev"
Requires-Dist: twine>=5.0; extra == "dev"
Requires-Dist: jsonschema>=4.18; extra == "dev"
Requires-Dist: hypothesis>=6.100; extra == "dev"
Provides-Extra: schema
Requires-Dist: jsonschema>=4.18; extra == "schema"
Provides-Extra: torch
Requires-Dist: torch>=2.0; extra == "torch"
Provides-Extra: monai
Requires-Dist: monai>=1.3; extra == "monai"
Requires-Dist: torch>=2.0; extra == "monai"
Provides-Extra: nifti
Requires-Dist: nibabel>=5; extra == "nifti"
Provides-Extra: dicom
Requires-Dist: pydicom>=2.4; extra == "dicom"
Provides-Extra: dicomseg
Requires-Dist: highdicom>=0.22; extra == "dicomseg"
Requires-Dist: pydicom>=2.4; extra == "dicomseg"
Provides-Extra: itk
Requires-Dist: SimpleITK>=2.3; extra == "itk"
Provides-Extra: interp
Requires-Dist: scipy>=1.10; extra == "interp"
Dynamic: license-file

# medh5

[![PyPI version](https://img.shields.io/pypi/v/medh5.svg)](https://pypi.org/project/medh5/)
[![Python versions](https://img.shields.io/pypi/pyversions/medh5.svg)](https://pypi.org/project/medh5/)
[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](LICENSE)
[![CI](https://github.com/XwK-P/medh5/actions/workflows/ci.yml/badge.svg?branch=main)](https://github.com/XwK-P/medh5/actions/workflows/ci.yml)
[![Typed](https://img.shields.io/badge/typed-mypy%20strict-informational.svg)](medh5/py.typed)
[![Code style: ruff](https://img.shields.io/badge/code%20style-ruff-000000.svg)](https://github.com/astral-sh/ruff)

**One medical imaging sample — a subject, at every timepoint, with all of its
ground truth — in a single self-describing HDF5 file.**

Multi-modality images, segmentation in five encodings, detection boxes,
keypoints, contours, meshes, classification, registration between visits,
provenance and quality records, and per-object integrity digests. Format
version **1.0**, with a [normative specification](https://medh5.readthedocs.io/en/latest/spec/medh5-1.0/) and a
[116-case conformance suite](https://medh5.readthedocs.io/en/latest/spec/conformance/) any implementation can run.

```python
import medh5

with medh5.open("case_0001.medh5") as s:
    s.identity.subject_id                              # "BRATS-GLI-01234"
    s.at("tp1").images["CT_tp1"].read(physical=True)   # HU, not raw counts
    s.annotations["organs"].dense(["liver", "spleen"]) # any encoding, one API
    s.transform_between("tp0", "tp1")                  # resolved via frames
    s.tracks("lesion")                                 # lesions joined across visits
```

## Install

```bash
pip install medh5
pip install "medh5[torch,nifti,dicom]"
```

Reading and writing needs only `h5py`, `hdf5plugin` and `numpy`. Extras:
`torch`, `monai`, `nifti`, `dicom`, `dicomseg`, `itk`, `schema`, `interp`.

## Documentation

**[medh5.readthedocs.io](https://medh5.readthedocs.io/)** — tutorials, how-to
guides, the Python and CLI reference, and the normative specification.

[Write your first sample](https://medh5.readthedocs.io/en/latest/tutorials/first-sample/) ·
[How-to guides](https://medh5.readthedocs.io/en/latest/guides/) ·
[Python API](https://medh5.readthedocs.io/en/latest/reference/python-api/) ·
[CLI](https://medh5.readthedocs.io/en/latest/reference/cli/) ·
[Specification](https://medh5.readthedocs.io/en/latest/spec/medh5-1.0/)

## What the format is for

- **One file per subject, not per scan** — every visit in one place, so
  longitudinal work has a referent and splitting by file cannot leak a patient.
- **Geometry is stated once and never guessed** — declared grids, boxes at voxel
  edges, and converters that refuse rather than invent.
- **Absence is not silence** — a class examined and not found is recorded as
  such, which is a different training signal from one nobody examined.
- **Every claim is checkable** — per-object digests, a Merkle `content_id` that
  survives recompression, a stable diagnostic-code table, and a 116-case
  conformance corpus.
- **Reading a patch is fast** — a 64³ multi-class patch in ~4 ms, and O(1)
  foreground sampling once `build_index()` has run.

[The reasoning behind each](https://medh5.readthedocs.io/en/latest/).

## Write a sample

```python
import numpy as np
import medh5
from medh5 import LabelClass, LabelSet

labels = LabelSet("demo-v1", version="1.0.0", classes=[
    LabelClass(1, "liver", "Liver", category="organ"),
    LabelClass(2, "spleen", "Spleen", category="organ"),
    LabelClass(3, "lesion", "Lesion", parents=[1], category="lesion"),
])

with medh5.create("case_0001.medh5", sample_id="case_0001",
                  subject_id="DEMO-0001") as w:
    w.label_set(labels)
    w.add_timepoint("tp0", label="baseline", days_from_baseline=0)
    w.add_grid("ct", shape=ct.shape, spacing=(2.0, 0.8, 0.8),
               origin=(-64.0, -38.4, -38.4), timepoint="tp0")
    w.add_image("CT", ct, grid="ct", modality="CT",
                value_type="quantitative", value_units="HU")
    w.add_segmentation("organs", grid="ct",
                       masks={"liver": liver, "lesion": lesion},
                       annotated_classes=["liver", "spleen", "lesion"])
    w.build_index()   # optional; foreground sampling is O(1) only with it
```

`annotated_classes` names the spleen although there is no spleen mask: that
records "we looked and found none". The encoding is chosen by measuring the
class overlap graph — liver and lesion overlap, so it picks one that can
represent that — and the write is atomic.

## Train on it

```python
from torch.utils.data import DataLoader
from medh5.torch import PatchDataset, collate, worker_init_fn
from medh5.sampling import PatchSampler

sampler = PatchSampler((96, 96, 96), strategy="balanced",
                       foreground_classes=["liver", "lesion"])
dataset = PatchDataset(paths, sampler, images=["CT"],
                       annotations={"organs": ["liver", "lesion"]},
                       samples_per_volume=8)

loader = DataLoader(dataset, batch_size=2, num_workers=8,
                    worker_init_fn=worker_init_fn, collate_fn=collate)
```

`worker_init_fn` drops handles inherited across a `fork`. It is recommended
rather than required: the handle cache is PID-keyed and re-checks ownership on
every access, so a forked worker abandons the parent's handles on first use
rather than reading through or closing them.

## Command line

```bash
medh5 info case.medh5                  # grids, images, annotations, coverage
medh5 validate case.medh5 --level strict
medh5 verify case.medh5                # digests and content_id
medh5 timeline case.medh5              # visits and intervals
medh5 track case.medh5 --class lesion  # per-lesion volumes across visits

medh5 dataset index studies/ -o cohort.json
medh5 dataset split cohort.json --group-by group_id --stratify-by site_id
medh5 dataset stats cohort.json --partition train --workers 8
medh5 dataset check cohort.json --deep

medh5 convert from-dicom /studies out/     # one sample per patient, all visits
medh5 convert from-nifti case.medh5 --image CT=ct.nii.gz
medh5 convert from-rtstruct plan.dcm case.medh5 --rasterize
medh5 migrate old/*.medh5 -o new/ --group-by subject

medh5 scrub out/*.medh5 --apply --date-shift-days -117
medh5 pack cohort/*.medh5 -o shard.medh5c
medh5 recompress cohort/*.medh5 --profile training
medh5 bench                                # reproduce the performance targets
medh5 conformance publish suite/           # the suite, for another implementation
```

## Interoperability

| Format | |
|---|---|
| **NIfTI** | affine and voxels bit-identical on round trip; RAS↔LPS is a sign flip, never a resample |
| **DICOM** | slices ordered by geometry, spacing measured between origins, modality LUT stored not applied, tags on an explicit allow-list; slices that disagree about orientation, spacing or rescale are refused rather than read off the first one |
| **DICOM SEG** | frames placed by geometry; segments matched by label, not number; import preserves overlap and `FRACTIONAL`, export writes `BINARY` |
| **RTSTRUCT** | contours stay contours; rasterisation is opt-in and recorded in provenance |
| **nnU-Net v2** | class ids kept; region labels become label-set DAG parents; `dataset.json` round-trips |
| **MONAI** | `to_metatensor` gives a `MetaTensor` with the correct affine |
| **0.x** | `medh5 migrate`, reporting every decision and every guess |

Every conversion returns a report distinguishing what it **decided** from the
data and where it **guessed** — the encoding chosen, the class ids minted, a
half-voxel convention changed, a timepoint order inferred rather than read.

COCO is deliberately unsupported: it has no world geometry, spacing or frame of
reference, so importing means inventing a grid and exporting means discarding
the geometry that makes a medical annotation reproducible.

## Reading it without medh5

```python
import h5py, json, hdf5plugin       # hdf5plugin only for blosc2 profiles

with h5py.File("case_0001.medh5") as f:
    doc = json.loads(f["meta"][()])
    doc["identity"]["subject_id"]
    dict(f["grids"]["ct"].attrs)     # spacing, origin, direction
    f["images"]["CT"][10:20]
```

`medh5 recompress --profile portable` writes gzip, readable by any HDF5 build.

## Versioning

The **format** is 1.0. A minor version may add optional objects, profiles,
encodings and diagnostic codes; it may not change what an existing one means
(spec §16). The **package** follows semantic versioning from 1.0.0.

0.x files are not readable by 1.0 and are not meant to be — `medh5 migrate`
converts them once. See [Converters](https://medh5.readthedocs.io/en/latest/guides/migrate-0x/).

## License

MIT
