Metadata-Version: 2.4
Name: librivox-mirror
Version: 0.1.0
Summary: Continuously mirror LibriVox audiobooks into ML-ready Hugging Face datasets
License-Expression: MIT
License-File: LICENSE
Classifier: Development Status :: 3 - Alpha
Classifier: Environment :: Console
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3.14
Classifier: Topic :: Multimedia :: Sound/Audio
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Dist: huggingface-hub[hf-xet]>=1.0
Requires-Dist: httpx>=0.28
Requires-Dist: internetarchive>=5.4.2
Requires-Dist: mutagen>=1.47
Requires-Dist: pyarrow>=21
Requires-Dist: pydantic>=2.11
Requires-Dist: rich>=14
Requires-Dist: tenacity>=9.1
Requires-Dist: typer>=0.17
Requires-Python: >=3.12
Description-Content-Type: text/markdown

# librivox-mirror

Fast, structured, continuously updated LibriVox audio mirror.

Original MP3s are mirrored to Hugging Face as deterministic WebDataset TARs with
compact Parquet indexes. Dedicated infrastructure handles the initial backfill;
GitHub Actions handles updates.

## Install

```console
uvx librivox-mirror --help
```

## Use

```console
librivox-mirror plan
librivox-mirror mirror 47
librivox-mirror backfill --start-id 1 --end-id 1000
```

Set `HF_DATASET_REPO` and `HF_TOKEN` to publish. Local SQLite checkpoints make
backfills resumable.

## Dataset

- `preview` (default): browser-playable original MP3 samples
- `sections`: one Parquet row per audio section
- `books`: one Parquet row per book
- `data/`: streaming WebDataset audio and sample metadata

Audio is never transcoded. Complete LibriVox and Internet Archive metadata and
checksums are preserved.

## Develop

```console
uv sync --locked --group dev --no-group integration
```

Python 3.12–3.14 is supported. See [CONTRIBUTING.md](CONTRIBUTING.md).

## Licenses

Code is MIT. Mirror-specific curation, indexes, metadata, and documentation are CC
BY 4.0. Original LibriVox audio remains public domain in the United States.
