Metadata-Version: 2.4
Name: audio-super-resolution
Version: 0.6.0
Summary: Easy to use audio super-resolution and bandwidth extension from CLI or as a Python package.
Project-URL: Homepage, https://github.com/Tinnci/python-audio-super-resolution
Project-URL: Repository, https://github.com/Tinnci/python-audio-super-resolution
Project-URL: Documentation, https://github.com/Tinnci/python-audio-super-resolution/blob/main/README.md
Project-URL: Issues, https://github.com/Tinnci/python-audio-super-resolution/issues
Author: SiTinc
License-Expression: MIT
License-File: LICENSE
Keywords: audio,audiosr,bandwidth-extension,music,speech,super-resolution
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: License :: OSI Approved :: MIT License
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3.14
Classifier: Topic :: Multimedia :: Sound/Audio
Classifier: Topic :: Multimedia :: Sound/Audio :: Analysis
Requires-Python: >=3.10
Requires-Dist: numpy>=1.23.5
Requires-Dist: scipy>=1.11
Requires-Dist: soundfile>=0.12
Requires-Dist: tqdm>=4.66
Provides-Extra: audiosr
Requires-Dist: audiosr==0.0.7; extra == 'audiosr'
Provides-Extra: download
Requires-Dist: huggingface-hub>=0.24; extra == 'download'
Provides-Extra: eval-fullref
Requires-Dist: pesq>=0.0.4; extra == 'eval-fullref'
Requires-Dist: pystoi>=0.4; extra == 'eval-fullref'
Provides-Extra: lavasr
Requires-Dist: torch>=2; extra == 'lavasr'
Provides-Extra: weights
Requires-Dist: safetensors>=0.4; extra == 'weights'
Description-Content-Type: text/markdown

<div align="center">

# Audio Super Resolution

[![PyPI version](https://badge.fury.io/py/audio-super-resolution.svg)](https://badge.fury.io/py/audio-super-resolution)
[![CI](https://github.com/Tinnci/python-audio-super-resolution/actions/workflows/ci.yml/badge.svg)](https://github.com/Tinnci/python-audio-super-resolution/actions/workflows/ci.yml)
[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](LICENSE)

</div>

Audio Super Resolution is a Python CLI and library for audio super-resolution and bandwidth extension. It ships with a deterministic `sinc-resample` baseline, optional external AudioSR support, and managed metadata for self-contained model backends.

The baseline package stays lightweight: normal inference is offline, model downloads are explicit, and heavyweight model dependencies live behind optional extras.

## Features

- CLI and Python API for single files, directory batches, dry runs, and recursive path-preserving output.
- Pluggable backend registry with `sinc-resample`, optional `audiosr`, and experimental `lavasr-compat`.
- Shared inference config for device, precision, chunking, preprocessing, seeds, and model cache paths.
- Import-light accelerator and runtime-provider metadata for CPU/CUDA/ROCm/XPU/MPS/DirectML planning.
- JSON run manifests, manifest comparison, and quality reports for regression workflows.
- Optional benchmark JSON reports for local accelerator/runtime validation.
- Explicit local weight resolution with multi-file manifests, size/SHA256 checks, and opt-in Hugging Face downloads.
- Pixi tasks for repeatable lint, format, typecheck, test, build, and package validation commands.

## Installation

Install from PyPI:

```sh
uv pip install audio-super-resolution
```

Install optional model/runtime extras only when needed. For example, LavaSR-compatible inference with
managed Hugging Face downloads uses:

```sh
uv pip install "audio-super-resolution[lavasr,download]"
```

Install the unreleased repository version from GitHub:

```sh
uv pip install "audio-super-resolution @ git+https://github.com/Tinnci/python-audio-super-resolution.git"
```

GitHub installs can include extras:

```sh
uv pip install "audio-super-resolution[lavasr,download] @ git+https://github.com/Tinnci/python-audio-super-resolution.git"
```

For local development:

```sh
git clone https://github.com/Tinnci/python-audio-super-resolution.git
cd python-audio-super-resolution
pixi install
```

Optional extras:

| Extra | Purpose |
| --- | --- |
| `audiosr` | External AudioSR wrapper. Use Python 3.10 because upstream dependencies are older. |
| `download` | Hugging Face model weight downloads. |
| `weights` | Optional safetensors loading helpers. |
| `lavasr` | Torch runtime for the experimental LavaSR-compatible backend. |

## CLI Quick Start

Enhance one file:

```sh
audio-super-res input.wav output.wav --target-sr 48000
```

If `output.wav` is omitted, the CLI writes next to the input as `input-sr48000.wav`.

Batch process a directory:

```sh
audio-super-res ./low-res-audio ./enhanced-audio --recursive --target-sr 48000
```

Preview or record a run:

```sh
audio-super-res ./low-res-audio ./enhanced-audio --recursive --dry-run --manifest plan.json
audio-super-res ./low-res-audio ./enhanced-audio --recursive --manifest run.json
audio-super-res --compare-manifests expected.json actual.json
```

Run a lightweight backend eval over clean references:

```sh
audio-super-res eval run --dataset evalsets/speech_clean_48k --backend sinc-resample --work-dir runs/sinc-work --output runs/sinc.json
audio-super-res eval compare runs/sinc.json runs/lavasr.json --output runs/comparison.json
```

List backends and models:

```sh
audio-super-res --list-backends
audio-super-res --list-models --list-format json
```

Run post-write quality checks:

```sh
audio-super-res input.wav output.wav --quality-report --fail-on-quality-issue
audio-super-res input.wav output.wav --quality-report-json quality.json
audio-super-res input.wav output.wav --benchmark-json benchmark.json
```

The shorter `audiosr` command is also available as an alias for `audio-super-res`.

## Models And Weights

Current backend status:

| Backend | Status |
| --- | --- |
| `sinc-resample` | Default deterministic baseline. |
| `audiosr` | Optional external package backend; upstream package owns its checkpoint behavior. |
| `lavasr-compat` | Experimental self-contained LavaSR v2 BWE path with managed weights. Gated real-weight download, torch smoke, and initial upstream parity validation pass. |

Use `audio-super-res --list-models --list-format json` for machine-readable comparison metadata, including task/domain, input and target sample rates, implementation family, I/O capabilities, accelerator and runtime-provider declarations, weight source/size/license, validation evidence, recommended use, and known limitations.

Accelerator install paths, `device=auto` behavior, runtime-provider selection, and gated benchmark guidance live in [docs/ACCELERATORS.md](docs/ACCELERATORS.md).

Future model candidates are tracked in the speech and general-audio reviews linked from [docs/README.md](docs/README.md). Candidate entries are not supported backends until they pass admission and validation.

Managed downloads are explicit. Normal enhancement only uses local verified files unless `--download-weights` is set:

```sh
audio-super-res --backend lavasr-compat --download-weights --prepare-model-cache
audio-super-res --backend lavasr-compat --verify-weights
```

Use an existing manifest:

```sh
audio-super-res input.wav output.wav \
  --backend lavasr-compat \
  --target-sr 48000 \
  --weights-manifest C:\path\to\lavasr-v2-bwe\manifest.json
```

Run the optional external AudioSR backend:

```sh
audio-super-res input.wav output.wav \
  --backend audiosr \
  --target-sr 48000 \
  --model-name basic \
  --device auto
```

## Python API

```python
from audio_super_resolution import AudioSuperResolver

resolver = AudioSuperResolver(target_sr=48000)
result = resolver.enhance("input.wav", "output.wav")

print(result.output_path)
print(result.sample_rate)
```

Batch planning and manifests:

```python
from audio_super_resolution import InferenceConfig, build_manifest, plan_enhancements

jobs = plan_enhancements("low-res-audio", "enhanced-audio", recursive=True)
manifest = build_manifest("dry-run", jobs, InferenceConfig(), backend="sinc-resample", target_sample_rate=48000)
```

Managed weights:

```python
from audio_super_resolution import (
    InferenceConfig,
    download_model_weights,
    resolve_model_weights,
    verify_model_weights,
)

download_model_weights("lavasr-v2-bwe")
verified = verify_model_weights("lavasr-v2-bwe")
weights = resolve_model_weights("lavasr-v2-bwe", InferenceConfig(model_cache_dir=verified.root_dir.parent))
model_path = weights.path_for("enhancer_v2/pytorch_model.bin")
```

## Development

```sh
pixi run lint
pixi run format-check
pixi run typecheck
pixi run test
pixi run build
pixi run metadata-check
pixi run wheel-check
```

Run optional real AudioSR integration only when model inference and upstream checkpoint handling are intended:

```sh
set AUDIO_SUPER_RESOLUTION_RUN_AUDIOSR_INTEGRATION=1
pixi run pytest tests/test_audiosr_integration.py
```

Real LavaSR weight download and torch smoke tests are also gated; see [tests/README.md](tests/README.md).

## Docker

```sh
docker build -t audio-super-resolution .
docker run --rm -v "%cd%":/workdir audio-super-resolution input.wav output.wav --target-sr 48000
```

On Unix-like shells, use `-v "$PWD":/workdir`.

## Project Docs

- [docs/README.md](docs/README.md): documentation map and ownership rules.
- [docs/ARCHITECTURE.md](docs/ARCHITECTURE.md): package layers, backend contract, and weight-management boundaries.
- [docs/ACCELERATORS.md](docs/ACCELERATORS.md): accelerator install strategy, runtime providers, and validation matrix.
- [docs/EVALUATION.md](docs/EVALUATION.md): eval/regression, golden compatibility, metrics, and manifests.
- [docs/BACKEND_CANDIDATES.md](docs/BACKEND_CANDIDATES.md): candidate evidence and implementation blockers.
- [docs/RELEASE.md](docs/RELEASE.md): merge, package, and release checklist.
- [ROADMAP.md](ROADMAP.md): prioritized next work and completed milestone summary.
- [CHANGELOG.md](CHANGELOG.md): release history and unreleased changes.
- [tests/README.md](tests/README.md): default and optional test strategy.
- [examples/](examples/): Python examples and sample JSON artifacts.

## Requirements

- Python 3.10 or newer
- Pixi 0.70 or newer for development; the committed `pixi.lock` uses lock-file version 7.
- libsndfile-compatible audio files for the default reader/writer

## License

This project is licensed under the MIT License. See [LICENSE](LICENSE) for details.

## Credits

Inspired by the project structure and user experience of [python-audio-separator](https://github.com/nomadkaraoke/python-audio-separator).
