Metadata-Version: 2.4
Name: vox-dia
Version: 0.2.15
Summary: Dia TTS adapter for Vox
Requires-Python: >=3.11
Requires-Dist: numpy<2.4,>=1.26.0
Requires-Dist: soundfile>=0.13.1
Requires-Dist: vox-runtime>=0.2.2
Description-Content-Type: text/markdown

# vox-dia

Dia TTS adapter package for Vox.

## Included adapter

- `dia-tts-torch` — Dia 1.6B text-to-speech backend

## Install

```bash
pip install vox-dia
```

## Runtime Dependencies

Dia uses an isolated target runtime at `$VOX_HOME/runtime/dia` for backend
packages that must not leak into the base Vox app environment.

The adapter package itself is intentionally lightweight and does not install
Torch. The current Dia backend uses Hugging Face Transformers and requires a
Vox runtime/image that already provides PyTorch with CUDA. CPU execution is not
supported by the official Dia Transformers runtime used by this adapter.

The current Vox adapter is classified for Linux x86_64 CUDA/Torch runtimes.
CPU, ONNX, and Spark/ARM NVIDIA paths are not currently production-supported by
this adapter. Upstream Dia 1.6B documentation says the model has only been
tested on GPUs with PyTorch/CUDA and that CPU support is future work:
https://github.com/nari-labs/dia#hardware-and-inference-speed

Plan for at least 12GiB of usable VRAM budget: upstream reports that the full
Dia 1.6B model requires around 10GB of VRAM, and Vox budgets the adapter's 10GB
model estimate plus the deployment's configured VRAM headroom. A server started
with `--max-vram 10GiB --vram-headroom 1GiB` will reject Dia at load time.

During `vox pull`, the adapter verifies or installs the Dia-capable
Transformers runtime into `$VOX_HOME/runtime/dia`. It uses the released
`transformers==4.57.6` runtime rather than a moving source checkout, so clean
pulls are reproducible. Model weights remain in the normal Vox model store.

## Use with Vox

```bash
vox pull dia-tts:1.6b
```

Dia prompt control is text-driven. Use speaker tags and non-verbal markers in
the input text, for example:

```text
[S1] This is Dia speaking. (laughs) [S2] Keep the tags sparse or artifacts can appear.
```

The adapter supports Dia's audio-prompt voice cloning path when Vox resolves a
stored cloned voice into both `reference_audio` and `reference_text`. Dia needs
the reference transcript for the reference clip; requests with reference audio
but no reference text are rejected clearly.

The adapter also exposes Dia generation parameters through Vox synthesis
`params`:

- `max_new_tokens` (integer, default `3072`, range `1..8192`)
- `guidance_scale` (number, default `3.0`, range `0..10`)
- `temperature` (number, default `1.8`, range `0..3`)
- `top_p` (number, default `0.9`, range `0..1`)
- `top_k` (integer, default `45`, range `0..200`)
