Metadata-Version: 2.4
Name: kittenml
Version: 0.9.0
Summary: Text-to-speech with voice cloning and expression control, plus ultra-lightweight models that run on CPU
Home-page: https://github.com/KittenML/KittenTTS
Author: KittenML
Author-email: 
License: Apache 2.0
Project-URL: Homepage, https://github.com/KittenML/KittenTTS
Project-URL: Repository, https://github.com/KittenML/KittenTTS
Project-URL: Issues, https://github.com/KittenML/KittenTTS/issues
Keywords: text-to-speech,tts,speech-synthesis,voice-cloning,neural-networks,onnx
Classifier: Topic :: Multimedia :: Sound/Audio :: Speech
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: torch>=2.6
Requires-Dist: torchaudio>=2.6
Requires-Dist: transformers>=4.55
Requires-Dist: safetensors
Requires-Dist: numpy
Requires-Dist: scipy
Requires-Dist: soundfile
Requires-Dist: librosa
Requires-Dist: huggingface_hub
Requires-Dist: einops
Requires-Dist: diffusers
Requires-Dist: omegaconf
Requires-Dist: hydra-core
Requires-Dist: s3tokenizer
Requires-Dist: resemble-perth
Requires-Dist: pyloudnorm
Requires-Dist: kitten-text-processing>=0.1.0
Requires-Dist: espeakng_loader
Requires-Dist: phonemizer
Requires-Dist: onnxruntime
Dynamic: home-page
Dynamic: license-file
Dynamic: requires-python

# Kitten TTS

<p align="center">
  <img width="607"   alt="Kitten TTS" src="https://github.com/user-attachments/assets/6e24bdc1-9750-4416-ad8b-275bdc30b798" />
</p>

<p align="center">
  <a href="https://huggingface.co/spaces/KittenML/KittenTTS-Demo"><img src="https://img.shields.io/badge/Demo-Hugging%20Face%20Spaces-orange" alt="Hugging Face Demo"></a>
  <a href="https://discord.com/invite/VJ86W4SURW"><img src="https://img.shields.io/badge/Discord-Join%20Community-5865F2?logo=discord&logoColor=white" alt="Discord"></a>
  <a href="https://kittenml.com"><img src="https://img.shields.io/badge/Website-kittenml.com-blue" alt="Website"></a>
  <a href="LICENSE"><img src="https://img.shields.io/badge/License-Apache_2.0-green.svg" alt="License"></a>
</p>

> ## **New:** Free Kitten TTS API available at  [https://platform.kittenml.com](https://platform.kittenml.com/)

Kitten TTS is an open-source text-to-speech library. Its flagship model, **KittenTTS 2**, is a
1.7B-parameter speech language model with in-context voice cloning and expression control: give it
five seconds of anyone's voice and it speaks your text in that voice.

The library also ships the original [**lightweight ONNX models**](docs/onnx-models.md),
15M-80M parameters, which run on CPU without a GPU. Both families load through the same
`KittenTTS(...)` constructor.

> **Status:** Developer preview -- APIs may change between releases.

**Commercial support is available.** For integration assistance, custom voices, or enterprise licensing, [contact us](https://docs.google.com/forms/d/e/1FAIpQLSc49erSr7jmh3H2yeqH4oZyRRuXm0ROuQdOgWguTzx6SMdUnQ/viewform?usp=preview).

## Table of Contents

- [Features](#features)
- [Available Models](#available-models)
- [Demo](#demo)
- [Quick Start](#quick-start)
- [Voice cloning](#voice-cloning)
- [Expression controls](#expression-controls)
- [Long text and streaming](#long-text-and-streaming)
- [Documentation](#documentation)
- [System Requirements](#system-requirements)
- [Roadmap](#roadmap)
- [Commercial Support](#commercial-support)
- [Community and Support](#community-and-support)
- [License](#license)

## Features

- **Voice cloning** -- Clone any speaker from 5-30 seconds of audio, no fine-tuning
- **47 built-in voices** -- Including the eight from KittenTTS 0.8 and nine non-English
- **Multilingual** -- Ten languages: English, Arabic, Chinese, French, German, Hindi, Italian, Portuguese, Russian, Spanish
- **Expression control** -- `[emotion]` tags, inline `<event>` tags, and `(((emphasis)))` spans
- **Decoding presets** -- Trade stability against expressiveness per request
- **Long-form text** -- Sentence-aware chunking with seamless joins
- **Text preprocessing** -- Numbers, currencies, dates, units and abbreviations expanded automatically
- **24 kHz output** -- High-quality audio at a standard sample rate
- **Also runs on CPU** -- Via the [lightweight ONNX models](docs/onnx-models.md), from 25 MB

## Available Models

**KittenTTS 2** -- speech language model, requires a GPU:

| Model | Parameters | Size | Voices | Download |
|---|---|---|---|---|
| kitten-tts-2 | 1.7B | 3.5 GB | 47 + cloning | [KittenML/kitten-tts-2](https://huggingface.co/KittenML/kitten-tts-2) |

**Lightweight ONNX** -- runs on CPU, no GPU required:

| Model | Parameters | Size | Download |
|---|---|---|---|
| kitten-tts-mini | 80M | 80 MB | [KittenML/kitten-tts-mini-0.8](https://huggingface.co/KittenML/kitten-tts-mini-0.8) |
| kitten-tts-micro | 40M | 41 MB | [KittenML/kitten-tts-micro-0.8](https://huggingface.co/KittenML/kitten-tts-micro-0.8) |
| kitten-tts-nano | 15M | 56 MB | [KittenML/kitten-tts-nano-0.8](https://huggingface.co/KittenML/kitten-tts-nano-0.8-fp32) |
| kitten-tts-nano (int8) | 15M | 25 MB | [KittenML/kitten-tts-nano-0.8-int8](https://huggingface.co/KittenML/kitten-tts-nano-0.8-int8) |

> **Note:** Some users have reported issues with the `kitten-tts-nano-0.8-int8` model. If you encounter problems, please [open an issue](https://github.com/KittenML/KittenTTS/issues).

## Demo

https://github.com/user-attachments/assets/d80120f2-c751-407e-a166-068dd1dd9e8d

### Try it online

Try Kitten TTS directly in your browser on [Hugging Face Spaces](https://huggingface.co/spaces/KittenML/KittenTTS-Demo).

## Quick Start

### Prerequisites

- Python 3.9 or later
- A CUDA GPU with roughly 8 GB free
- About 4 GB of disk space for the model

### Installation

```bash
pip install kittenml
```

That is the whole install. KittenTTS 2 is the default model, voice cloning is included, and no
Hugging Face login is needed -- every weight the model uses ships in its own repository. It also
pulls the small ONNX runtime, so the [lightweight models](docs/onnx-models.md) work from the
same install.

### Basic usage

```python
from kittenml import KittenTTS
import soundfile as sf

m = KittenTTS("KittenML/kitten-tts-2")

audio = m.generate("One day, a little girl named Lily found a needle in her room.",
                   voice="Bruno")
sf.write("output.wav", audio, m.sample_rate)
```

`m.available_voices` lists all 47 built-in voices, described in
[voices and expression](docs/voices-and-expression.md). Bella, Jasper, Luna, Bruno, Rosie, Hugo,
Kiki and Leo are the same speakers as in KittenTTS 0.8, so code written against the ONNX models
keeps working.

```python
# Trade stability against expressiveness
audio = m.generate("Hello, world.", voice="Luna", preset="expressive")

# Save directly to a file
m.generate_to_file("Hello, world.", "output.wav", voice="Bruno")
```

## Voice cloning

Pass `reference=` instead of `voice=` — same method, the recording just replaces the built-in
speaker. Give it 5-30 seconds of a single speaker. The transcript is part of the prompt, but you
do not have to type it; Whisper fills it in when omitted.

```python
audio = m.generate("This is my own voice, cloned.", reference="my_voice.wav")

# Supplying the transcript skips the Whisper pass
audio = m.generate("This is my own voice.", reference="my_voice.wav",
                   reference_text="what is actually said in the clip")
```

The reference feeds the model by two independent routes -- a speaker embedding through the
model's projection head, and the clip itself as codec tokens in the prompt -- so identity
survives even when one route is weak.

Measured on the built-in voices: a generated clip scores 0.49-0.72 speaker similarity against its
own reference and 0.01-0.18 against the other 37, and cloning an unseen recording scores 0.81
against that recording.

## Expression controls

```python
audio = m.generate(
    "[joyful] We actually won the grant <laugh> I can (((hardly))) believe it!",
    voice="Kiki",
    preset="expressive",
)
```

A leading `[emotion]` tag, inline `<event>` tags and `(((emphasis)))` spans reach the model as
markup rather than being spoken, and automatically enable its expression conditioning. Ten
emotions and ten vocal events are recognised -- see
[voices and expression](docs/voices-and-expression.md) for the full lists and what is not
covered.

## Long text and streaming

Long input is split on sentence boundaries and synthesized chunk by chunk, then joined with
silence trimming and short edge fades so the seams are inaudible. This is automatic: the model is
reliable on short inputs but truncates or drifts into repetition when asked for a whole script in
one pass.

To start playing before the whole thing is ready, stream it:

```python
for chunk in m.generate_stream(long_text, voice="Luna"):
    play(chunk)          # each chunk is a numpy array at m.sample_rate
```

It takes the same arguments as `generate`, so `reference=` streams a cloned voice too.

Streaming is **chunk-level, not token-level**: a chunk is generated and vocoded in full before it
is yielded, so the first chunk still costs its own generation time. On an A100, a 936-character
passage yielded its first 21 s of audio after 17 s and finished 54 s of audio in 42 s of wall
clock -- so playback keeps ahead of generation, but there is a real initial delay.

Two consequences worth knowing:

- **Short text does not stream.** Input that fits in one chunk (under roughly 380 characters,
  and short trailing pieces get merged into their neighbour) yields exactly one chunk, so
  `generate_stream` behaves like `generate`.
- **Chunks are yielded raw.** `generate` post-processes the seams -- trimming each segment's edge
  silence, adding short fades and one consistent pause -- which a streaming caller cannot do
  without waiting for the next chunk. Concatenating streamed chunks directly gives slightly
  rougher joins than `generate` on the same text.

## Documentation

| | |
|---|---|
| [API reference](docs/api.md) | Every argument to `generate`, streaming, and the advanced knobs |
| [Voices and expression](docs/voices-and-expression.md) | The 47 voices, emotion and vocal-event tags, the ten languages |
| [Decoders](docs/decoders.md) | How audio is decoded, and the smaller quantised decoders |
| [Text normalization](docs/text-normalization.md) | How written text becomes spoken text |
| [Architecture](docs/architecture.md) | What the model is, package layout, vendored components |
| [Lightweight ONNX models](docs/onnx-models.md) | The CPU models, 15M-80M parameters, and their API |

## System Requirements

**KittenTTS 2**

- **Operating system:** Linux or Windows (CUDA)
- **Python:** 3.9 or later
- **Hardware:** CUDA GPU with roughly 8 GB free
- **Disk space:** About 4 GB for the model

**Lightweight ONNX models**

- **Operating system:** Linux, macOS, or Windows
- **Python:** 3.8 or later
- **Hardware:** Runs on CPU; no GPU required
- **Disk space:** 25-80 MB depending on model variant

A virtual environment (conda, venv, or similar) is recommended to avoid dependency conflicts.

## Roadmap

- [ ] Release optimized inference engine
- [ ] Release mobile SDK
- [x] Release higher quality TTS models (KittenTTS 2)
- [ ] Release multilingual TTS
- [ ] Release KittenASR
- [ ] Need anything else? [Let us know](https://github.com/KittenML/KittenTTS/issues)

## Commercial Support

We offer commercial support for teams integrating Kitten TTS into their products. This includes integration assistance, custom voice development, and enterprise licensing.

[Contact us](https://docs.google.com/forms/d/e/1FAIpQLSc49erSr7jmh3H2yeqH4oZyRRuXm0ROuQdOgWguTzx6SMdUnQ/viewform?usp=preview) or email info@stellonlabs.com to discuss your requirements.

## Community and Support

- **Discord:** [Join the community](https://discord.com/invite/VJ86W4SURW)
- **Website:** [kittenml.com](https://kittenml.com)
- **Custom support:** [Request form](https://docs.google.com/forms/d/e/1FAIpQLSc49erSr7jmh3H2yeqH4oZyRRuXm0ROuQdOgWguTzx6SMdUnQ/viewform?usp=preview)
- **Email:** info@stellonlabs.com
- **Issues:** [GitHub Issues](https://github.com/KittenML/KittenTTS/issues)

## License

This project is licensed under the [Apache License 2.0](LICENSE).
