Metadata-Version: 2.5
Name: MyanmarTTS
Version: 1.0.2
Summary: A from-scratch Burmese (Myanmar) text-to-speech model (CC0).
Project-URL: Homepage, https://huggingface.co/freococo/MyanmarTTS
Project-URL: Repository, https://huggingface.co/freococo/MyanmarTTS
Author: freococo
License-File: LICENSE
Keywords: burmese,myanmar,speech,text-to-speech,tts
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Topic :: Multimedia :: Sound/Audio :: Speech
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Python: >=3.9
Requires-Dist: huggingface-hub>=0.20
Requires-Dist: numba
Requires-Dist: numpy<2.3
Requires-Dist: safetensors
Requires-Dist: scipy
Requires-Dist: soundfile
Requires-Dist: torch>=2.0
Requires-Dist: torchaudio>=2.0
Requires-Dist: torchdiffeq
Description-Content-Type: text/markdown

---
license: cc0-1.0
language:
  - my
tags:
  - text-to-speech
  - tts
  - burmese
  - myanmar
  - stabletts
  - from-scratch
pipeline_tag: text-to-speech
---

# MyanmarTTS

A from-scratch Burmese (Myanmar) text-to-speech model. 31M parameters, ~63 MB (fp16).

- **Language**: Burmese (မြန်မာဘာသာ)
- **Training data**: ~1.34M samples (news + real-world audio)
- **Training steps**: 81,000
- **Training hardware**: A100-80GB (Colab Pro+), ~13 hours, ~88 compute units
- **Inference hardware**: Free Colab T4, any modern GPU, or CPU (slower)
- **Architecture**: StableTTS (DiT + flow matching) + Vocos vocoder
- **License**: **CC0 1.0** (public domain, no attribution required)

---

## Sample Output

**Sample 0**

**Text:** "မြန်မာလူမျိုးများဟာ အလွန် ယဥ်ကျေးသိမ်မွေ့ပြီး ဧည့်သည်များကို ပျူပျူငှာငှာ လှိုက်လှိုက်လှဲလှဲနဲ့ ကြိုဆိုကြပါတယ်"

<audio controls src="https://huggingface.co/freococo/MyanmarTTS/resolve/main/samples/sample_0.wav"></audio>

**Sample 1**

**Text:** "မင်္ဂလာပါရှင် ကျွန်မကတော့ မြန်မာလူမျိုး ကရင်တိုင်းရင်းသူ အမျိုးသမီးလေး တစ်ဦး ဖြစ်ပါတယ်"

<audio controls src="https://huggingface.co/freococo/MyanmarTTS/resolve/main/samples/sample_1.wav"></audio>

**Sample 2**

**Text:** "ဒီနေ့ ကျွန်မတို့ရဲ့ တီတီအက်စ် စနစ်သစ်လေး မော်ဒယ်အသစ်လေးတစ်ခုကို အောင်မြင်စွာ လေ့ကျင့် သင်ကြားနိုင်ခဲ့ပါတယ်"

<audio controls src="https://huggingface.co/freococo/MyanmarTTS/resolve/main/samples/sample_2.wav"></audio>

**Sample 3**

**Text:** "လူသားတိုင်း လူသားတိုင်း ကိုယ်စိတ်နှစ်ဖြာ ကျန်းမာရွှင်လန်းပြီး စီးပွားလာဘ်လာဘတွေ ဒီရေအလား ကြီးပွား တိုးတက်နိုင်ကြပါစေ"

<audio controls src="https://huggingface.co/freococo/MyanmarTTS/resolve/main/samples/sample_3.wav"></audio>

**Sample 4**

**Text:** "ဒီအသံထုတ်စနစ်လေးကို အသုံးပြုသူတိုင်း ကျန်းမာချမ်းသာပြီး လိုရာဆန္ဒတွေ တလုံးတဝတည်း ပြည့်စုံနိုင်ကြပါစေ"

<audio controls src="https://huggingface.co/freococo/MyanmarTTS/resolve/main/samples/sample_4.wav"></audio>

---


## Reference Audio

The reference voice used for voice cloning during inference:

<audio controls src="https://huggingface.co/freococo/MyanmarTTS/resolve/main/samples/sample_0.wav"></audio>

## Quick Start

```python
import torch, soundfile as sf
from huggingface_hub import hf_hub_download
from api import StableTTSAPI
from burmese import burmese_to_ipa2

# Download model assets
model_path = hf_hub_download("freococo/MyanmarTTS", "model_fp16.pt")
vocos_path = hf_hub_download("freococo/MyanmarTTS", "vocos.pt")
ref_path   = hf_hub_download("freococo/MyanmarTTS", "samples/sample_0.wav")

# Initialize model
model = StableTTSAPI(model_path, vocos_path, "vocos").to("cuda")
model.g2p_mapping["burmese"] = burmese_to_ipa2

# Inference
text = "လူသားတွေ အားလုံးကို အရမ်း ချစ်ပါတယ်ရှင့်"
audio, _ = model.inference(text, ref_path, "burmese", step=32, solver="dopri5", cfg=3.0)
sf.write("output.wav", audio.squeeze(0).cpu().numpy(), 44100)
```

> *Translation: "I love all human beings very much." (female polite form)*

---

## Files

| File | Size | Purpose | License |
| :--- | :--- | :--- | :--- |
| `model_fp16.pt` | 63 MB | Default model weights | CC0 |
| `model_fp32.pt` | 126 MB | Full precision weights | CC0 |
| `vocos.pt` | 57 MB | Mel-to-waveform vocoder | MIT (KdaiP) |
| `vocab.txt` | 98 tokens | Burmese character-level vocab | CC0 |
| `config.json` | -- | Mel + model configuration | CC0 |
| `symbols.py`, `burmese.py` | -- | Text frontend & G2P | MIT (adapted) |
| `api.py` | -- | Inference wrapper API | MIT (adapted) |
| `samples/` | -- | Demo audio WAV files | CC0 |
| `transcripts.json` | -- | Sample text/audio mapping | CC0 |
| `NOTICE` | -- | Full license summary | -- |

---

## Training Progression

See [`checkpoint_comparison`](checkpoint_comparison) for sample audio generated at steps **8k**, **33k**, **44k**, **55k**, and **81k**.

---

## Acknowledgments

This work would not exist without the generous open-source community and AI assistance:

- **StableTTS by KdaiP** — The DiT + flow-matching architecture and training code (MIT).
- **Vocos** — Pretrained mel-to-wav vocoder (MIT).
- **DeepSeek AI** — Provided AI pair-programming and engineering assistance throughout the project. From data pipeline design and architecture choices to resolving CUDA OOM bottlenecks, DeepSeek's guidance was instrumental at every stage.
- **The Burmese open-data community** — For providing the audio corpora that made training possible.

### Special Thanks
To **DeepSeek AI** — a true engineering partner from the first line of code to the final deployment. This model exists because of that collaboration.

---

## License

- **Model weights and generated audio**: [CC0 1.0 Universal](https://creativecommons.org/publicdomain/zero/1.0/) (Public Domain).
- **Supporting code**: Adapted from [KdaiP/StableTTS](https://github.com/KdaiP/StableTTS) (MIT).
- **Vocoder**: From KdaiP/StableTTS1.1 (MIT).

*See `NOTICE` for additional details.*