Metadata-Version: 2.4
Name: generic-audio-transcriber
Version: 0.1.0
Summary: Drop-in speech-to-text for any Python project: audio bytes in, JSON-ready text out.
Author: Eduardo Milani
License: MIT
Project-URL: Repository, https://github.com/EduardoMilani8/generic-audio-transcriber
Keywords: speech-to-text,transcription,whisper,audio,asr
Classifier: License :: OSI Approved :: MIT License
Classifier: Programming Language :: Python :: 3
Classifier: Topic :: Multimedia :: Sound/Audio :: Speech
Requires-Python: >=3.9
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: faster-whisper>=1.0
Requires-Dist: av<19,>=11
Provides-Extra: dev
Requires-Dist: pytest>=8; extra == "dev"
Dynamic: license-file

# generic-audio-transcriber

Drop-in speech-to-text for any Python project. Audio bytes in, JSON-ready text out.

```python
from audio_transcriber import transcribe

result = transcribe(audio_bytes)   # mp3, wav, webm, ogg, m4a... detected automatically
print(result.text)                 # "Hello, this is a test."
print(result.to_json())            # {"text": "...", "language": "en", ...}
```

No API keys, no servers, no cloud. It runs locally on the CPU, and the model is
downloaded automatically the first time you use it.

## Why

Many apps need the same small feature: the user sends or records audio, and the app
needs the text. This package is that feature as one function call, so you can add it to
a project without learning anything about speech recognition.

Under the hood it uses [faster-whisper](https://github.com/SYSTRAN/faster-whisper)
(OpenAI's Whisper model), which is accurate, supports about 99 languages, and runs well
on a plain CPU. You never have to train or tune a model.

## Install

```bash
pip install git+https://github.com/EduardoMilani8/generic-audio-transcriber.git
```

Requires Python 3.9+. You do **not** need to install ffmpeg: audio decoding is bundled.

## Usage

### From bytes (the main use case)

```python
from audio_transcriber import transcribe

result = transcribe(audio_bytes)
```

### From a file path or file object

```python
transcribe("recording.mp3")

with open("recording.wav", "rb") as f:
    transcribe(f)
```

### Inside a web endpoint

```python
from fastapi import FastAPI, UploadFile
from audio_transcriber import transcribe

app = FastAPI()

@app.post("/transcribe")
async def endpoint(file: UploadFile):
    return transcribe(await file.read()).to_dict()
```

### The result

```python
result.text                  # full transcript
result.language              # detected language code, e.g. "pt"
result.language_probability  # confidence of the language detection
result.duration              # audio length in seconds
result.segments              # timed pieces: .start, .end, .text
result.to_dict()             # plain dict
result.to_json()             # JSON string (non-ASCII kept readable)
```

Silent audio gives an empty `text`, not an error.

## Options

```python
transcribe(
    audio,
    model="small",         # tiny | base | small | medium | large-v3
    language=None,         # "pt", "en", ... None = auto-detect
    device="cpu",          # "cpu" | "cuda" | "auto"
    compute_type="int8",   # "int8" is light; "float16" suits GPUs
    beam_size=5,
    vad_filter=True,       # skip silence, reduces made-up text
)
```

Choosing a model is a trade-off between weight and accuracy:

| Model | Size on disk | Speed | Accuracy |
|---|---|---|---|
| `tiny` | smallest | fastest | lowest |
| `base` | small | fast | fair |
| `small` (default) | medium | good | good |
| `medium` / `large-v3` | large | slow on CPU | best |

The model is loaded once per process and reused, so only the first call is slow.
Setting `language` explicitly is faster and more accurate than auto-detection.

## Errors

```python
from audio_transcriber import transcribe, TranscriptionError

try:
    transcribe(data)
except TranscriptionError as e:
    ...  # empty input or audio that cannot be decoded
```

## Command line

```bash
audio-transcriber recording.mp3            # prints JSON
audio-transcriber recording.mp3 --text     # prints plain text
audio-transcriber recording.mp3 --model base --language pt
cat recording.mp3 | audio-transcriber -    # read bytes from stdin
```

## Development

```bash
python -m venv .venv && source .venv/bin/activate
pip install -e ".[dev]"
pytest                                # fast unit tests, no model needed
RUN_SLOW=1 pytest -m slow             # end-to-end tests, downloads the tiny model
```

## Notes

- The default device is `cpu` because it works on every machine. Pass `device="cuda"`
  if you have a working CUDA setup and want GPU speed.
- `av` is pinned below version 19 because faster-whisper still uses an argument that
  newer releases removed. The pin can be lifted once faster-whisper fixes it.

## License

MIT
