Metadata-Version: 2.4
Name: paratran
Version: 0.6.0
Summary: CLI, REST interface, and MCP server for audio transcription using parakeet-mlx
Author: Brian Sunter
License-Expression: MIT
Project-URL: Homepage, https://github.com/briansunter/paratran
Project-URL: Repository, https://github.com/briansunter/paratran
Keywords: transcription,asr,speech-to-text,mlx,apple-silicon
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: Operating System :: MacOS
Classifier: Programming Language :: Python :: 3
Classifier: Topic :: Multimedia :: Sound/Audio :: Speech
Requires-Python: >=3.11
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: parakeet-mlx<1,>=0.5.0
Requires-Dist: fastapi<1,>=0.115
Requires-Dist: uvicorn[standard]<1,>=0.30
Requires-Dist: python-multipart<1,>=0.0.9
Requires-Dist: mcp<2,>=1.0
Dynamic: license-file

# Paratran

CLI, REST API, and MCP server for audio transcription on Apple Silicon, powered by [parakeet-mlx](https://github.com/senstella/parakeet-mlx).

The default model ([parakeet-tdt-0.6b-v3](https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3)) achieves 6.34% average WER across 8 English benchmarks and supports 25 languages. Runs ~30x faster than Whisper on Apple Silicon via MLX.

## Requirements

- macOS with Apple Silicon (M1/M2/M3/M4)
- Python 3.11+
- ~2 GB memory for the default model

## Quick Start

Transcribe audio files directly:

```bash
uvx paratran recording.wav
```

Or start the REST API server and transcribe via client mode (no model reload per file):

```bash
uvx paratran serve
uvx paratran -s http://localhost:8000 recording.wav
```

## Install

### uv (recommended)

```bash
uv tool install paratran
```

### pip

```bash
pip install paratran
```

### From source

```bash
git clone https://github.com/briansunter/paratran.git
cd paratran
uv sync
uv run paratran
```

## CLI Usage

```bash
# Transcribe a single file
paratran recording.wav

# Transcribe multiple files with verbose output
paratran -v file1.wav file2.mp3 file3.m4a

# Output as SRT subtitles
paratran --output-format srt recording.wav

# Output all formats (txt, json, srt, vtt)
paratran --output-format all --output-dir ./output recording.wav

# Use beam search decoding
paratran --decoding beam recording.wav

# Custom model and cache directory
paratran --model mlx-community/parakeet-tdt-1.1b-v2 --cache-dir /Volumes/Storage/models recording.wav
```

### Client Mode

Use `--server` / `-s` to send files to a running paratran server instead of transcribing locally. This avoids model loading time on every invocation — start the server once, then transcribe instantly.

```bash
# Start the server (loads model once)
paratran serve

# Transcribe via the server
paratran -s http://localhost:8000 recording.wav

# Output and transcription options work in client mode
paratran -s http://localhost:8000 --output-format all --output-dir ./output -v recording.wav

# Set the server URL via environment variable
export PARATRAN_SERVER=http://localhost:8000
paratran recording.wav  # automatically uses the server
```

### CLI Options

| Flag | Default | Description |
|------|---------|-------------|
| `-s`, `--server` | | URL of a running paratran server |
| `--api-key` | | Bearer token for an authenticated server |
| `--timeout` | `60` | Server request timeout in seconds |
| `--model` | `mlx-community/parakeet-tdt-0.6b-v3` | HF model ID or local path |
| `--cache-dir` | HuggingFace default | Model cache directory |
| `--output-dir` | `.` | Output directory |
| `--output-format` | `txt` | `txt`, `json`, `srt`, `vtt`, or `all` |
| `--decoding` | `greedy` | `greedy` or `beam` |
| `--chunk-duration` | `120` | Chunk duration in seconds (0 to disable) |
| `--overlap-duration` | `15` | Overlap between chunks |
| `--beam-size` | `5` | Beam size (beam decoding) |
| `--length-penalty` | `0.013` | Length penalty (beam decoding) |
| `--patience` | `3.5` | Patience (beam decoding) |
| `--duration-reward` | `0.67` | Duration reward (beam decoding) |
| `--max-words` | | Max words per sentence |
| `--silence-gap` | | Split at silence gaps (seconds) |
| `--max-duration` | | Max sentence duration (seconds) |
| `--fp32` | | Use FP32 precision instead of BF16 |
| `-v` | | Verbose output |

Environment variables: `PARATRAN_MODEL`, `PARATRAN_MODEL_DIR`, `PARATRAN_SERVER`, `PARATRAN_API_KEY`.

When using client mode, configure `--model` and `--cache-dir` on the running server; those options do not change a remote server.

## REST API Server

```bash
# Start server with local-only defaults
paratran serve

# Custom host, port, and model cache
paratran serve --host 127.0.0.1 --port 9000 --cache-dir /Volumes/Storage/models

# Expose the server only with an API key
paratran serve --host 0.0.0.0 --api-key "$PARATRAN_API_KEY"
```

The server defaults to `127.0.0.1`, limits uploads to 512 MB, and processes one transcription at a time. Non-loopback hosts require `--api-key`.

## API

The REST API is compatible with the [OpenAI Audio Transcription API](https://platform.openai.com/docs/api-reference/audio/createTranscription).

### `GET /health`

```bash
curl http://localhost:8000/health
```

```json
{
  "status": "ok",
  "model": "mlx-community/parakeet-tdt-0.6b-v3",
  "model_dir": "/Volumes/Storage/models"
}
```

### `POST /v1/audio/transcriptions`

Upload an audio file (wav, mp3, flac, m4a, ogg, webm):

```bash
curl http://localhost:8000/v1/audio/transcriptions \
  -F "file=@recording.m4a" \
  -F "model=parakeet"
```

#### OpenAI-compatible parameters

| Parameter | Default | Description |
|-----------|---------|-------------|
| `file` | *(required)* | Audio file to transcribe |
| `model` | | Accepted for compatibility; uses the configured model |
| `response_format` | `json` | `json`, `text`, `srt`, `vtt`, or `verbose_json` |
| `language` | | Accepted for compatibility; language is auto-detected |
| `prompt` | | Accepted for compatibility; prompts are not applied |
| `temperature` | | Accepted for compatibility; temperature is not applied |

#### Paratran-specific parameters

| Parameter | Default | Description |
|-----------|---------|-------------|
| `decoding` | `greedy` | `greedy` or `beam` |
| `beam_size` | `5` | Beam size (beam decoding) |
| `length_penalty` | `0.013` | Length penalty (beam decoding) |
| `patience` | `3.5` | Patience (beam decoding) |
| `duration_reward` | `0.67` | Duration reward (beam decoding) |
| `max_words` | | Max words per sentence |
| `silence_gap` | | Split at silence gaps (seconds) |
| `max_duration` | | Max sentence duration (seconds) |
| `chunk_duration` | `120` | Chunk duration for long audio (seconds); `0` disables chunking |
| `overlap_duration` | `15.0` | Overlap between chunks (seconds) |
| `fp32` | `false` | Use FP32 instead of BF16 |

#### Response formats

**`json`** (default):
```json
{"text": "Hello world, this is a test."}
```

**`verbose_json`** (also includes `processing_time`):
```json
{
  "task": "transcribe",
  "duration": 3.52,
  "processing_time": 0.176,
  "text": "Hello world, this is a test.",
  "segments": [
    {"id": 0, "start": 0.0, "end": 3.52, "text": "Hello world, this is a test."}
  ],
  "words": [
    {"word": "Hello", "start": 0.0, "end": 0.48},
    {"word": " world", "start": 0.48, "end": 0.8}
  ]
}
```

**`text`**: Returns plain text. **`srt`** / **`vtt`**: Returns subtitles.

Interactive API docs are available at `http://localhost:8000/docs`.

## MCP Server

Paratran includes an MCP server so Claude Code, Claude Desktop, or any MCP client can transcribe audio files directly. Supports both stdio and streamable HTTP transports.

### Claude Code (stdio)

Add to `.claude/settings.json`:

```json
{
  "mcpServers": {
    "paratran": {
      "command": "uvx",
      "args": ["--from", "paratran", "paratran-mcp"]
    }
  }
}
```

### Claude Desktop (stdio)

Add to `~/Library/Application Support/Claude/claude_desktop_config.json`:

```json
{
  "mcpServers": {
    "paratran": {
      "command": "uvx",
      "args": ["--from", "paratran", "paratran-mcp"]
    }
  }
}
```

Optionally set `PARATRAN_MODEL_DIR` in the `env` block to customize the model cache location.

### Streamable HTTP

Run the MCP server over HTTP for local or multi-client access:

```bash
paratran-mcp --transport streamable-http --host 127.0.0.1 --port 8000
```

The MCP endpoint is available at `http://localhost:8000/mcp`. For HTTP MCP on a non-loopback host, pass both `--allowed-root` and `--api-key`; the key is accepted as an `Authorization: Bearer` token. Loopback HTTP can also be protected with `--api-key` when multiple local clients share the server.

### MCP Tool

The `transcribe` tool accepts an absolute file path and all the same transcription options as the REST interface. `--allowed-root` restricts paths to a directory, which is recommended for HTTP MCP.

## License

MIT
