Metadata-Version: 2.4
Name: jackai-stt-mcp
Version: 0.1.1
Summary: Local MCP server that transcribes audio files with OpenAI — point it at a path, get the text back
Author: JackAI
License: MIT
Project-URL: Homepage, https://github.com/jack-ai-net/stt-mcp
Project-URL: Issues, https://github.com/jack-ai-net/stt-mcp/issues
Keywords: mcp,speech-to-text,transcription,whisper,openai,arabic
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: MIT License
Classifier: Programming Language :: Python :: 3
Classifier: Topic :: Multimedia :: Sound/Audio :: Speech
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: mcp[cli]<2.0,>=1.2.0
Requires-Dist: httpx>=0.27.0
Requires-Dist: pydantic>=2.6.0
Provides-Extra: dev
Requires-Dist: pytest>=8.0; extra == "dev"
Requires-Dist: pytest-anyio>=0.0.0; extra == "dev"
Requires-Dist: anyio>=4.0; extra == "dev"
Dynamic: license-file

# jackai-stt-mcp

[![tests](https://github.com/jack-ai-net/stt-mcp/actions/workflows/test.yml/badge.svg)](https://github.com/jack-ai-net/stt-mcp/actions/workflows/test.yml) [![PyPI](https://img.shields.io/pypi/v/jackai-stt-mcp)](https://pypi.org/project/jackai-stt-mcp/) [![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT)

Transcribe audio by asking for it in plain language:

> **You:** *transcribe ~/Downloads/voice.opus*
>
> **Claude:** تمام تمام، موضوع اللغة العربية إن شاء الله محلول…

An MCP server that runs on your own machine, so the assistant can open the file
directly. Nothing to run, no commands to memorise — you mention a file and the
assistant does the rest.

No uploading, no copying, no base64. Your OpenAI key never leaves your computer.

Arabic works well, including dialect.

## Install

Add this to your MCP client's config. `uvx` fetches and runs the package on
first use — nothing to install by hand.

```json
{
  "servers": {
    "jackai-stt": {
      "command": "uvx",
      "args": ["jackai-stt-mcp"],
      "env": {
        "OPENAI_API_KEY": "sk-proj-..."
      }
    }
  }
}
```

Where the config lives:

| Client | File |
|---|---|
| VS Code | `~/Library/Application Support/Code/User/mcp.json` (macOS) |
| Claude Desktop | `~/Library/Application Support/Claude/claude_desktop_config.json` |
| Claude Code | `.mcp.json` in the project, or `claude mcp add` |
| Cursor | `~/.cursor/mcp.json` |

Restart the client, then ask it to transcribe something.

Get an API key at [platform.openai.com/api-keys](https://platform.openai.com/api-keys).
The account needs credit; transcription runs about **$0.0045 per minute**.

## Usage

Nothing to run, no syntax to learn. Talk to the assistant the way you normally
would and mention the file:

> *transcribe ~/Downloads/voice.opus*
>
> *what does the voice note on my desktop say?*
>
> *read ./recordings/call.m4a and summarise what the client is asking for*
>
> *transcribe meeting.mp3 and tell me who said what*

The assistant recognises that it needs this tool and fills in the arguments
from what you said. You never call it yourself or write out its parameters.

That last example needs speaker labels, which only one model produces. Mention
it in passing and the assistant picks it:

> *transcribe meeting.mp3 with the diarize model*

Same for a language it keeps mishearing (*"it's Egyptian Arabic"*), or names it
should spell correctly (*"the speakers are Ahmad and Sara"*). Ordinary words —
the table below is just what those words map onto.

### What the assistant fills in

You don't set these by hand; this is here so you know what it can control.

| Argument | Default | Notes |
|---|---|---|
| `file_path` | — | Path on this machine. Absolute, relative, or `~/...`. |
| `audio_url` | — | Public http(s) URL, downloaded then transcribed. |
| `audio_base64` | — | For short clips. Prefer `file_path`. |
| `model` | `gpt-transcribe` | See the table below. |
| `language` | auto-detect | ISO-639-1 (`ar`, `en`). Set only if detection is wrong. |
| `prompt` | — | Names or jargon likely in the audio, to steer spelling. |

Exactly one audio source per request.

### Models

| Model | Cost/min | Speaker labels |
|---|---|---|
| `gpt-transcribe` *(default)* | $0.0045 | no |
| `gpt-4o-mini-transcribe` | $0.003 | no |
| `gpt-4o-transcribe` | $0.006 | no |
| `gpt-4o-transcribe-diarize` | $0.006 | **yes** |
| `whisper-1` | $0.006 | no |

`gpt-transcribe` is OpenAI's recommended model: cheaper than `whisper-1` and
more accurate. There is no reason to pick `whisper-1` unless you need it
specifically.

### Formats

`flac` `m4a` `mp3` `mp4` `mpeg` `mpga` `oga` `ogg` `wav` `webm` — up to 25 MB.

`.opus` files work too. OpenAI rejects that extension even though the bytes are
what it accepts as `.ogg`, so this server renames it in flight. WhatsApp voice
notes are all `.opus`, which is exactly the case that would otherwise fail.

## Why it runs locally

A remote MCP server cannot read a file you attached in chat. There is no
mechanism in the protocol that carries attachments to a server — it has been
[an open issue](https://github.com/modelcontextprotocol/modelcontextprotocol/issues/155)
since early 2025, and the proposals to add one are still drafts. Base64 through
a tool argument works in theory but dies on client argument-size limits after a
second or two of audio.

Running on your machine avoids all of it. The server is a process you own,
reading a file you own, with a key you hold.

## Development

```bash
git clone https://github.com/jack-ai-net/stt-mcp
cd stt-mcp
python3 -m venv .venv && .venv/bin/pip install -e ".[dev]"
.venv/bin/python -m pytest
```

Tests stub the OpenAI call, so they need no API key and cost nothing.

## License

MIT — see [LICENSE](LICENSE).

Built by [JackAI](https://jack-ai.net).
