Metadata-Version: 2.5
Name: voice-agent-kit
Version: 0.1.0
Summary: An open-source, provider-agnostic realtime voice agent framework in Python
Project-URL: Homepage, https://github.com/sandipan-ai95/voice-agent-kit
Project-URL: Documentation, https://github.com/sandipan-ai95/voice-agent-kit/tree/main/docs
Project-URL: Repository, https://github.com/sandipan-ai95/voice-agent-kit
Project-URL: Issues, https://github.com/sandipan-ai95/voice-agent-kit/issues
Author-email: Sandidas <sandidas@example.com>
License: Apache-2.0
License-File: LICENSE
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: Apache Software License
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Multimedia :: Sound/Audio :: Speech
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Python: >=3.10
Requires-Dist: aiohttp>=3.9.0
Requires-Dist: pydantic>=2.0.0
Requires-Dist: pyyaml>=6.0.0
Provides-Extra: all
Requires-Dist: openai>=1.30.0; extra == 'all'
Requires-Dist: websockets>=13.0; extra == 'all'
Provides-Extra: deepgram
Requires-Dist: websockets>=13.0; extra == 'deepgram'
Provides-Extra: dev
Requires-Dist: build>=1.1.0; extra == 'dev'
Requires-Dist: mypy>=1.9.0; extra == 'dev'
Requires-Dist: pytest-asyncio>=0.23.0; extra == 'dev'
Requires-Dist: pytest-cov>=4.1.0; extra == 'dev'
Requires-Dist: pytest>=8.0.0; extra == 'dev'
Requires-Dist: ruff>=0.3.0; extra == 'dev'
Requires-Dist: twine>=5.0.0; extra == 'dev'
Provides-Extra: elevenlabs
Requires-Dist: websockets>=13.0; extra == 'elevenlabs'
Provides-Extra: local-stt
Requires-Dist: faster-whisper>=1.0.0; extra == 'local-stt'
Provides-Extra: openai
Requires-Dist: openai>=1.30.0; extra == 'openai'
Provides-Extra: websocket
Requires-Dist: websockets>=13.0; extra == 'websocket'
Description-Content-Type: text/markdown

# voice-agent-kit

[![CI](https://github.com/sandidas/voice-agent-kit/actions/workflows/ci.yml/badge.svg)](https://github.com/sandidas/voice-agent-kit/actions)
[![License: Apache-2.0](https://img.shields.io/badge/License-Apache%202.0-blue.svg)](https://opensource.org/licenses/Apache-2.0)
[![Python 3.10+](https://img.shields.io/badge/python-3.10+-blue.svg)](https://www.python.org/downloads/)

An open-source, provider-agnostic, low-latency realtime voice agent framework in Python.

Build conversational voice AI for **customer support**, **telephony**, **browser voice apps**, and **self-hosted local enterprise clusters** without vendor lock-in or cloud intermediary proxies.

---

## Key Highlights

- **Provider-agnostic:** swap STT, LLM and TTS providers via config or `register_provider` without changing orchestration code.
- **OpenAI-compatible LLMs:** one SSE-streaming adapter for Ollama, vLLM, LM Studio, OpenRouter, OpenAI and gateways.
- **Runs fully local (no API key):** faster-whisper STT (optional extra) → Ollama → macOS `say` TTS (macOS only). No portable local TTS yet.
- **OS-agnostic core:** pure-Python `py3-none-any` wheel with no mandatory OS, audio-hardware or GPU dependencies; platform-specific code lives only in optional providers ([docs/PLATFORMS.md](docs/PLATFORMS.md)).
- **No project server:** audio and credentials go directly to the endpoints you configure.
- **Barge-in:** user speech cancels the in-flight LLM/TTS response and drains queued output (regression-tested).
- **Tool calling:** `@tool` decorator with JSON-schema extraction, argument validation and deadlines.
- **Evaluation framework:** `voice-agent eval` runs YAML scenarios in strictly separated `synthetic` / `local` / `remote` modes and reports WER/CER, latencies, reliability, barge-in and reply script; never ranks providers.
- **Deterministic offline tests:** fake STT/LLM/TTS/VAD for GPU-free, key-free CI.
- **Languages:** BCP-47 language metadata and per-provider language declarations; English/Hindi/Bengali scenarios included. **No language is verified for a real model yet** (see [docs/LANGUAGES.md](docs/LANGUAGES.md)).

Not implemented: telephony transports (only interface stubs in `voice_agent.telephony`), microphone I/O, ML-based VAD.

---

## Architecture Overview

```mermaid
flowchart TD
    subgraph Transports ["Transport Layer"]
        WS["WebSocketServerTransport (one shared conversation)"]
        WSS["WebSocketSessionServer (one agent per connection)"]
        MEM[In-memory / custom Transport]
    end

    subgraph CoreEngine ["VoiceAgent Orchestrator"]
        VAD[Energy VAD]
        STT[STT Adapter]
        BUS[Async Event Bus]
        LLM[Streaming LLM / OpenAI-Compatible]
        TOOLS[Tool Engine]
        TTS[TTS Adapter]
        DRAIN[Barge-In Cancellation]
    end

    Transports <-->|Audio Chunks| VAD
    VAD -->|Voice Activity| STT
    STT -->|Transcripts| BUS
    BUS <--> LLM
    LLM <--> TOOLS
    LLM -->|Tokens| TTS
    TTS -->|PCM Audio Chunks| Transports
    VAD -.->|Interruption Signal| DRAIN
    DRAIN -.->|Cancel Task & Flush Queue| Transports
```

---

## 30-Second Quickstart

### 1. Installation

```bash
# Core package (ultra-lightweight, < 25MB)
pip install voice-agent-kit

# With OpenAI support
pip install "voice-agent-kit[openai]"

# With the WebSocket server transports
pip install "voice-agent-kit[websocket]"

# With all adapters
pip install "voice-agent-kit[all]"
```

### 2. Basic Agent (Pure Python)

```python
import asyncio
import os
from voice_agent import VoiceAgent, tool
from voice_agent.providers import OpenAICompatibleLLM, DeepgramSTT, ElevenLabsTTS
from voice_agent.transports import WebSocketSecurityConfig, WebSocketServerTransport


@tool(name="check_order", description="Look up shipping status of an order")
def check_order(order_id: str) -> dict:
    return {"order_id": order_id, "status": "Out for delivery"}


agent = VoiceAgent(
    stt=DeepgramSTT(api_key="..."),
    llm=OpenAICompatibleLLM(
        base_url="http://localhost:8000/v1",  # Local vLLM or Ollama instance
        model="Qwen/Qwen2.5-7B-Instruct",
        api_key="EMPTY",
    ),
    tts=ElevenLabsTTS(api_key="..."),
    system_prompt="You are an AI customer concierge. Keep answers concise.",
    # Set VOICE_AGENT_WS_TOKEN before starting. Bind to loopback; use a TLS proxy for remote clients.
    transport=WebSocketServerTransport(
        host="127.0.0.1",
        port=8765,
        security=WebSocketSecurityConfig(auth_token=os.environ["VOICE_AGENT_WS_TOKEN"]),
    ),
)

agent.register_tool(check_order)

if __name__ == "__main__":

    async def main() -> None:
        await agent.start()
        try:
            await asyncio.Event().wait()
        finally:
            await agent.stop()

    asyncio.run(main())
```

### 3. One conversation vs. many callers

`WebSocketServerTransport` belongs to **one** `VoiceAgent` and therefore **one shared conversation**.
With the default `max_connections=1` it serves a single caller. `max_connections` on this transport
controls how many sockets may join **that same conversation**; it does **not** create independent
callers. With `max_connections > 1`, every socket shares conversation history, agent state, tools,
provider state, and the broadcast audio output (a `UserWarning` is emitted for this configuration).
Use it only when several sockets really are the same session.

For **independent concurrent callers**, use `WebSocketSessionServer`, which builds a new `VoiceAgent`
per connection. Its `max_connections` bounds the number of concurrent isolated sessions:

```python
from voice_agent.transports import WebSocketConnectionTransport, WebSocketSessionServer


def make_agent(transport: WebSocketConnectionTransport) -> VoiceAgent:
    return VoiceAgent(stt=..., llm=..., tts=..., transport=transport)  # fresh providers/state per caller


server = WebSocketSessionServer(make_agent, security=WebSocketSecurityConfig(max_connections=8))
await server.start()
```

| | `WebSocketServerTransport` | `WebSocketSessionServer` |
| --- | --- | --- |
| Agents | one, shared by all sockets | one per connection (built by your factory) |
| Transport | one, output broadcast to every socket | one `WebSocketConnectionTransport` per connection |
| `max_connections > 1` means | extra sockets in the same shared conversation | independent, isolated callers |
| History / tools / audio | shared | isolated per caller |
| Session reset | when the first socket joins an idle transport and when the last one leaves | each session starts fresh and is stopped when its connection ends |

`WebSocketSessionServer` also rejects (close code 1011) a session whose agent uses an `EventBus`, VAD,
memory, pipeline, STT/LLM/TTS provider, or tool still owned by another active session, bounds factory +
`agent.start()` by `security.session_setup_timeout_s`, cancels setup if the caller disconnects or the
server stops, and cancels any `agent.stop()` exceeding `security.session_shutdown_timeout_s`. Build
providers and tools inside the factory; share one only if it declares `session_shareable = True`
(`@tool(session_shareable=True)` for tools), i.e. keeps no per-caller state. Ownership also covers mutable
objects reachable through wrappers (adapters, partials, closures, bound methods), so a fresh wrapper per
session around one shared stateful backend is rejected too; see [docs/API_DESIGN.md](docs/API_DESIGN.md).

`voice-agent run --ws` uses `WebSocketServerTransport`. See
[examples/websocket_server.py](examples/websocket_server.py) for both server types (`--sessions` for the per-caller server).

---

## Command Line Interface (CLI)

The package includes a comprehensive diagnostics and testing CLI:

```bash
# 1. Inspect Python runtime, installed codecs, and credential status
voice-agent doctor

# 2. Run deterministic turn-taking and barge-in test scenarios
voice-agent test --scenario barge-in

# 3. Benchmark latency: SYNTHETIC (fake providers, framework overhead only) or REAL (live providers)
voice-agent benchmark --trials 5
voice-agent benchmark --mode real --config prod.yaml --audio speech_16k.wav --network "describe it"

# 4. Run interactive local conversation demo
voice-agent demo

# 5. Launch agent declaratively from a YAML file (optionally over authenticated WebSocket)
voice-agent run config.yaml
VOICE_AGENT_WS_TOKEN=... voice-agent run config.yaml --ws
```

---

## Declarative YAML Configuration

```yaml
version: "1.0"
agent:
  name: "support-agent"
  system_prompt: "You are a customer service voice assistant."

stt:
  provider: "deepgram"
  model: "nova-2"

llm:
  provider: "openai-compatible"
  base_url: "http://localhost:8000/v1"
  model: "Qwen/Qwen2.5-7B-Instruct"

tts:
  provider: "elevenlabs"
  voice_id: "21m00Tcm4TlvDq8ikWAM"

realtime:
  allow_interruptions: true
  sample_rate: 16000
```

---

## Development & Testing

```bash
# Clone the repository
git clone https://github.com/sandidas/voice-agent-kit.git
cd voice-agent-kit

# Create virtual environment
python3 -m venv .venv
source .venv/bin/activate

# Install in editable mode with development dependencies
pip install -e ".[dev,all]"

# Run offline unit and e2e test suite
pytest tests/unit tests/e2e
```

---

## Documentation Index

- [Ecosystem Research & Analysis](docs/RESEARCH.md)
- [Architecture & ADRs](docs/ARCHITECTURE.md)
- [API Design Specification](docs/API_DESIGN.md)
- [Provider System Specification](docs/PROVIDER_SYSTEM.md)
- [Realtime Audio Pipeline](docs/REALTIME_PIPELINE.md)
- [Testing Strategy](docs/TESTING.md)
- [Security & Telemetry Policy](docs/SECURITY.md)
- [Deployment Models](docs/DEPLOYMENT.md)
- [Benchmarking Methodology](docs/BENCHMARKING.md)
- [Project Roadmap](docs/ROADMAP.md)
- [Developer & Agent Guidelines](AGENTS.md)
- [Contributing](docs/CONTRIBUTING.md)
- [Evaluation](docs/EVALUATION.md) · [Languages](docs/LANGUAGES.md) · [Local models](docs/LOCAL_MODELS.md)
- [Adding a provider](docs/ADDING_A_PROVIDER.md) · [Adding a language](docs/ADDING_A_LANGUAGE.md)
- [Platform support](docs/PLATFORMS.md)

---

## License

This project is licensed under the **Apache License 2.0**. See the [LICENSE](LICENSE) file for details.

---

## Project Status (v0.1, pre-release)

This is **not production-ready**. See [`docs/FINAL_AUDIT.md`](docs/FINAL_AUDIT.md) for which
items are verified and which are not. Live cloud adapters (Deepgram, ElevenLabs, OpenAI) have
**not** been exercised against real APIs by the maintainers. A fully local pipeline
(faster-whisper `base` + Ollama llama3 + macOS `say`) has been run end to end; results, including
unusable Hindi/Bengali recognition with Whisper `base`, are in [`docs/EVALUATION.md`](docs/EVALUATION.md).

