Metadata-Version: 2.4
Name: voice-agent-56
Version: 0.3.0
Summary: Reusable voice agent sessions with local speech and OpenRouter adapters
License-Expression: Apache-2.0
Project-URL: Repository, https://github.com/Raghav-56/voice-agent
Project-URL: Issues, https://github.com/Raghav-56/voice-agent/issues
Requires-Python: >=3.12
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: numpy>=2.5.3
Provides-Extra: kokoro
Requires-Dist: kokoro-onnx>=0.6.1; extra == "kokoro"
Provides-Extra: piper
Requires-Dist: piper-tts>=1.8.0; extra == "piper"
Provides-Extra: stt
Requires-Dist: faster-whisper>=1.2.1; extra == "stt"
Requires-Dist: sherpa-onnx==1.13.8; extra == "stt"
Requires-Dist: sherpa-onnx-core==1.13.8; extra == "stt"
Provides-Extra: openrouter
Requires-Dist: httpx>=0.28.1; extra == "openrouter"
Requires-Dist: httpx-sse>=0.4.3; extra == "openrouter"
Dynamic: license-file

# voice-agent

Reusable Python voice-agent components with Kokoro and Piper TTS, local STT,
and an asynchronous OpenRouter LLM adapter. A turn-based conversation session connects these adapters.

## Setup

Use uv and Python 3.12 or newer. Add only the providers you need to a consuming
project, for example:

```bash
uv add "voice-agent-56[openrouter,stt]"
```

From this repository checkout, install the TTS providers with:

```bash
uv sync --locked --extra kokoro --extra piper
```

Select one extra to install only that provider. Add dependencies with `uv add`
and commit `pyproject.toml` with `uv.lock`. Run scripts with `uv run`.

## TTS

```python
from voice_agent.tts import create

engine = create("piper:en_US-lessac-medium", models_dir="/path/to/models")
for audio, sample_rate in engine.stream("Hello. How can I help?"):
    print(len(audio), sample_rate)
```

Chunks contain mono float32 NumPy audio at the returned sample rate.
Streaming is sentence-based. The application owns playback and resampling.

Supply model files through `models_dir`; they are not stored in Git:

- Kokoro needs `kokoro-v1.0.onnx` and `voices-v1.0.bin`. Its default voice is `af_heart`.
- Piper needs `<voice>.onnx` and `<voice>.onnx.json`. Its default voice is `en_US-lessac-medium`.

## STT

Install local recognition with `uv sync --locked --extra stt`.
The adapters accept mono float32 samples at 16 kHz and return a `Transcript`.
The application owns decoding, resampling, and recording limits.

```python
from voice_agent.stt import create

engine = create("whisper", models_dir="models/stt", num_threads=2)
result = engine.transcribe(samples, language="en")
print(result.text)
```

Choose `whisper` for multilingual Whisper base or `zipformer` for English.
Whisper also accepts `language=None` for detection. Zipformer accepts only
English. Calls return completed-utterance results. They do not provide a live
microphone API. Whisper includes segment timestamps; Zipformer currently returns
text without segment timestamps. Callers must serialize access to each engine.

Model directories under `models_dir` are `whisper-base` and `zipformer-en-full`.
The checked-out model assets are in `models/stt/`.
See [STT research and local measurements](https://github.com/Raghav-56/voice-agent/blob/main/docs/stt-options.md) for the exact
models, versions, limitations, and benchmark command. Model files stay outside Git.

## LLM through OpenRouter

In this repository checkout, install the provider and configure your key:

```bash
uv sync --locked --extra openrouter
cp .env.example .env
# Set OPENROUTER_API_KEY in .env, then run:
uv run --env-file .env --extra openrouter python scripts/chat_llm.py
```

The default model is `stealth/space-bunny-alpha`. This is the model described in
[the supplied AI/ML API documentation](https://docs.aimlapi.com/api-references/text-models-llm/stealth/space-bunny-alpha),
and [OpenRouter lists the same model ID](https://openrouter.ai/stealth/space-bunny-alpha).
Requests go directly to `https://openrouter.ai/api/v1/chat/completions` and require
an OpenRouter key, not an AI/ML API key. Set `OPENROUTER_MODEL` to change models,
or pass `model=` explicitly. The library reads environment variables; the
`uv run --env-file` option loads the example configuration for the script.

```python
import asyncio
from contextlib import aclosing
from voice_agent.llm import create

async def main():
    messages = [
        {"role": "system", "content": "Reply briefly in natural spoken language. Avoid markdown."},
        {"role": "user", "content": "Hello. Can you help me?"},
    ]
    async with create("openrouter") as engine:
        async with aclosing(engine.stream(messages)) as response:
            async for text in response:
                print(text, end="", flush=True)
        # Or await engine.complete(messages) for one complete string.

asyncio.run(main())
```

`stream()` yields text deltas, which may be fragments of words. Buffer them into
sentences before passing them to TTS. The application owns the system prompt,
conversation history, sentence buffering, and playback. This adapter supports
text messages with system, user, and assistant roles; tool execution and full
audio conversation orchestration are not implemented.

Provider options are `api_key`, `model`, `max_tokens` with a default of 2048,
`reasoning_effort` with a default of `low`, and `timeout` with a default of 60
seconds of network inactivity. The token budget also covers reasoning. Use a
larger budget if the provider hits its limit. Set `reasoning_effort=None` to omit
the reasoning option for another model. Only answer content is yielded.

Use `aclosing` around a stream if you might stop early. Cancelling the task or
closing the iterator closes its HTTP response. Use the engine as an async
context manager, or call `await engine.aclose()` when finished. Requests are not
retried automatically. `LLMError` reports HTTP failures, timeouts, stream errors,
truncation, and empty answers. Partial text may already have been yielded when
an error occurs; the application decides what to play or retain in history.

Run the provider tests without credentials or model downloads:

```bash
uv run --locked --extra openrouter python -m unittest discover -s tests -v
```

## Conversation sessions

`voice_agent.session.ConversationSession` provides turn-based orchestration and
isolated in-memory history. Supply a frozen `SessionConfig`, an asynchronous LLM
adapter, and async transcription and synthesis callbacks. The caller owns file
decoding, native-model concurrency, transport, deadlines, and session expiry.

For an application-supplied greeting, call `begin()` and consume
`speak(text, generation)` before the first caller turn. Like `turn()`, it emits
audio and completion events and requires a playback acknowledgement. It does not
add a fabricated caller message to history.

Call `begin()` before consuming `turn(audio, generation)`. It emits stage,
transcript, text delta, WAV audio, and completion events. Call `acknowledge` with
the generation only after playback completes. Cancel invalidates old generations
and closes active LLM work. It does not stop synchronous native inference that
is already running. Interrupted replies are marked in history.

`voice_agent.endpoint.EndpointDetector` accepts 20 ms mono float32 frames at
16 kHz and emits speech-start and completed-utterance events. It uses an RMS
threshold with configurable start, silence, pre-roll, and maximum-duration
limits. It does not distinguish caller speech from echo or background voices;
test and tune it against real call recordings before enabling barge-in.

This session implementation uses completed English utterances and waits
for a full generated reply before synthesis. It does not implement streaming
microphone input, automatic barge-in, or a remote worker protocol.

An application may pass typed tool schemas and an `execute_tool` callback to
`ConversationSession`. The OpenRouter adapter uses a non-streaming completion
when tools are available. The session emits `tool_request`, awaits the callback,
and emits `tool_result`. It ends that turn without synthesizing model text when
a tool is requested. The application must authorize the request and provide any
speech the caller should hear before executing it. Only one tool request is
supported per turn. The package does not perform campaign or call-control actions.
The live integration is in the sibling `voice-lab` application.

## Design

See [architecture](https://github.com/Raghav-56/voice-agent/blob/main/docs/architecture.md),
[provider research](https://github.com/Raghav-56/voice-agent/blob/main/docs/voice-agent-landscape.md),
and [NVIDIA model notes](https://github.com/Raghav-56/voice-agent/blob/main/docs/voice-agent-options.md).
Applications supply transport, context, and domain tools. Keep application setup
in the consuming repository.

The package is licensed under [Apache-2.0](https://github.com/Raghav-56/voice-agent/blob/main/LICENSE).
