Metadata-Version: 2.5
Name: dialt-sdk
Version: 0.15.0
Summary: Python SDK for the Dialt realtime voice and text API.
Project-URL: Documentation, https://dialt.com/docs/api/
Project-URL: Homepage, https://dialt.com/
Author: Dialt
License-Expression: Apache-2.0
License-File: LICENSE
License-File: NOTICE
Keywords: dialt,realtime,speech,voice,voice-ai
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Typing :: Typed
Requires-Python: >=3.11
Requires-Dist: numpy>=1.26
Requires-Dist: websockets>=12.0
Provides-Extra: webrtc
Requires-Dist: aiortc>=1.15.0; extra == 'webrtc'
Description-Content-Type: text/markdown

# dialt-sdk

The headless Python SDK for the
[Dialt realtime voice API](https://dialt.com/docs/api/). Dialt manages
speech recognition, turn-taking, interruption handling, speech generation, reasoning and tool
orchestration through a single API; this SDK connects telephony bridges, services, evaluations
and custom devices to it.

`dialt-sdk` replaces the deprecated `converse-sdk` distribution. New code imports from `dialt`.
Existing applications can install the final `converse-sdk` compatibility release while migrating imports from `converse_sdk` to `dialt`.

```sh
uv add dialt-sdk
```

The session-loop fragment below assumes your media layer supplies `mic_frame` and `play_audio`,
and your application supplies `run_tool`:

```python
import os

from dialt import ConverseMode, ConverseSession, ToolDefinition

lookup_tool: ToolDefinition = {
    "name": "lookup_order",
    "description": "Look up an order by ID.",
    "parameters": {"type": "object", "properties": {"order_id": {"type": "string"}}},
    "read_only": True,
    "expected_duration": "instant",
    "status_label": "order lookup",
    "deferred": True,
    "deferred_timeout": 7200,
    "notify_on_complete": True,
}

session = await ConverseSession.connect(
    "wss://dialt.com/ws",
    api_key=os.environ["CONVERSE_API_KEY"],
    mode=ConverseMode(
        instructions="Help callers with their orders.",
        greeting="Hello, how can I help?",
        tools=[lookup_tool],
    ),
)

async with session:
    await session.send_audio(mic_frame)  # PCM16 little-endian mono at 16 kHz
    await session.inject_context(
        "Claude Code finished. Tell the user briefly.", role="context", reply=True)

    async for event in session.events():
        if event.type == "audio":
            play_audio(event.audio)  # Float32 mono at 16 kHz
        elif event.type == "tool_call":
            result = await run_tool(event.data["name"], event.data["args"])
            await session.send_tool_result(
                event.data["id"], result, outcome="succeeded", verified=True)
        elif event.type == "tool_deferred_resume":
            resume_host_job(event.data["handle"])
```

`deferred: True` makes the tool a background job: the call can outlive the voice turn, the
conversation carries on while the host works, and with `notify_on_complete` the agent speaks the
result when it lands — even if the caller has moved on to another topic. For work that should
release the voice turn immediately, call `send_tool_deferred(id, handle, status_label=...)` before
continuing it in the background. The handle identifies that call, not your worker: mint a fresh one
per call (e.g. `f"job-{call_id}"`) and check the returned acknowledgement — re-using a live handle
is rejected as `handle_in_use`, and a rejected defer leaves the call on the ordinary tool timeout,
so it expires while you think it is backgrounded. Keep a handle-to-worker map if several calls
should feed one long-running worker.
For proactive host announcements, call
`await session.inject_context(text, role="context", reply=True, message_id="job-42")`. It returns
the broker's authoritative acknowledgement (`accepted` plus optional `retryable`/`detail`). For a
typed user turn, the same `message_id` and `input_source="text"` appear on the final ASR event;
spoken input uses `input_source="voice"`. The role defaults to `"context"`, `reply` defaults to
`False`, and omitted message IDs are generated by the SDK. Text is limited to 1–2000 characters.
An accepted `message_id` is an idempotency key for the logical session. After an acknowledgement
loss, retry the identical payload and ID (including after resume) to replay delivery proof without
injecting a duplicate turn. Reusing an accepted ID with different content is rejected; the broker
retains up to 512 accepted receipts per logical session.

Save `session.resume_token` from the connected session; after a transport loss, pass it as
`resume_token=...` to the replacement `ConverseSession.connect(...)` call so pending jobs are
re-announced as `tool_deferred_resume` events within the broker's bounded resume window.

The SDK deliberately does not own capture, playback, pacing or echo cancellation. Live playback
integrations must implement the
[playback contract](https://dialt.com/docs/api/websocket/#playback-contract).

For a text session, keep the same conversation configuration and replace the media loop with
committed turns:

```python
async with await ConverseSession.connect(
    "wss://dialt.com/ws",
    api_key=os.environ["CONVERSE_API_KEY"],
    mode=ConverseMode(modality="text", instructions="Help customers with their orders."),
) as session:
    await session.send_text("Where is order A123?")
    async for event in session.events():
        if event.type == "utterance":
            print(event.data["text"])
```

Text mode is WebSocket-only and emits no audio. Model behavior, tools, greeting, history and
conversation lifecycle events remain the same.

`ConverseMode(end_call=True)` lets the agent end the session itself through the managed
`end_call(farewell)` tool: the farewell is spoken, `session_end_requested` arrives with the
`farewell`, and the server closes after a short grace unless the user speaks. Off by default; the
host then ends the session with `wrap_up` or by closing.

See the [complete Python reference](https://dialt.com/docs/api/python/#python-sdk)
and the [Twilio bridge quickstart](https://dialt.com/docs/api/twilio/).

### Hosted evals

`dialt.evals` creates cases and starts runs on the Dialt evals dashboard with your
account API key (`ck_...`). Cases are JSON files in your repository; `upsert_case` keeps the
hosted copy in step by name, so a re-push updates rather than duplicates.

```python
from dialt import EvalsClient, load_cases

evals = EvalsClient(api_key=os.environ["CONVERSE_API_KEY"])
cases = evals.upsert_cases(load_cases("evals/"))          # one case per *.json file
run = evals.start_run([c["id"] for c in cases], modality="text")
print(evals.dashboard_url(run["id"]))
result = evals.wait(run["id"])                            # polls until terminal
print(result["status"], [a["status"] for a in result["attempts"]])
```

A case has `name`, `starter`, `target` (`instructions`, optional `tools`), `simulator`
(`instructions`), `fixtures`, `checks` and `limits`; the field reference is in the evals guide.
`converse-recipes` ships a `converse-evals push evals/` command built on this client.

### Relaying two sessions (simulations)

`dialt.relay` cross-pipes two sessions so one can play the user for the other: the
building blocks behind `converse-recipes`' `converse-sim` and the hosted evals in the Dialt
webapp.

```python
from dialt.relay import TextTurnRelay, VoiceTurnRelay

# text: forward each committed target utterance to the simulator once tool work has settled
relay = TextTurnRelay(simulator.send_text)
relay.utterance(event.data["text"])      # on the target's `utterance`
relay.working(active)                    # on `working`
relay.done()                             # on `done`: forwards after TURN_RELAY_SETTLE_S

# voice: stream the target's audio straight into the simulator's mic, then close the turn
relay = VoiceTurnRelay(simulator)
await relay.audio(event.audio)           # on each `audio` event
relay.input_committed()                  # on the simulator's `asr`: stops the silence tail early
```

`VoiceTurnRelay.done()` feeds up to `voice_tail_s()` of silence so the receiving side's
endpointer commits the turn (bounded by `CARTESIA_TURN_END_TIMEOUT_MS` when set, else Ink's
5.6 s default). Both relays take `on_error` to surface background failures.

### WebRTC transport (experimental)

Experimental: the API is stable, but this transport is newly shipped and still being hardened on
real networks; `ws` remains the default and recommended fallback.

Pass `transport="webrtc"` to `ConverseSession.connect(...)` to carry the session over WebRTC (UDP)
instead of the default WebSocket. Requires `pip install "dialt-sdk[webrtc]"` (or `uv add`).
`ws` remains the default; most headless callers are fine on it. See the
[Python guide's WebRTC section](https://dialt.com/docs/api/python/#python-webrtc).

Licensed under the [Apache License 2.0](LICENSE). This license applies to the SDK, not to the
hosted Dialt service, its models, or its server-side implementation. Runtime dependencies
remain under their own licenses.
