Voice Agent Evaluation

PASSING 10 / 10 tests
LATEST RUN Just now
System status

Run Summary

ALL EVALUATORS PASSING BASELINE DATA
100% PASS RATE
EVALUATION HEALTHOPTIMAL

Evaluation Pass Rate

100%

10 of 10 evaluation cases passed

Average Score 1.000
Avg Latency ~0.01 ms
Dataset 10 cases
Evaluators Multiple
EVALUATION PIPELINE
  1. USER INPUT
  2. TRANSCRIPTION
  3. INTENT DETECTION
  4. TOOL SELECTION
  5. TOOL EXECUTION
  6. RESPONSE
  7. TRACE
  8. EVALUATORS
  9. METRICS
Test Runs

Agent Execution Pipeline

End-to-end path of the latest run across the agent's execution stages.

TRACE 20:41:02
Session Run History 0 RUNS
RUN ID SCENARIO TIME STATUS SCORE LATENCY
INPUT
User request
OK
TRANSCRIPT
Speech → text
OK
INTENT
order_status
OK
TOOL
get_order
OK
RESULT
shipped
OK
RESPONSE
Generated response
OK
EVALUATION
PASS
PASS

Execution path shown for the order status request; evaluators score the completed trace.

Test Run Explorer

Test Cases

Every case in the current dataset with its evaluation outcome. Select a row to inspect the full case detail.

10 / 10 PASS BASELINE
10 cases
STATUS TEST CASE CATEGORY SCORE LATENCY EVALUATION
Metric Matrix

Evaluation Metrics

Behavioral and response-level evaluation across the agent pipeline.

8 METRICS BASELINE
Evidence Check

Groundedness

Claims in the agent response are verified against what the tool actually returned.

10 / 10 BASELINE

TOOL EVIDENCE

status:shipped
eta:September 28

SUPPORTED CLAIMS

shipped
September 28
order 1234
GROUNDING SCORE 100%

Response claims are checked against information returned by the tool.

AGENT RESPONSE

Your order 1234 is shipped. The estimated delivery date is September 28.

Observability

Agent Trace

Step-by-step spans for the order status request. Select a step to expand its payload.

7 STEPS
"Where is my order 1234?"
"Where is my order 1234?"
intent: order_status
get_order(order_id: 1234)
{ status: shipped, eta: September 28 }
"Your order 1234 is shipped. The estimated delivery date is September 28."
result: PASS
Scenario Lab

Scenario Lab

Create controlled agent situations and evaluate behavior.

API OFFLINE 0 RUNS
Runs execute the real Python agent + Agentic Evals pipeline
Failure Injection

Faults modify real agent behavior; each is detected by a specific evaluator.

No override — each scenario runs its own fault.

Performance

Execution Latency

Local execution time per test case, plotted on a 0 - 0.03 ms axis.

~0.01 ms AVG BASELINE

Latency shown here represents local deterministic execution and is not representative of production STT/LLM/TTS latency.

Failure Handling

Reliability and Safe Recovery

When required information is missing, the agent must ask instead of calling a tool with incomplete arguments.

PASS BASELINE
INPUT
User request

The request arrives without the detail a tool requires.

"Where is my order?"
MISSING INFORMATION
No order number

The order identifier needed by get_order is absent, so no tool is called.

SAFE RECOVERY
No tool invocation

The agent avoids an incomplete or invented tool call.

CLARIFICATION REQUEST
Ask the user

A clarifying question returns the conversation to a resolvable state.

"Could you please provide your order number?"
AGENT RESPONSE

Could you please provide your order number?

EVALUATION PASS RECOVERY
Scope

Limitations

  • STT is simulated
  • TTS is not implemented
  • Dataset is small and synthetic
  • Agent logic is deterministic
  • Tools are local simulations
  • Latency is local execution latency
  • No real network/tool failure simulation

AGENTIC EVALS / VOICE AGENT OBSERVATORY / VOICE SUPPORT AGENT