Run Summary
Evaluation Pass Rate
100%
10 of 10 evaluation cases passed
- USER INPUT
- TRANSCRIPTION
- INTENT DETECTION
- TOOL SELECTION
- TOOL EXECUTION
- RESPONSE
- TRACE
- EVALUATORS
- METRICS
Agent Execution Pipeline
End-to-end path of the latest run across the agent's execution stages.
| RUN ID | SCENARIO | TIME | STATUS | SCORE | LATENCY |
|---|
Execution path shown for the order status request; evaluators score the completed trace.
Test Cases
Every case in the current dataset with its evaluation outcome. Select a row to inspect the full case detail.
| STATUS | TEST CASE | CATEGORY | SCORE | LATENCY | EVALUATION |
|---|
Evaluation Metrics
Behavioral and response-level evaluation across the agent pipeline.
Groundedness
Claims in the agent response are verified against what the tool actually returned.
TOOL EVIDENCE
SUPPORTED CLAIMS
Response claims are checked against information returned by the tool.
AGENT RESPONSE
Your order 1234 is shipped. The estimated delivery date is September 28.
Agent Trace
Step-by-step spans for the order status request. Select a step to expand its payload.
Scenario Lab
Create controlled agent situations and evaluate behavior.
Faults modify real agent behavior; each is detected by a specific evaluator.
No override — each scenario runs its own fault.
Scenario
Execution Latency
Local execution time per test case, plotted on a 0 - 0.03 ms axis.
Latency shown here represents local deterministic execution and is not representative of production STT/LLM/TTS latency.
Reliability and Safe Recovery
When required information is missing, the agent must ask instead of calling a tool with incomplete arguments.
The request arrives without the detail a tool requires.
"Where is my order?"The order identifier needed by get_order is absent, so no tool is called.
The agent avoids an incomplete or invented tool call.
A clarifying question returns the conversation to a resolvable state.
"Could you please provide your order number?"Could you please provide your order number?
Limitations
- STT is simulated
- TTS is not implemented
- Dataset is small and synthetic
- Agent logic is deterministic
- Tools are local simulations
- Latency is local execution latency
- No real network/tool failure simulation