Agentic Evals Documentation
A standalone, framework-agnostic evaluation and scoring engine for LLM and agent outputs.
What is agentic-evals?
agentic-evals is a Python package for evaluating the quality of LLM and agent outputs programmatically. It runs standalone and is not tied to any agent framework or tracing format.
You provide:
- Test cases — expectations (exact match, substring, JSON Schema, tool calls, latency, etc.)
- Evaluation samples — actual outputs and execution traces
- Custom evaluators (optional) — your own scoring logic or LLM judges
agentic-evals returns:
- Per-case scores — individual judgments with explanations
- Summary metrics — pass rate, average score, cost, latency
- Pass/fail gates — CI-ready exit codes for release decisions
Installation
pip install agentic-evals
Optional: for voice agent evaluation examples with FastAPI:
pip install "agentic-evals[voice]"
Your first eval in 60 seconds
The simplest way to write an eval:
from agentic_evals import Eval, equals
Eval(
"my_first_eval",
data=[
{"input": "France", "expected": "Paris"},
{"input": "Japan", "expected": "Tokyo"},
],
task=lambda x: my_agent(x["input"]),
scores=[equals],
)
Save as my_eval.py and run:
python my_eval.py
Or discover and run all eval files:
agentic-evals run
Core Concepts
EvalTrace and EvalSpan
The minimal trace structure agentic-evals inspects — deliberately not tied to any specific tracing tool or framework.
class EvalTrace:
trace_id: str
spans: list[EvalSpan]
total_latency_ms: float
estimated_cost_usd: float | None
metadata: dict[str, Any]
class EvalSpan:
tool_name: str | None
attributes: dict[str, Any]
You build these from whatever trace data you have (AgenticLens, OpenTelemetry, custom instrumentation, etc.):
from agentic_evals import EvalTrace, EvalSpan
trace = EvalTrace(
trace_id="run-123",
spans=[
EvalSpan(tool_name="calculator", attributes={"tool_args": {"expr": "2+2"}}),
EvalSpan(tool_name="search", attributes={}),
],
total_latency_ms=250,
estimated_cost_usd=0.015,
metadata={"turn_count": 2},
)
TestSuite and TestCase
Define expectations for your agent outputs:
| Field |
Type |
Purpose |
| expected_output |
str |
Exact match (whitespace-trimmed) |
| expected_contains |
list[str] |
Required substrings (case-insensitive) |
| output_json_schema |
dict |
JSON Schema Draft 2020-12 validation |
| required_output_fields |
list[str] |
Dotted paths in JSON (e.g., "data.total") |
| required_tools |
list[str] |
Tool calls that must happen |
| forbidden_tools |
list[str] |
Tool calls that must NOT happen |
| required_tool_arguments |
dict[str, list[str]] |
Specific arguments a tool must receive |
| max_latency_ms |
float |
Latency threshold in milliseconds |
| max_cost_usd |
float |
Cost threshold in USD |
| max_turns |
int |
Maximum conversation turns (from trace metadata) |
| evaluators |
list[EvaluatorConfig] |
Custom evaluators to run |
from agentic_evals import TestSuite, TestCase
suite = TestSuite(
name="support-agent",
version="1.0",
cases=[
TestCase(
id="refund-status",
name="Check refund status",
expected_contains=["Your refund", "processing"],
required_tools=["lookup_order"],
max_latency_ms=1000,
),
],
)
Evaluators
An evaluator is anything with a .name property and an .evaluate(context) -> list[Score] method.
Built-in types:
CallableEvaluator — Wrap any Python function
LLMJudgeEvaluator — Provider-neutral LLM judge
BusinessRuleEvaluator — Named convenience for custom logic
Scores and Reports
Each evaluation produces a Score:
class Score(BaseModel):
name: str # e.g., "factuality", "exact_match"
value: float # 0.0 to 1.0
passed: bool # True if value >= threshold
required: bool = True
explanation: str # Why it passed/failed
evaluator_type: str = "deterministic" # or "llm_judge"
metadata: dict[str, Any] = {}
Multiple scores combine into a CaseEvaluation (one per test case), and those roll up into an EvaluationReport:
class EvaluationReport(BaseModel):
suite_name: str
suite_version: str
created_at: datetime
summary: EvaluationSummary # Pass rate, avg score, cost, latency
cases: list[CaseEvaluation] # Per-case details
APIs
Eval() Quickstart
The fastest, most opinionated API. Perfect for simple cases.
def Eval(
name: str,
data: list[dict] | Callable[[], list[dict]],
task: Callable[[Any], str],
scores: list[Callable],
threshold: float = 1.0,
) -> EvalResult
Parameters:
| Parameter |
Type |
Description |
| name |
str |
Eval name (appears in output) |
| data |
list[dict] or callable |
Test data with "input"/"expected" keys, or callable that returns it |
| task |
callable |
Function to evaluate, takes input dict, returns string |
| scores |
list[callable] |
Scorer functions; each returns 0-1 or bool |
| threshold |
float |
Min score for a case to pass (default 1.0) |
Built-in scorers for Eval():
equals — Exact match
contains — Case-sensitive substring
icontains — Case-insensitive substring
levenshtein — Similarity (0-1)
matches — Regex match
numeric_close — 1% relative tolerance
TestSuite API
More structured, trace-aware evaluation with custom expectations.
def evaluate_suite(
suite: TestSuite,
samples: list[EvaluationSample],
registry: EvaluatorRegistry | None = None,
) -> EvaluationReport
Returns a full EvaluationReport with per-case scores, summary metrics, pass rates, cost/latency.
Scorers Library
Pre-built evaluators organized by type:
Text Scorers (No LLM)
from agentic_evals import scorers
# Deterministic checks
scorer = scorers.text.exact_match
scorer = scorers.text.levenshtein_similarity
scorer = scorers.text.valid_json
scorer = scorers.text.json_diff
scorer = scorers.text.regex_match
Rubric Scorers (LLM-Graded)
from agentic_evals import LLMRubricEvaluator, FACTUALITY, default_registry
# Register an LLM rubric evaluator
registry = default_registry()
registry.register(
LLMRubricEvaluator(
"factuality_check",
FACTUALITY,
complete_fn=your_model_call, # Your function to call LLM
)
)
Trajectory Scorers (Trace-Aware)
from agentic_evals import scorers
scorer = scorers.trajectory.tool_call_precision
scorer = scorers.trajectory.tool_call_recall
scorer = scorers.trajectory.no_redundant_tool_calls
scorer = scorers.trajectory.trajectory_efficiency
Release Gates
Turn an evaluation report into a pass/fail CI decision:
def evaluate_gate(
report: EvaluationReport,
config: GateConfig,
) -> GateDecision
from agentic_evals import evaluate_gate, GateConfig
decision = evaluate_gate(
report,
GateConfig(
min_pass_rate=0.95, # At least 95% of cases pass
max_average_latency_ms=1500, # Avg latency under 1.5s
max_total_cost_usd=0.25, # Total cost under $0.25
)
)
if not decision.passed:
print(f"Release blocked: {decision.reasons}")
exit(1)
Live Targets
Run a suite against a real system (trusted code or HTTP endpoint) instead of pre-recorded samples:
def run_live_suite(
suite: TestSuite,
target: LiveTarget,
registry: EvaluatorRegistry | None = None,
) -> EvaluationReport
Python Target (executes local code):
from agentic_evals import PythonTarget, run_live_suite
report = run_live_suite(
suite,
PythonTarget(callable_path="my_module:my_agent"),
# my_module.py must have a my_agent(input, case=None) function
)
HTTP Target (calls a remote endpoint):
from agentic_evals import HTTPTarget, run_live_suite
report = run_live_suite(
suite,
HTTPTarget(
url="http://localhost:8000/evaluate",
method="POST",
headers={"Authorization": "Bearer token"},
),
)
Advanced Topics
Custom Evaluators
Write your own evaluation logic with full access to the test case, output, and trace.
from agentic_evals import (
LLMJudgeEvaluator, CallableEvaluator,
EvaluationContext, Score, EvaluatorRegistry
)
# LLM-based judge
def my_judge(context: EvaluationContext) -> Score:
output = context.sample.output
expected = context.case.expected_output
# Call your LLM here
judgment = your_llm_model(
f"Is this output correct? {output} vs {expected}"
)
return Score(
name="llm_judgment",
value=1.0 if judgment else 0.0,
passed=judgment,
explanation="LLM agrees" if judgment else "LLM disagrees",
)
registry = EvaluatorRegistry()
registry.register(LLMJudgeEvaluator("my_judge", my_judge))
Then use it in your test case:
case = TestCase(
id="case-1",
name="Custom judge",
evaluators=[EvaluatorConfig(name="my_judge", threshold=0.8)],
)
Eval Packs
Bundle a TestSuite with the scorers it needs into one shareable YAML/JSON file:
name: my-eval-pack
version: 1.0
required_scorers:
- levenshtein_similarity
- factuality
cases:
- id: case-1
name: Check if output is factual
expected_output: Paris
evaluators:
- name: factuality
threshold: 0.9
Load and use it:
from agentic_evals import load_pack, evaluate_suite, default_registry
pack = load_pack("path/to/pack.yaml")
report = evaluate_suite(pack.to_suite(), samples, registry=default_registry())
Built-in packs:
tool-use-correctness — Evaluate tool selection and arguments
json-output-contract — Validate JSON structure
factual-qa — QA with factuality rubric
from agentic_evals import load_builtin_pack
pack = load_builtin_pack("tool-use-correctness")
Evaluator Registry
Manage evaluators centrally so test cases can reference them by name:
from agentic_evals import (
EvaluatorRegistry, default_registry,
LLMJudgeEvaluator, CallableEvaluator
)
# Start with defaults
registry = default_registry()
# Add custom evaluators
def my_rule(context):
return Score(...)
registry.register(
CallableEvaluator("my_rule", my_rule)
)
# Get an evaluator
evaluator = registry.get("my_rule")
Next Steps