Agentic Evals Documentation
A standalone, framework-agnostic evaluation and scoring engine for LLM and agent outputs.
What is agentic-evals?
agentic-evals is a Python package for evaluating the quality of LLM and agent outputs programmatically. It runs standalone and is not tied to any agent framework or tracing format.
You provide:
- Test cases — expectations (exact match, substring, JSON Schema, tool calls, latency, etc.)
- Evaluation samples — actual outputs and execution traces
- Custom evaluators (optional) — your own scoring logic or LLM judges
agentic-evals returns:
- Per-case scores — individual judgments with explanations
- Summary metrics — pass rate, average score, cost, latency, broken down by metric and by tag
- Pass/fail gates — CI-ready exit codes for release decisions
Installation
pip install agentic-evals
Optional: for voice agent evaluation examples with FastAPI:
pip install "agentic-evals[voice]"
Your first eval in 60 seconds
The simplest way to write an eval:
from agentic_evals import Eval, equals
Eval(
"my_first_eval",
data=[
{"input": "France", "expected": "Paris"},
{"input": "Japan", "expected": "Tokyo"},
],
task=my_agent,
scores=[equals],
)
Save as my_eval.py and run:
python my_eval.py
Or discover and run all eval files:
agentic-evals run
Core Concepts
EvalTrace and EvalSpan
The minimal trace structure agentic-evals inspects — deliberately not tied to any specific tracing tool or framework.
class EvalTrace:
trace_id: str
spans: list[EvalSpan]
total_latency_ms: float
estimated_cost_usd: float | None
metadata: dict[str, Any]
class EvalSpan:
tool_name: str | None
attributes: dict[str, Any]
You build these from whatever trace data you have (AgenticLens, OpenTelemetry, custom instrumentation, etc.):
from agentic_evals import EvalTrace, EvalSpan
trace = EvalTrace(
trace_id="run-123",
spans=[
EvalSpan(tool_name="calculator", attributes={"tool_args": {"expr": "2+2"}}),
EvalSpan(tool_name="search", attributes={}),
],
total_latency_ms=250,
estimated_cost_usd=0.015,
metadata={"turn_count": 2},
)
TestSuite and TestCase
Define expectations for your agent outputs:
| Field |
Type |
Purpose |
| expected_output |
str |
Exact match (whitespace-trimmed) |
| expected_contains |
list[str] |
Required substrings (case-insensitive) |
| output_json_schema |
dict |
JSON Schema Draft 2020-12 validation |
| required_output_fields |
list[str] |
Dotted paths in JSON (e.g., "data.total") |
| required_tools |
list[str] |
Tool calls that must happen |
| forbidden_tools |
list[str] |
Tool calls that must NOT happen |
| required_tool_arguments |
dict[str, list[str]] |
Argument names a tool call must include |
| expected_tool_arguments |
dict[str, dict[str, Any]] |
Argument values a tool call must carry (compared without type coercion) |
| required_tool_order |
list[str] |
Tools whose first calls must come in this order |
| max_latency_ms |
float |
Latency threshold in milliseconds |
| max_cost_usd |
float |
Cost threshold in USD |
| max_turns |
int |
Maximum conversation turns (from trace metadata) |
| evaluators |
list[EvaluatorConfig] |
Custom evaluators to run |
| tags |
list[str] |
Labels for grouping results in the report (e.g., "refunds") |
| metadata |
dict[str, Any] |
Free-form data for custom evaluators; copied onto the case result |
from agentic_evals import TestSuite, TestCase
suite = TestSuite(
name="support-agent",
version="1.0",
cases=[
TestCase(
id="refund-status",
name="Check refund status",
expected_contains=["Your refund", "processing"],
required_tools=["lookup_order"],
max_latency_ms=1000,
tags=["refunds"],
),
],
)
Tool expectations can also state what a tool was called with, and in what order. Arguments are read from each span's attributes["tool_args"]:
TestCase(
id="publish-draft",
name="Saves the draft before publishing it",
required_tools=["save_draft", "publish"],
expected_tool_arguments={"publish": {"visibility": "internal"}},
required_tool_order=["save_draft", "publish"],
)
expected_tool_arguments passes when one call to the tool carries every listed value. required_tool_order passes when each listed tool is first called before the next; calls to other tools in between are fine.
Evaluators
An evaluator is anything with a .name property and an .evaluate(context) -> list[Score] method.
Built-in types:
CallableEvaluator — Wrap any Python function
LLMJudgeEvaluator — Provider-neutral LLM judge
BusinessRuleEvaluator — Named convenience for custom logic
Scores and Reports
Each evaluation produces a Score:
class Score(BaseModel):
name: str # e.g., "factuality", "contains:refund"
metric: str # Grouping key; defaults to name ("contains")
value: float # 0.0 to 1.0
passed: bool # True if value >= threshold
required: bool = True
skipped: bool = False # True when the check did not apply
explanation: str # Why it passed/failed
evaluator_type: str = "deterministic" # or "llm_judge"
metadata: dict[str, Any] = {}
Multiple scores combine into a CaseEvaluation (one per test case), and those roll up into an EvaluationReport:
class EvaluationReport(BaseModel):
suite_name: str
suite_version: str
created_at: datetime
summary: EvaluationSummary # Pass rate, avg score, cost, latency
cases: list[CaseEvaluation] # Per-case details
Breakdowns by metric and tag
An overall pass rate does not show where the failures are. The summary breaks results down two ways:
report = evaluate_suite(suite, samples)
for tag, stats in report.summary.tags.items():
print(f"{tag}: {stats.passed_cases}/{stats.total_cases} cases passed")
for metric, stats in report.summary.metrics.items():
print(f"{metric}: {stats.passed}/{stats.total} checks passed")
| Field |
Type |
Contents |
| summary.tags |
dict[str, TagSummary] |
Per tag: total_cases, passed_cases, failed_cases, pass_rate, average_score |
| summary.metrics |
dict[str, MetricSummary] |
Per metric: total, passed, failed, skipped, pass_rate, average_score |
Checks whose names carry a per-case detail share one metric: contains:refund and contains:shipped both roll up under contains. Each CaseEvaluation also carries its case's tags and metadata, so a saved report can be filtered or regrouped on its own.
Skipped checks
When a check does not apply to a case, return Score.skip(...) instead of a pass. A skipped score never fails the case and stays out of averages and pass rates; it is counted in MetricSummary.skipped.
def cites_order_id(context: EvaluationContext) -> Score:
order_id = context.case.metadata.get("order_id")
if order_id is None:
return Score.skip("cites_order_id", "Case has no order id to cite.")
cited = order_id in context.sample.output
return Score(
name="cites_order_id",
value=float(cited),
passed=cited,
explanation=f"Order id {order_id} cited: {cited}.",
)
APIs
Eval() Quickstart
The fastest, most opinionated API. Perfect for simple cases.
def Eval(
name: str,
data: list[dict] | Callable[[], list[dict]],
task: Callable[[Any], str],
scores: list[Callable],
threshold: float = 1.0,
) -> EvalResult
Parameters:
| Parameter |
Type |
Description |
| name |
str |
Eval name (appears in output) |
| data |
list[dict] or callable |
Test data with "input"/"expected" keys, or callable that returns it |
| task |
callable |
Function under test; called with each row's input value, returns the output |
| scores |
list[callable] |
Scorer functions; each returns 0-1 or bool |
| threshold |
float |
Min score for a case to pass (default 1.0) |
Built-in scorers for Eval():
equals — Exact match
contains — Case-sensitive substring
icontains — Case-insensitive substring
levenshtein — Similarity (0-1)
matches — Regex match
numeric_close — 1% relative tolerance
TestSuite API
More structured, trace-aware evaluation with custom expectations.
def evaluate_suite(
suite: TestSuite,
samples: list[EvaluationSample],
registry: EvaluatorRegistry | None = None,
) -> EvaluationReport
Returns a full EvaluationReport with per-case scores, summary metrics, pass rates, cost/latency, and breakdowns by metric and tag.
Scorers Library
Pre-built evaluators organized by type:
Text Scorers (No LLM)
from agentic_evals import scorers
# Deterministic checks
scorer = scorers.text.exact_match
scorer = scorers.text.levenshtein_similarity
scorer = scorers.text.valid_json
scorer = scorers.text.json_diff
scorer = scorers.text.regex_match
Rubric Scorers (LLM-Graded)
from agentic_evals import LLMRubricEvaluator, FACTUALITY, default_registry
# Register an LLM rubric evaluator
registry = default_registry()
registry.register(
LLMRubricEvaluator(
"factuality_check",
FACTUALITY,
complete_fn=your_model_call, # Your function to call LLM
)
)
When no built-in template fits, describe the criterion in plain language. make_rubric() builds the prompt, the verdict letters and their scores:
from agentic_evals import LLMRubricEvaluator, make_rubric
concise = make_rubric("concise", "The answer is at most two sentences and has no preamble.")
tone = make_rubric(
"tone",
"The reply is courteous and does not blame the reader.",
levels=[ # best first; scores from 0 to 1
("Courteous throughout.", 1.0),
("Neutral: neither courteous nor rude.", 0.5),
("Rude, dismissive, or blames the reader.", 0.0),
],
)
registry.register(LLMRubricEvaluator("concise", concise, complete_fn=your_model_call))
registry.register(LLMRubricEvaluator("tone", tone, complete_fn=your_model_call))
The default scale is pass/fail. Pass with_reference=True to show the judge a reference next to the output.
Trajectory Scorers (Trace-Aware)
from agentic_evals import scorers
scorer = scorers.trajectory.tool_call_precision
scorer = scorers.trajectory.tool_call_recall
scorer = scorers.trajectory.no_redundant_tool_calls
scorer = scorers.trajectory.trajectory_efficiency
Release Gates
Turn an evaluation report into a pass/fail CI decision:
def evaluate_gate(
report: EvaluationReport,
config: GateConfig,
) -> GateDecision
from agentic_evals import evaluate_gate, GateConfig
decision = evaluate_gate(
report,
GateConfig(
min_pass_rate=0.95, # At least 95% of cases pass
max_average_latency_ms=1500, # Avg latency under 1.5s
max_total_cost_usd=0.25, # Total cost under $0.25
)
)
if not decision.passed:
print(f"Release blocked: {decision.reasons}")
exit(1)
A gate can hold one slice of the report to a stricter bar than the suite as a whole, keyed by case tag or by Score.metric:
GateConfig(
min_pass_rate=0.95,
min_tag_pass_rate={"safety": 1.0}, # Every safety-tagged case passes
min_metric_pass_rate={"forbidden_tool": 1.0}, # No forbidden tool call, anywhere
)
A tag or metric named in the config but missing from the report fails the gate, so a renamed tag cannot quietly switch a check off.
Live Targets
Run a suite against a real system (trusted code or HTTP endpoint) instead of pre-recorded samples:
def run_live_suite(
suite: TestSuite,
target: LiveTarget,
registry: EvaluatorRegistry | None = None,
) -> EvaluationReport
Python Target (executes local code):
from agentic_evals import PythonTarget, run_live_suite
report = run_live_suite(
suite,
PythonTarget(callable_path="my_module:my_agent"),
# my_module.py must have a my_agent(input, case=None) function
)
HTTP Target (calls a remote endpoint):
from agentic_evals import HTTPTarget, run_live_suite
report = run_live_suite(
suite,
HTTPTarget(
url="http://localhost:8000/evaluate",
method="POST",
headers={"Authorization": "Bearer token"},
),
)
Advanced Topics
Custom Evaluators
Write your own evaluation logic with full access to the test case, output, and trace.
from agentic_evals import (
LLMJudgeEvaluator, CallableEvaluator,
EvaluationContext, Score, EvaluatorRegistry
)
# LLM-based judge
def my_judge(context: EvaluationContext) -> Score:
output = context.sample.output
expected = context.case.expected_output
# Call your LLM here
judgment = your_llm_model(
f"Is this output correct? {output} vs {expected}"
)
return Score(
name="llm_judgment",
value=1.0 if judgment else 0.0,
passed=judgment,
explanation="LLM agrees" if judgment else "LLM disagrees",
)
registry = EvaluatorRegistry()
registry.register(LLMJudgeEvaluator("my_judge", my_judge))
Then use it in your test case:
case = TestCase(
id="case-1",
name="Custom judge",
evaluators=[EvaluatorConfig(name="my_judge", threshold=0.8)],
)
Eval Packs
Bundle a TestSuite with the scorers it needs into one shareable YAML/JSON file:
name: my-eval-pack
description: Checks answers for factual consistency.
required_scorers:
- factuality
suite:
name: my-eval-pack
version: "1"
cases:
- id: case-1
name: Check if output is factual
input: What is the capital of France?
evaluators:
- name: factuality
threshold: 0.8
config:
reference: Paris is the capital of France.
Load and use it:
from pathlib import Path
from agentic_evals import (
FACTUALITY, LLMRubricEvaluator, default_registry, evaluate_suite, load_pack,
)
registry = default_registry()
registry.register(LLMRubricEvaluator("factuality", FACTUALITY, complete_fn=your_model_call))
pack = load_pack(Path("path/to/pack.yaml"), registry=registry)
report = evaluate_suite(pack.to_suite(), samples, registry=registry)
Built-in packs:
tool-use-correctness — Evaluate tool selection and arguments
json-output-contract — Validate JSON structure
factual-qa — QA with factuality rubric
from agentic_evals import load_builtin_pack
pack = load_builtin_pack("tool-use-correctness")
Evaluator Registry
Manage evaluators centrally so test cases can reference them by name:
from agentic_evals import (
EvaluatorRegistry, default_registry,
LLMJudgeEvaluator, CallableEvaluator
)
# Start with defaults
registry = default_registry()
# Add custom evaluators
def my_rule(context):
return Score(...)
registry.register(
CallableEvaluator("my_rule", my_rule)
)
# Get an evaluator
evaluator = registry.get("my_rule")
Next Steps