Agentic Evals Documentation

A standalone, framework-agnostic evaluation and scoring engine for LLM and agent outputs.

What is agentic-evals?

agentic-evals is a Python package for evaluating the quality of LLM and agent outputs programmatically. It runs standalone and is not tied to any agent framework or tracing format.

You provide:

  • Test cases — expectations (exact match, substring, JSON Schema, tool calls, latency, etc.)
  • Evaluation samples — actual outputs and execution traces
  • Custom evaluators (optional) — your own scoring logic or LLM judges

agentic-evals returns:

  • Per-case scores — individual judgments with explanations
  • Summary metrics — pass rate, average score, cost, latency
  • Pass/fail gates — CI-ready exit codes for release decisions

Installation

pip install agentic-evals

Optional: for voice agent evaluation examples with FastAPI:

pip install "agentic-evals[voice]"

Your first eval in 60 seconds

The simplest way to write an eval:

from agentic_evals import Eval, equals

Eval(
    "my_first_eval",
    data=[
        {"input": "France", "expected": "Paris"},
        {"input": "Japan", "expected": "Tokyo"},
    ],
    task=lambda x: my_agent(x["input"]),
    scores=[equals],
)

Save as my_eval.py and run:

python my_eval.py

Or discover and run all eval files:

agentic-evals run

Core Concepts

EvalTrace and EvalSpan

The minimal trace structure agentic-evals inspects — deliberately not tied to any specific tracing tool or framework.

class EvalTrace: trace_id: str spans: list[EvalSpan] total_latency_ms: float estimated_cost_usd: float | None metadata: dict[str, Any] class EvalSpan: tool_name: str | None attributes: dict[str, Any]

You build these from whatever trace data you have (AgenticLens, OpenTelemetry, custom instrumentation, etc.):

from agentic_evals import EvalTrace, EvalSpan

trace = EvalTrace(
    trace_id="run-123",
    spans=[
        EvalSpan(tool_name="calculator", attributes={"tool_args": {"expr": "2+2"}}),
        EvalSpan(tool_name="search", attributes={}),
    ],
    total_latency_ms=250,
    estimated_cost_usd=0.015,
    metadata={"turn_count": 2},
)

TestSuite and TestCase

Define expectations for your agent outputs:

Field Type Purpose
expected_output str Exact match (whitespace-trimmed)
expected_contains list[str] Required substrings (case-insensitive)
output_json_schema dict JSON Schema Draft 2020-12 validation
required_output_fields list[str] Dotted paths in JSON (e.g., "data.total")
required_tools list[str] Tool calls that must happen
forbidden_tools list[str] Tool calls that must NOT happen
required_tool_arguments dict[str, list[str]] Specific arguments a tool must receive
max_latency_ms float Latency threshold in milliseconds
max_cost_usd float Cost threshold in USD
max_turns int Maximum conversation turns (from trace metadata)
evaluators list[EvaluatorConfig] Custom evaluators to run
from agentic_evals import TestSuite, TestCase

suite = TestSuite(
    name="support-agent",
    version="1.0",
    cases=[
        TestCase(
            id="refund-status",
            name="Check refund status",
            expected_contains=["Your refund", "processing"],
            required_tools=["lookup_order"],
            max_latency_ms=1000,
        ),
    ],
)

Evaluators

An evaluator is anything with a .name property and an .evaluate(context) -> list[Score] method.

Built-in types:

  • CallableEvaluator — Wrap any Python function
  • LLMJudgeEvaluator — Provider-neutral LLM judge
  • BusinessRuleEvaluator — Named convenience for custom logic

Scores and Reports

Each evaluation produces a Score:

class Score(BaseModel):
    name: str                    # e.g., "factuality", "exact_match"
    value: float                 # 0.0 to 1.0
    passed: bool                 # True if value >= threshold
    required: bool = True
    explanation: str             # Why it passed/failed
    evaluator_type: str = "deterministic"  # or "llm_judge"
    metadata: dict[str, Any] = {}

Multiple scores combine into a CaseEvaluation (one per test case), and those roll up into an EvaluationReport:

class EvaluationReport(BaseModel):
    suite_name: str
    suite_version: str
    created_at: datetime
    summary: EvaluationSummary  # Pass rate, avg score, cost, latency
    cases: list[CaseEvaluation] # Per-case details

APIs

Eval() Quickstart

The fastest, most opinionated API. Perfect for simple cases.

def Eval( name: str, data: list[dict] | Callable[[], list[dict]], task: Callable[[Any], str], scores: list[Callable], threshold: float = 1.0, ) -> EvalResult

Parameters:

Parameter Type Description
name str Eval name (appears in output)
data list[dict] or callable Test data with "input"/"expected" keys, or callable that returns it
task callable Function to evaluate, takes input dict, returns string
scores list[callable] Scorer functions; each returns 0-1 or bool
threshold float Min score for a case to pass (default 1.0)

Built-in scorers for Eval():

  • equals — Exact match
  • contains — Case-sensitive substring
  • icontains — Case-insensitive substring
  • levenshtein — Similarity (0-1)
  • matches — Regex match
  • numeric_close — 1% relative tolerance

TestSuite API

More structured, trace-aware evaluation with custom expectations.

def evaluate_suite( suite: TestSuite, samples: list[EvaluationSample], registry: EvaluatorRegistry | None = None, ) -> EvaluationReport

Returns a full EvaluationReport with per-case scores, summary metrics, pass rates, cost/latency.

Scorers Library

Pre-built evaluators organized by type:

Text Scorers (No LLM)

from agentic_evals import scorers

# Deterministic checks
scorer = scorers.text.exact_match
scorer = scorers.text.levenshtein_similarity
scorer = scorers.text.valid_json
scorer = scorers.text.json_diff
scorer = scorers.text.regex_match

Rubric Scorers (LLM-Graded)

from agentic_evals import LLMRubricEvaluator, FACTUALITY, default_registry

# Register an LLM rubric evaluator
registry = default_registry()
registry.register(
    LLMRubricEvaluator(
        "factuality_check",
        FACTUALITY,
        complete_fn=your_model_call,  # Your function to call LLM
    )
)

Trajectory Scorers (Trace-Aware)

from agentic_evals import scorers

scorer = scorers.trajectory.tool_call_precision
scorer = scorers.trajectory.tool_call_recall
scorer = scorers.trajectory.no_redundant_tool_calls
scorer = scorers.trajectory.trajectory_efficiency

Release Gates

Turn an evaluation report into a pass/fail CI decision:

def evaluate_gate( report: EvaluationReport, config: GateConfig, ) -> GateDecision
from agentic_evals import evaluate_gate, GateConfig

decision = evaluate_gate(
    report,
    GateConfig(
        min_pass_rate=0.95,           # At least 95% of cases pass
        max_average_latency_ms=1500,  # Avg latency under 1.5s
        max_total_cost_usd=0.25,      # Total cost under $0.25
    )
)

if not decision.passed:
    print(f"Release blocked: {decision.reasons}")
    exit(1)

Live Targets

Run a suite against a real system (trusted code or HTTP endpoint) instead of pre-recorded samples:

def run_live_suite( suite: TestSuite, target: LiveTarget, registry: EvaluatorRegistry | None = None, ) -> EvaluationReport

Python Target (executes local code):

from agentic_evals import PythonTarget, run_live_suite

report = run_live_suite(
    suite,
    PythonTarget(callable_path="my_module:my_agent"),
    # my_module.py must have a my_agent(input, case=None) function
)

HTTP Target (calls a remote endpoint):

from agentic_evals import HTTPTarget, run_live_suite

report = run_live_suite(
    suite,
    HTTPTarget(
        url="http://localhost:8000/evaluate",
        method="POST",
        headers={"Authorization": "Bearer token"},
    ),
)

Advanced Topics

Custom Evaluators

Write your own evaluation logic with full access to the test case, output, and trace.

from agentic_evals import (
    LLMJudgeEvaluator, CallableEvaluator,
    EvaluationContext, Score, EvaluatorRegistry
)

# LLM-based judge
def my_judge(context: EvaluationContext) -> Score:
    output = context.sample.output
    expected = context.case.expected_output

    # Call your LLM here
    judgment = your_llm_model(
        f"Is this output correct? {output} vs {expected}"
    )

    return Score(
        name="llm_judgment",
        value=1.0 if judgment else 0.0,
        passed=judgment,
        explanation="LLM agrees" if judgment else "LLM disagrees",
    )

registry = EvaluatorRegistry()
registry.register(LLMJudgeEvaluator("my_judge", my_judge))

Then use it in your test case:

case = TestCase(
    id="case-1",
    name="Custom judge",
    evaluators=[EvaluatorConfig(name="my_judge", threshold=0.8)],
)

Eval Packs

Bundle a TestSuite with the scorers it needs into one shareable YAML/JSON file:

name: my-eval-pack
version: 1.0
required_scorers:
  - levenshtein_similarity
  - factuality

cases:
  - id: case-1
    name: Check if output is factual
    expected_output: Paris
    evaluators:
      - name: factuality
        threshold: 0.9

Load and use it:

from agentic_evals import load_pack, evaluate_suite, default_registry

pack = load_pack("path/to/pack.yaml")
report = evaluate_suite(pack.to_suite(), samples, registry=default_registry())

Built-in packs:

  • tool-use-correctness — Evaluate tool selection and arguments
  • json-output-contract — Validate JSON structure
  • factual-qa — QA with factuality rubric
from agentic_evals import load_builtin_pack

pack = load_builtin_pack("tool-use-correctness")

Evaluator Registry

Manage evaluators centrally so test cases can reference them by name:

from agentic_evals import (
    EvaluatorRegistry, default_registry,
    LLMJudgeEvaluator, CallableEvaluator
)

# Start with defaults
registry = default_registry()

# Add custom evaluators
def my_rule(context):
    return Score(...)

registry.register(
    CallableEvaluator("my_rule", my_rule)
)

# Get an evaluator
evaluator = registry.get("my_rule")