Deterministic scoring
Exact match, substring, regex, JSON Schema, JSON diff, numeric ranges and Levenshtein similarity. No model call required.
Agentic Evals · v0.6.0
Score LLM and agent outputs with deterministic checks, LLM-as-judge rubrics and trace-aware scorers — then turn the results into a pass/fail release gate for CI.
Capabilities
Start with a plain input/expected row. Move to traces, tool calls and release gates when you need them — it's the same engine underneath.
Exact match, substring, regex, JSON Schema, JSON diff, numeric ranges and Levenshtein similarity. No model call required.
Rubric scoring with ten built-in templates, from factuality to SQL correctness. You supply complete_fn — any provider works.
Score the path, not just the answer: tool-call precision and recall, redundant calls, trajectory efficiency, latency, cost and turns.
Set thresholds on pass rate, average score, latency and total cost. The CLI exits 0 or 1, so CI can block the merge.
Run a suite against a Python callable or an HTTP endpoint instead of pre-recorded samples.
Share a suite and the scorers it needs as one YAML or JSON file. Three packs ship built in.
Quickstart
Three ways in, depending on how much structure you need.
# support_eval.py
from agentic_evals import Eval, contains
from my_app import load_rows, support_agent
def cites_order_id(output: str) -> bool:
return "ORD-" in output
Eval(
"support_answers",
data=lambda: load_rows("evals/support.jsonl"),
task=support_agent,
scores=[contains, cites_order_id],
threshold=0.9,
)
Score. Run agentic-evals run to discover every eval file and get a single exit code.
from agentic_evals import (
GateConfig, PythonTarget, TestCase, TestSuite,
evaluate_gate, run_live_suite,
)
suite = TestSuite(
name="refunds",
version="3",
cases=[
TestCase(
id="refund-status",
name="Looks up the order before answering",
input={"message": "Where is my refund for ORD-1042?"},
required_tools=["lookup_order"],
forbidden_tools=["issue_refund"],
required_output_fields=["status"],
max_latency_ms=2000,
max_cost_usd=0.02,
),
],
)
report = run_live_suite(suite, PythonTarget(callable_path="my_app:run_case"))
decision = evaluate_gate(report, GateConfig(
min_pass_rate=0.95,
max_average_latency_ms=1500,
max_total_cost_usd=0.25,
))
if not decision.passed:
raise SystemExit(f"Release gate failed: {decision.reasons}")
from agentic_evals import (
FACTUALITY, SECURITY, EvaluatorConfig, LLMRubricEvaluator,
TestCase, default_registry, evaluate_suite,
)
from my_app.llm import complete # (prompt: str) -> str
registry = default_registry()
registry.register(LLMRubricEvaluator("factuality", FACTUALITY, complete_fn=complete))
registry.register(LLMRubricEvaluator("security", SECURITY, complete_fn=complete))
case = TestCase(
id="policy-answer",
name="Answers from the policy without leaking internals",
evaluators=[
EvaluatorConfig(name="factuality", threshold=0.8),
EvaluatorConfig(name="security", threshold=0.9),
EvaluatorConfig(name="tool_call_recall", threshold=1.0),
],
)
report = evaluate_suite(suite, samples, registry=registry)
complete_fn and mix rubric judges with deterministic and trajectory scorers on the same case.
pip install agentic-evalsPython 3.10+Use cases
Works with any agent framework — anything that produces an output, and optionally a trace, can be scored.
Grade answers against references with closed-QA and factuality rubrics, backed by deterministic checks.
Confirm the agent called the right tools with the right arguments — and avoided the ones it shouldn't touch.
Block merges when pass rate drops or latency and cost budgets are exceeded.
Check multi-step conversations for the right lookup, a structured response and a bounded number of turns.
Score outputs for moderation, security and PII leakage with dedicated rubric templates.
Track cost and latency per case and across the run. Missing cost data is reported as unavailable, not zero.
Start evaluating
Open source, framework-agnostic, and ready to run in CI.