Agentic Evals · v0.6.0

Know your agent is ready before it ships.

Score LLM and agent outputs with deterministic checks, LLM-as-judge rubrics and trace-aware scorers — then turn the results into a pass/fail release gate for CI.

Framework-agnostic Provider-neutral judges MIT licensed
Evaluate / CIGate
$ agentic-evals run agentic-evals :: support_answers -------------------------------- [PASS] refund-status contains=1.00 [PASS] order-lookup tool_recall=1.00 [PASS] json-contract valid_json=1.00 [FAIL] escalation latency 2.4s > 2.0s 3/4 passed (75.0%) gate: min_pass_rate 0.95 not met exit 1
DeterministicLLM judgeTrajectory

One engine for every kind of check.

Start with a plain input/expected row. Move to traces, tool calls and release gates when you need them — it's the same engine underneath.

01

Deterministic scoring

Exact match, substring, regex, JSON Schema, JSON diff, numeric ranges and Levenshtein similarity. No model call required.

02

LLM-as-judge

Rubric scoring with ten built-in templates, from factuality to SQL correctness. You supply complete_fn — any provider works.

03

Trace-aware

Score the path, not just the answer: tool-call precision and recall, redundant calls, trajectory efficiency, latency, cost and turns.

04

Release gates

Set thresholds on pass rate, average score, latency and total cost. The CLI exits 0 or 1, so CI can block the merge.

05

Live targets

Run a suite against a Python callable or an HTTP endpoint instead of pre-recorded samples.

06

Eval packs

Share a suite and the scorers it needs as one YAML or JSON file. Three packs ship built in.

From first eval to CI gate.

Three ways in, depending on how much structure you need.

# support_eval.py
from agentic_evals import Eval, contains
from my_app import load_rows, support_agent


def cites_order_id(output: str) -> bool:
    return "ORD-" in output


Eval(
    "support_answers",
    data=lambda: load_rows("evals/support.jsonl"),
    task=support_agent,
    scores=[contains, cites_order_id],
    threshold=0.9,
)
The fast path A scorer is any function returning a bool, a float from 0 to 1, or a Score. Run agentic-evals run to discover every eval file and get a single exit code.
from agentic_evals import (
    GateConfig, PythonTarget, TestCase, TestSuite,
    evaluate_gate, run_live_suite,
)

suite = TestSuite(
    name="refunds",
    version="3",
    cases=[
        TestCase(
            id="refund-status",
            name="Looks up the order before answering",
            input={"message": "Where is my refund for ORD-1042?"},
            required_tools=["lookup_order"],
            forbidden_tools=["issue_refund"],
            required_output_fields=["status"],
            max_latency_ms=2000,
            max_cost_usd=0.02,
        ),
    ],
)

report = run_live_suite(suite, PythonTarget(callable_path="my_app:run_case"))
decision = evaluate_gate(report, GateConfig(
    min_pass_rate=0.95,
    max_average_latency_ms=1500,
    max_total_cost_usd=0.25,
))
if not decision.passed:
    raise SystemExit(f"Release gate failed: {decision.reasons}")
Declarative suites Expectations cover tool calls, output fields, JSON Schema, latency, cost and turn count. Run against a live Python or HTTP target, then gate the release on the report.
from agentic_evals import (
    FACTUALITY, SECURITY, EvaluatorConfig, LLMRubricEvaluator,
    TestCase, default_registry, evaluate_suite,
)
from my_app.llm import complete  # (prompt: str) -> str

registry = default_registry()
registry.register(LLMRubricEvaluator("factuality", FACTUALITY, complete_fn=complete))
registry.register(LLMRubricEvaluator("security", SECURITY, complete_fn=complete))

case = TestCase(
    id="policy-answer",
    name="Answers from the policy without leaking internals",
    evaluators=[
        EvaluatorConfig(name="factuality", threshold=0.8),
        EvaluatorConfig(name="security", threshold=0.9),
        EvaluatorConfig(name="tool_call_recall", threshold=1.0),
    ],
)

report = evaluate_suite(suite, samples, registry=registry)
Bring your own model The package never calls a model itself. Pass a complete_fn and mix rubric judges with deterministic and trajectory scorers on the same case.
pip install agentic-evalsPython 3.10+
16Deterministic and trajectory scorers in the default registry
10Built-in LLM rubric templates
3Shareable eval packs included
0 / 1CI exit code from a single command

Where teams put it to work.

Works with any agent framework — anything that produces an output, and optionally a trace, can be scored.

QA

Answer quality

Grade answers against references with closed-QA and factuality rubrics, backed by deterministic checks.

Tools

Tool use

Confirm the agent called the right tools with the right arguments — and avoided the ones it shouldn't touch.

CI

Release gates

Block merges when pass rate drops or latency and cost budgets are exceeded.

Support

Customer support

Check multi-step conversations for the right lookup, a structured response and a bounded number of turns.

Safety

Safety and privacy

Score outputs for moderation, security and PII leakage with dedicated rubric templates.

Cost

Cost and latency

Track cost and latency per case and across the run. Missing cost data is reported as unavailable, not zero.

Put a gate in front of your next release.

Open source, framework-agnostic, and ready to run in CI.