Metadata-Version: 2.4
Name: simulhausen
Version: 0.1.0
Summary: Local simulation testing for agno agents: an LLM plays the user, an LLM judge grades the dialogue criterion by criterion.
Keywords: agents,agno,llm,testing,simulation,evaluation,llm-as-a-judge,pytest
Author: Ilya Lubenets
Author-email: Ilya Lubenets <lubenets.ilya.igorevich@gmail.com>
License-Expression: Apache-2.0
License-File: LICENSE
License-File: NOTICE
Classifier: Development Status :: 4 - Beta
Classifier: Framework :: Pytest
Classifier: Intended Audience :: Developers
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3.14
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Topic :: Software Development :: Testing
Classifier: Typing :: Typed
Requires-Dist: agno>=3.0,<4
Requires-Dist: pydantic>=2.10
Requires-Dist: rich>=13
Requires-Python: >=3.12
Project-URL: Homepage, https://github.com/ormeilu/simulhausen
Project-URL: Repository, https://github.com/ormeilu/simulhausen
Project-URL: Issues, https://github.com/ormeilu/simulhausen/issues
Project-URL: Changelog, https://github.com/ormeilu/simulhausen/blob/main/CHANGELOG.md
Description-Content-Type: text/markdown

# simulhausen

[![PyPI](https://img.shields.io/pypi/v/simulhausen)](https://pypi.org/project/simulhausen/)
[![Python](https://img.shields.io/pypi/pyversions/simulhausen)](https://pypi.org/project/simulhausen/)
[![CI](https://github.com/ormeilu/simulhausen/actions/workflows/ci.yml/badge.svg)](https://github.com/ormeilu/simulhausen/actions/workflows/ci.yml)
[![License](https://img.shields.io/badge/license-Apache--2.0-blue)](https://github.com/ormeilu/simulhausen/blob/main/LICENSE)

Simulation testing for [agno](https://github.com/agno-agi/agno) agents that runs entirely on your side.
An LLM plays the user and holds a multi-turn dialogue with your agent, team or workflow. Then an LLM judge
grades the result against your criteria, one verdict and one reason per criterion.

The API follows [LangWatch Scenario](https://github.com/langwatch/scenario), minus the cloud: no LangWatch
account, no tracing backend, no litellm. The user simulator and the judge are plain agno agents on whatever
model you configure, and the reports land in your terminal, in a JSON file and in a markdown comment for your
pull request.

```python
import simulhausen as sim

result = await sim.run(
    name="refund: duplicated charge",
    description="The user was charged twice for order 1042 and wants the extra charge refunded.",
    agents=[
        sim.agno_adapter(lambda: support_agent),
        sim.UserSimulatorAgent(),
        sim.JudgeAgent(criteria=[
            "The agent looks the order up before promising anything",
            "The agent does not invent a refund amount",
        ]),
    ],
)
assert result.success, result.failure_summary()
```

## Install

```bash
uv add --dev simulhausen
```

or `pip install simulhausen`. Python 3.12+ and agno 3.x. The pytest plugin registers itself on install.

## Quick start

Configure the model for the simulator and the judge once, in `conftest.py`. Keep model construction inside the
fixture, so collecting tests never needs credentials:

```python
# conftest.py
import pytest

import simulhausen as sim


@pytest.fixture(scope="session", autouse=True)
def configure_simulations() -> None:
    from agno.models.openai import OpenAIChat

    sim.configure(model=OpenAIChat(id="gpt-4.1-mini"), max_turns=8)
```

Then write scenarios as ordinary async tests:

```python
# test_support.py
import pytest

import simulhausen as sim


def _support_agent():
    from myapp.agents import support_agent  # lazy: imported only when the test runs

    return support_agent


def _no_lookup_yet(state: sim.ScenarioState) -> None:
    assert not state.has_tool_call("lookup_order"), state.tool_call_names()


@pytest.mark.simulation
async def test_refund_asks_for_order_number() -> None:
    result = await sim.run(
        name="refund: missing order number",
        description="The user wants a refund but does not say which order it is about.",
        agents=[
            sim.agno_adapter(_support_agent, user_id="user-42"),
            sim.UserSimulatorAgent(),
            sim.JudgeAgent(criteria=["The agent asks for the order number before anything else"]),
        ],
        script=[
            sim.user("i want my money back"),
            sim.agent(),
            _no_lookup_yet,
            sim.judge(),
        ],
    )
    assert result.success, result.failure_summary()
```

`agno_adapter` takes the agent itself or a factory, and forwards extra keyword arguments to `arun`
(`session_state=...`, `user_id=...`). The whole scenario runs in one agno session. A runnable version lives in
[`examples/`](https://github.com/ormeilu/simulhausen/tree/main/examples).

## Scripts

Without `script`, a run is `[proceed()]` followed by the judge's verdict. With a script you decide what happens
and in which order:

| Step | What it does |
|---|---|
| `user("text")` | user message with fixed text |
| `user()` | user message written by the simulator |
| `agent()` | call the agent under test |
| `agent("text")` | inject an agent reply without calling the agent |
| `message(role, "text")` | append any message as is, to seed earlier context |
| `proceed(turns=None, on_turn=..., on_step=...)` | simulator and agent talk on their own (see below) |
| `judge()` | final verdict on the judge's criteria; later steps still run |
| `judge(criteria=[...], additional_context=...)` | checkpoint: fails the scenario on the spot if a criterion fails, goes on otherwise |
| `succeed(reasoning)` / `fail(reasoning)` | end the scenario right here |
| any callable | gets the `ScenarioState`; sync or async; raise `AssertionError` to fail |

`proceed()` stops at the first of these: `turns` turns played, `max_turns` reached, the simulator reporting the
goal as achieved, or the judge concluding it has seen enough. The judge is not consulted before `min_turns`
agent turns. If it concludes while some criteria are still inconclusive and turns remain, the dialogue goes on
to collect the missing evidence.

When a script has a judge with criteria but no `judge()` step, the judge runs after the last step. It also runs
again there if the dialogue went on after its last verdict, so the final verdict always covers the whole
dialogue; a new final verdict replaces the previous one, while checkpoint verdicts accumulate.
`additional_context` gives the judge evidence the dialogue cannot show, such as a database row the agent was
supposed to write.

## Verdicts and statuses

The judge answers per criterion: `passed`, `failed` or `inconclusive`, each with a restated requirement and its
reasoning. A criterion phrased as a prohibition ("the agent must not X") passes when X did not happen.
`inconclusive` means the evidence was not there, for example because the dialogue ended before the moment the
criterion depends on.

Every result gets one of three statuses:

- `PASS`: every criterion passed and no assertion failed.
- `FAIL`: a criterion failed, a script assertion failed, the judge score was below threshold, or `fail()` ran.
- `WARN`: the scenario could not be verified. Either the harness broke (`SimulationHarnessError`: the judge or
  the simulator returned no structured output, the model backend kept failing), or some criteria came back
  inconclusive and none failed.

WARN is kept out of the pass rate on purpose. A broken judge says nothing about your agent, and counting it as
a failure turns the pass rate into a measure of helper-LLM reliability. `pass_rate = PASS / (PASS + FAIL)`, and
the WARN count is reported next to it. Harness errors are still raised, so a plain pytest run fails on them.

`ScenarioResult` carries `verdicts`, `passed_criteria`, `failed_criteria`, `inconclusive_criteria`, the judge's
`reasoning`, `messages`, `tool_calls`, `turns`, `total_time` and `agent_time`.

For an overall quality gate on top of the criteria, `JudgeAgent(criteria, score_threshold=7)` also grades the
whole dialogue 1-10 with agno's `JudgeScorer`. Calibrate the threshold on known good and bad runs first.

## Checking tools

`agno_adapter` collects every `ToolExecution` of the run, including those of team members, so assertions work
for teams too:

```python
state.has_tool_call("lookup_order")         # ran cleanly at least once
state.tool_call_names()                     # names in call order, with repeats
state.has_tool_calls_in_order(["search", "lookup_order"])  # subsequence; other calls may sit between
state.last_tool_call("lookup_order")        # the ToolExecution, or None
```

Only clean executions count: a call that errored, was rejected or is paused waiting for confirmation does not
satisfy `has_tool_call`. That matches `agno.scorer.ToolCallScorer`.

The judge sees the same evidence: the transcript, the RAG references the agent retrieved (with source URLs) and
a capped summary of tool executions, errors included. Each block is fenced with a random per-call tag and the
judge is told never to follow instructions inside them.

## Writing scenarios that hold up

One scenario checks one behaviour from one starting state. When it fails, the name alone should tell you what
broke.

- `name` is the outcome in a few words, readable in a PR comment: `refund: asks for the order number`.
- `description` is for the simulator: starting conditions, the user's goal, how they behave. It is not a
  criterion.
- Each criterion is a single observable requirement. "The answer is good" is not one.

Check in code whatever code can check: tool calls and their order, exact identifiers, an expected error, a call
that must not happen. Leave semantics to the judge: did the agent make something up, explain a limitation,
stick to what the tools returned.

Spell out the user messages that matter for the regression with `user("...")`. Use `user()` or `proceed()` only
where variety in the user's wording is part of what you test. Most regressions need one or two user/agent
pairs; every extra turn costs money and adds variance.

Do not paper over a reproducible failure with reruns. Look at `failure_summary()`, the tool evidence and the
transcript in the JSON report first: they tell an agent bug from a broken backend, a wrong assertion or a vague
criterion.

## Running

Simulation tests call real models, so keep them out of the default run:

```toml
[tool.pytest.ini_options]
addopts = "-m 'not simulation'"
asyncio_mode = "auto"
```

```bash
uv run pytest -m simulation                    # failures fail the run
uv run pytest -m simulation --sim-benchmark \
  --sim-report reports/simulations.json \
  --sim-report-md reports/simulations.md       # always exits 0, writes reports
uv run pytest -m simulation -k refund -s --sim-debug  # you type the user's messages
```

Benchmark mode is for CI jobs that should stay green and post the outcome instead. After any run with
simulations, a rich table with every scenario, its verdicts, tools and dialogue is printed. With
`pytest-rerunfailures` only the last attempt is reported.

Parallel runs work with `pytest-xdist` (`-n 4`): each worker is its own process with its own event loop, and
the controller merges every worker's results into one report. Prefer that to `asyncio.gather` inside a test,
because agno agents and HTTP clients are usually process-wide singletons. When an agent's async clients bind to
the first event loop they see, run simulations on one session-scoped loop:

```python
# conftest.py
def pytest_collection_modifyitems(items):
    for item in items:
        if item.get_closest_marker("simulation"):
            item.add_marker(pytest.mark.asyncio(loop_scope="session"), append=False)
```

### Reports in pull requests

The markdown report starts with a hidden `<!-- simulhausen-report -->` marker and a headline such as
`🟡 88% ▓▓▓▓▓▓▓▓▓░ · PASS 154 / FAIL 21`. The emoji follows the pass rate (🟢 from 90% with no WARN, 🟡 from
75%, 🔴 below), since one failure out of 175 is noise, not a regression. The bar shows passed, not verified and
failed scenarios out of all of them. On GitHub Actions:

```yaml
- run: uv run pytest -m simulation --sim-benchmark --sim-report-md reports/simulations.md
- if: github.event_name == 'pull_request'
  run: gh pr comment ${{ github.event.pull_request.number }} --body-file reports/simulations.md
  env:
    GH_TOKEN: ${{ github.token }}
```

The report links the CI job when `CI_JOB_URL` (GitLab) or the `GITHUB_*` run variables are set.

## Configuration

```python
sim.configure(
    model=...,             # agno Model for the simulator and the judge (required)
    parser_model=...,      # optional agno Model that turns their answers into the schema
    max_turns=10,          # ceiling for proceed() and unscripted runs
    min_turns=1,           # the judge may not end the dialogue earlier
    verbose=True,          # log every message and verdict (logger "simulhausen")
    debug=False,           # ask for the user's messages on stdin
    user_name="User",      # labels in reports
    agent_name="Agent",
)
```

`run()` accepts `max_turns`, `min_turns`, `verbose` and `debug` per scenario. `UserSimulatorAgent` and
`JudgeAgent` accept their own `model` and `parser_model`, plus `system_prompt` to replace the built-in prompt
([`simulhausen.prompts`](https://github.com/ormeilu/simulhausen/blob/main/src/simulhausen/prompts.py)). The simulator also takes a `persona`.

Set `parser_model` when your backend breaks structured output in thinking mode. Some vLLM setups with a
reasoning parser return empty or broken content when `response_format` meets reasoning; with a parser model,
agno sends `response_format` only to the parser call. The helper agents also retry an off-schema answer twice,
with a stricter instruction each time.

## Custom adapters

`agno_adapter` covers agno agents, teams and workflows. It retries backend failures (a run that ended in
`RunStatus.error`, a proxy's HTML error page) and raises `SimulationBackendError` when they persist, while a
guardrail refusal (`InputCheckError`/`OutputCheckError`) counts as the agent's answer. For agents with
`output_schema`, pass `content_fn` to render the structured reply as text.

Anything else is a function or an `AgentAdapter` subclass:

```python
async def call_my_api(input: sim.AgentInput) -> str:
    return await my_client.chat(input.thread_id, input.last_new_user_message_str())

adapter = sim.function_adapter(call_my_api)


class MyAdapter(sim.AgentAdapter):
    async def call(self, input: sim.AgentInput) -> sim.AgentTurn:
        reply = await my_agent.respond(input.messages)
        return sim.AgentTurn(content=reply.text, tools=reply.tool_executions)
```

`AgentInput` has the whole transcript (`messages`), what changed since the agent last spoke (`new_messages`)
and the live `scenario_state`.

## Compared with LangWatch Scenario

Ported: the script steps (`user`, `agent`, `message`, `proceed`, `judge` with inline criteria and
`additional_context`, `succeed`, `fail`), per-criterion `passed`/`failed`/`inconclusive` verdicts, the two-phase
judge (an argument-free continue/conclude decision, then the verdict), `min_turns`, simulator `persona` and
`system_prompt`, debug mode, `total_time`/`agent_time`.

Different: models are agno `Model` objects instead of litellm strings; the simulator reports `goal_achieved` so
unscripted dialogues end without filler; harness failures and inconclusive scenarios are WARN and stay out of
the pass rate; a script that ends without a judge verdict passes if its assertions passed; the judge also gets
RAG references and tool executions.

Not ported: anything that needs LangWatch (tracing, events, evaluators, remote traces), voice and realtime
agents, the red-team agents, the joblib cache and long-transcript discovery tools.

## Development

```bash
uv sync
just ci          # ruff, ty, pytest with coverage
just precommit   # every prek hook
```

## License

Apache-2.0, see [LICENSE](https://github.com/ormeilu/simulhausen/blob/main/LICENSE). simulhausen began as a local reimplementation of LangWatch Scenario
(Apache-2.0, Copyright LangWatch); its API and parts of its prompts follow Scenario. See [NOTICE](https://github.com/ormeilu/simulhausen/blob/main/NOTICE).
