Metadata-Version: 2.4
Name: persona-lattice
Version: 0.1.0.dev3
Summary: pytest for your user population — coverage-based persona testing for conversational AI systems
License: Apache-2.0
Keywords: testing,llm,simulation,personas,evaluation,guardrails
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: Apache Software License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Software Development :: Testing
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: httpx>=0.27
Requires-Dist: pyyaml>=6.0
Provides-Extra: personahub
Requires-Dist: datasets; extra == "personahub"
Provides-Extra: dev
Requires-Dist: pytest; extra == "dev"
Requires-Dist: ruff; extra == "dev"
Requires-Dist: mypy; extra == "dev"
Requires-Dist: types-PyYAML; extra == "dev"
Dynamic: license-file

# persona-lattice

**pytest for your user population.**

Coverage-based persona testing for conversational AI systems: factorize your
simulated users into a lattice of *communication style x behavioral
disposition x domain*, drive every cell against your system, and gate CI on
per-cell pass rates.

> **0.1.0.dev0 — API still settling.** Interfaces may change between dev
> releases; pin exact versions if you depend on this today.

## Why

Simulated-user tools exist, but they score one scenario at a time. What they
leave open is *coverage semantics*: a factorized population, balanced cell
assignment with synthetic backfill, **bidirectional expectations** (attack
cells must trigger defenses; benign cells carry false-positive budgets),
pluggable product-side oracles instead of judge-only scoring, and per-cell
regression reporting built for CI.

## Install

```bash
pip install persona-lattice              # core
pip install 'persona-lattice[personahub]'  # + PersonaHub corpus sampling
```

## Quickstart

```python
from persona_lattice import Lattice, dispositions, PersonaHubSource
from persona_lattice.llm import OpenAICompatibleClient
from persona_lattice.sut import OpenAICompatibleSUT
from persona_lattice.oracles import GuardrailOracle, LLMJudgeOracle

llm = OpenAICompatibleClient(
    base_url="https://api.example.com/v1", api_key="...", model="your-model"
)

lattice = Lattice(
    styles=PersonaHubSource(),
    dispositions=dispositions.default(),  # or a custom subset / your own
    domains={"support": SUPPORT_CFG, "sales": SALES_CFG},
    affinity={"sales": {"cooperative", "impatient", "adversarial_pricing"}},
)

run = lattice.run(
    sut=OpenAICompatibleSUT(llm),  # or implement SUTAdapter for your product API
    oracles=[GuardrailOracle(false_positive_budget=1), LLMJudgeOracle(llm)],
    persona_client=llm,
    target_n=120,
    seed=42,
    workers=8,
)

run.save_report("report.json")
assert run.cell("support", "adversarial_pii").pass_rate >= 0.8  # CI gate
```

For a zero-network end-to-end demo, run
[`examples/toy_support_bot.py`](examples/toy_support_bot.py).

## Features

- **Coverage lattice** — style x disposition x domain with an affinity map for
  which combinations are valid; seeded round-robin balancing; synthetic-persona
  backfill so every required cell is always populated.
- **9 built-in dispositions** (6 non-cooperative), each carrying its expected
  outcomes: guardrail triggers, safety gates, sentiment direction, turn bounds.
- **Persona sources** — inline, JSONL, or seeded sampling from the public
  PersonaHub corpus; plus an LLM classifier to map corpus personas onto
  dispositions.
- **Pluggable SUT adapters** — `CallableSUT` for in-process functions,
  `OpenAICompatibleSUT` for any chat endpoint, or implement `SUTAdapter`
  against your product's API to surface real guardrail events and halts.
- **Bidirectional oracles** — `GuardrailOracle` (attacks must trigger; benign
  cells get a false-positive budget) and `LLMJudgeOracle` (sentiment + turn
  bounds) on a soft-assertion engine (`in_range`, `one_of`, `at_most`).
- **Parallel runner + reports** — thread-pooled execution, a formal JSON report
  schema, a self-contained HTML domain x disposition pass-rate heatmap, and
  `run.cell(...).pass_rate` accessors for CI gates.
- **Experimental pytest plugin** — parametrize tests over lattice cells.

Planned for v0.2: multilingual cells, a simulator-fidelity validator, and
report diffing across runs.

## License

Apache-2.0 — see [LICENSE](LICENSE).
