Metadata-Version: 2.3
Name: kitty-evals
Version: 0.1.0
Summary: Evaluator toolkit for LLM outputs: pluggable LLM-as-a-judge with structured verdicts
Author: WhiteFox0-0
Author-email: WhiteFox0-0 <swayammmm0@gmail.com>
Requires-Dist: questionary>=2.1
Requires-Python: >=3.13
Description-Content-Type: text/markdown

# kitty-evals

Evaluator toolkit for LLM outputs: an LLM-as-a-judge with pluggable providers, structured verdicts, and concurrent batches.

## Quickstart

```bash
uvx kitty_evals init
```

Opens a blue multi-select of the available evaluators, asks where to put them
(default `evaluation/`), and copies the packages in:

```
evaluation/
└── llm_judge/
    ├── __init__.py
    ├── base.py
    ├── config.py
    ├── ...
```

Non-interactive / scripted:

```bash
uvx kitty_evals init --list            # show the catalog
uvx kitty_evals init llm_judge         # copy one evaluator
uvx kitty_evals init --yes             # every ready evaluator
uvx kitty_evals init -d path/to/dir    # explicit destination
```

## Use a judge

```python
from kitty_evals import Judge

judge = Judge("gpt-4o", temperature=0.2, score_range=(1, 5))
verdict = judge.evaluate("The model said ...", context="task prompt")
print(verdict.score, verdict.rationale)

# Batch: a JSON list (or any iterable of dicts) — results in input order
verdicts = judge.evaluate_many("cases.json", max_concurrency=20)
```

Each item needs a `prompt`; optional `context`, `reference`, plus any extra keys,
which are passed to the judge as a JSON variable block.

## Install for development

```bash
uv sync
uv run pytest
```
