Metadata-Version: 2.4
Name: rupsycho
Version: 1.0.0
Summary: R.U.Psycho: robust, unified and reproducible psychometric testing of language models.
License: MIT
License-File: LICENSE
Keywords: LLM,psychometrics,social science,questionnaires,reproducibility
Author: Julian Schelb
Author-email: julian.schelb@uni-konstanz.de
Requires-Python: >=3.10
Classifier: Development Status :: 5 - Production/Stable
Classifier: Intended Audience :: Science/Research
Classifier: License :: OSI Approved :: MIT License
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3.14
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Topic :: Scientific/Engineering :: Information Analysis
Classifier: Typing :: Typed
Provides-Extra: all
Provides-Extra: configurator
Provides-Extra: deepseek
Provides-Extra: dev
Provides-Extra: docs
Provides-Extra: google
Provides-Extra: huggingface
Provides-Extra: models
Provides-Extra: notebook
Provides-Extra: ollama
Provides-Extra: openai
Provides-Extra: quantization
Provides-Extra: test
Requires-Dist: accelerate (>=0.25.0) ; extra == "all"
Requires-Dist: accelerate (>=0.25.0) ; extra == "huggingface"
Requires-Dist: accelerate (>=0.25.0) ; extra == "models"
Requires-Dist: aiofiles (>=24.1.0)
Requires-Dist: bitsandbytes (>=0.43.0) ; (sys_platform != "darwin") and (extra == "all")
Requires-Dist: bitsandbytes (>=0.43.0) ; extra == "quantization"
Requires-Dist: build (>=1.6.1) ; extra == "dev"
Requires-Dist: hypothesis (>=6.168.3) ; extra == "dev"
Requires-Dist: hypothesis (>=6.168.3) ; extra == "test"
Requires-Dist: ipykernel (>=7.3.0) ; extra == "dev"
Requires-Dist: ipywidgets (>=8.1.5,<9.0.0) ; extra == "all"
Requires-Dist: ipywidgets (>=8.1.5,<9.0.0) ; extra == "notebook"
Requires-Dist: langchain-core (>=0.3.78,<0.4.0)
Requires-Dist: langchain-deepseek (>=0.1.4,<1.0.0) ; extra == "all"
Requires-Dist: langchain-deepseek (>=0.1.4,<1.0.0) ; extra == "deepseek"
Requires-Dist: langchain-deepseek (>=0.1.4,<1.0.0) ; extra == "models"
Requires-Dist: langchain-google-genai (>=2.1.12,<3.0.0) ; extra == "all"
Requires-Dist: langchain-google-genai (>=2.1.12,<3.0.0) ; extra == "google"
Requires-Dist: langchain-google-genai (>=2.1.12,<3.0.0) ; extra == "models"
Requires-Dist: langchain-huggingface (>=0.3.1,<0.4.0) ; extra == "all"
Requires-Dist: langchain-huggingface (>=0.3.1,<0.4.0) ; extra == "huggingface"
Requires-Dist: langchain-huggingface (>=0.3.1,<0.4.0) ; extra == "models"
Requires-Dist: langchain-ollama (>=0.3.10,<0.4.0) ; extra == "all"
Requires-Dist: langchain-ollama (>=0.3.10,<0.4.0) ; extra == "models"
Requires-Dist: langchain-ollama (>=0.3.10,<0.4.0) ; extra == "ollama"
Requires-Dist: langchain-openai (>=0.3.35,<0.4.0) ; extra == "all"
Requires-Dist: langchain-openai (>=0.3.35,<0.4.0) ; extra == "configurator"
Requires-Dist: langchain-openai (>=0.3.35,<0.4.0) ; extra == "models"
Requires-Dist: langchain-openai (>=0.3.35,<0.4.0) ; extra == "openai"
Requires-Dist: mkdocs (>=1.6.1,<2.0.0) ; extra == "dev"
Requires-Dist: mkdocs (>=1.6.1,<2.0.0) ; extra == "docs"
Requires-Dist: mkdocs-material (>=9.7.7) ; extra == "dev"
Requires-Dist: mkdocs-material (>=9.7.7) ; extra == "docs"
Requires-Dist: mkdocstrings[python] (>=1.0.6) ; extra == "dev"
Requires-Dist: mkdocstrings[python] (>=1.0.6) ; extra == "docs"
Requires-Dist: mypy (>=2.3.1) ; extra == "dev"
Requires-Dist: nbclient (>=0.11.0) ; extra == "dev"
Requires-Dist: nbformat (>=5.11.1) ; extra == "dev"
Requires-Dist: numpy (>=1.26.0,<3.0.0)
Requires-Dist: openai (>=1.104.2,<2.0.0) ; extra == "all"
Requires-Dist: openai (>=1.104.2,<2.0.0) ; extra == "configurator"
Requires-Dist: pandas (>=2.1.1,<3.0.0)
Requires-Dist: pandas-stubs ; extra == "dev"
Requires-Dist: poethepoet (>=0.48.0) ; extra == "dev"
Requires-Dist: pre-commit (>=4.6.2) ; extra == "dev"
Requires-Dist: pydantic (>=2.9.0,<3.0.0)
Requires-Dist: pypdf (>=5.1.0,<7.0.0) ; extra == "all"
Requires-Dist: pypdf (>=5.1.0,<7.0.0) ; extra == "configurator"
Requires-Dist: pytest (>=9.1.1,<10.0.0) ; extra == "dev"
Requires-Dist: pytest (>=9.1.1,<10.0.0) ; extra == "test"
Requires-Dist: pytest-cov (>=7.1.0) ; extra == "dev"
Requires-Dist: pytest-cov (>=7.1.0) ; extra == "test"
Requires-Dist: pytest-timeout (>=2.4.0) ; extra == "dev"
Requires-Dist: pytest-timeout (>=2.4.0) ; extra == "test"
Requires-Dist: python-semantic-release (>=10.7.0) ; extra == "dev"
Requires-Dist: ruff (>=0.16.10,<0.17.0) ; extra == "dev"
Requires-Dist: sentencepiece (>=0.2.0) ; extra == "all"
Requires-Dist: sentencepiece (>=0.2.0) ; extra == "huggingface"
Requires-Dist: sentencepiece (>=0.2.0) ; extra == "models"
Requires-Dist: streamlit (>=1.40.2,<2.0.0) ; extra == "all"
Requires-Dist: streamlit (>=1.40.2,<2.0.0) ; extra == "configurator"
Requires-Dist: tabulate (>=0.9.0,<1.0.0)
Requires-Dist: torch (>=2.2.0,<3.0.0) ; extra == "all"
Requires-Dist: torch (>=2.2.0,<3.0.0) ; extra == "huggingface"
Requires-Dist: torch (>=2.2.0,<3.0.0) ; extra == "models"
Requires-Dist: tqdm (>=4.64.0,<5.0.0)
Requires-Dist: transformers (>=4.40.0,<5.0.0) ; extra == "all"
Requires-Dist: transformers (>=4.40.0,<5.0.0) ; extra == "huggingface"
Requires-Dist: transformers (>=4.40.0,<5.0.0) ; extra == "models"
Requires-Dist: types-aiofiles ; extra == "dev"
Requires-Dist: types-tabulate ; extra == "dev"
Requires-Dist: types-tqdm (>=4.70.0.20260906) ; extra == "dev"
Project-URL: Changelog, https://github.com/julianschelb/rupsycho/blob/main/CHANGELOG.md
Project-URL: Documentation, https://julianschelb.github.io/rupsycho/
Project-URL: Homepage, https://github.com/julianschelb/rupsycho
Project-URL: Issues, https://github.com/julianschelb/rupsycho/issues
Project-URL: Paper, https://arxiv.org/abs/2503.10229
Project-URL: Repository, https://github.com/julianschelb/rupsycho
Description-Content-Type: text/markdown

# R.U.Psycho

[![CI](https://github.com/julianschelb/rupsycho/actions/workflows/ci.yml/badge.svg)](https://github.com/julianschelb/rupsycho/actions/workflows/ci.yml)
[![Docs](https://github.com/julianschelb/rupsycho/actions/workflows/docs.yml/badge.svg)](https://julianschelb.github.io/rupsycho/)
[![Python](https://img.shields.io/badge/python-3.10%20%7C%203.11%20%7C%203.12%20%7C%203.13%20%7C%203.14-blue.svg)](https://github.com/julianschelb/rupsycho/blob/main/pyproject.toml)
[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://github.com/julianschelb/rupsycho/blob/main/LICENSE)
[![arXiv](https://img.shields.io/badge/arXiv-2503.10229-b31b1b.svg)](https://arxiv.org/abs/2503.10229)

**R.U.Psycho** (*Robust Unified Psychometric Testing of Language Models*) is a framework for
designing and running **robust and reproducible psychometric experiments on generative language
models**, with limited coding expertise required.

Documentation: <https://julianschelb.github.io/rupsycho/>

Paper: [*R.U.Psycho? Robust Unified Psychometric Testing of Language Models*](https://arxiv.org/abs/2503.10229)
(Schelb, Borin, Garcia & Spitz, 2025)

## Why?

Instabilities in model outputs, sensitivity to prompt design and generation parameters, and the
sheer number of model versions make psychometric studies of language models hard to reproduce.
R.U.Psycho turns a whole study into **one declarative configuration**: the questionnaire, the
personas the model answers as, the models and their parameters, the prompt template and the
random seeds. The package runs every *model × seed × persona × item* combination, turns the
free-text answers into scorable responses and computes item and scale scores.

## Features

- **Declarative experiments** in a single, shareable JSON file (secrets are never exported)
- **Seeds that work**: every back-end receives the seed in the way it supports (see
  [Reproducibility](https://julianschelb.github.io/rupsycho/tutorials/reproducibility/))
- **Many back-ends:** local & remote Hugging Face, Ollama, OpenAI, Google, DeepSeek, any LangChain runnable
- **Robust runs:** a `RunSummary` instead of silent failures, an error policy, opt-in concurrency for
  API models with a deterministic result order, re-runnable experiments
- **Post-processing and scoring:** cleaners, refusal / "as an AI" validators, rule- and model-based
  judges, then weights, reverse-keyed items and per-trait scale scores
- **Command line:** `rupsycho run | validate | prompt | postprocess | examples | configurator`
- **Configurator app** with LLM-assisted questionnaire import from PDF
- **Light core:** `import rupsycho` takes milliseconds; model back-ends are optional extras

## Installation

```bash
pip install git+https://github.com/julianschelb/rupsycho.git          # core: configs, run loop, parsers, scoring
pip install "rupsycho[huggingface] @ git+https://github.com/julianschelb/rupsycho.git"   # + local Hugging Face models
```

Extras: `huggingface` (PyTorch, Transformers), `openai`, `ollama`, `google`, `deepseek`, `models`
(all back-ends), `configurator` (Streamlit app), `notebook`, `quantization`, `all`. Using a
back-end whose extra is missing raises an error that names the extra to install.

Requires Python 3.10 or newer (tested on 3.10 – 3.14).

## Quick start

```python
import rupsycho as rup

# A bundled example: five Big Five items answered by two personas
experiment = rup.load_example_experiment("bfi", seeds=[1, 2, 3])
experiment.print_assembled_prompt(item_idx=1, persona_idx=0)   # exactly what the model sees

summary = experiment.run()                      # needs the huggingface extra for the example model
print(summary)                                  # e.g. "30 model calls in 41.2s"
answers = experiment.get_answers_as_dataframe() # one row per model x seed x persona x item
```

Use your own model instead of the example's:

```python
from langchain_openai import ChatOpenAI          # pip install "rupsycho[openai]"

experiment = rup.load_example_experiment("bfi", models={}, seeds=[1, 2, 3])
experiment.add_model(ChatOpenAI(model="gpt-4o-mini"), identifier="gpt-4o-mini")
experiment.run(max_concurrency=8)                # API calls run in parallel, order is preserved
```

From the shell:

```bash
rupsycho examples copy bfi bfi.json
rupsycho validate bfi.json
rupsycho run bfi.json -o results.csv --seeds 1 2 3
rupsycho postprocess bfi.json results.csv -o processed.csv
```

Then score the judged answers:

```python
import pandas as pd
from rupsycho import scoring

processed = pd.read_csv("processed.csv")
scored = scoring.score_answers(processed, experiment)
print(scoring.scale_scores(scored))              # mean score per model, persona, seed and trait
```

See the [Getting Started guide](https://julianschelb.github.io/rupsycho/getting-started/), the
[tutorials](https://julianschelb.github.io/rupsycho/tutorials/running-experiments/) and the
notebooks in [`examples/`](https://github.com/julianschelb/rupsycho/tree/main/examples).

## Development

```bash
pip install -e ".[dev,models,configurator]"
pre-commit install --hook-type pre-commit --hook-type commit-msg

poe check              # lint + format check + mypy + offline tests
poe docs               # serve the docs locally
poe test-integration   # tests that download real models
```

See [CONTRIBUTING.md](https://github.com/julianschelb/rupsycho/blob/main/CONTRIBUTING.md) and the
[development guide](https://julianschelb.github.io/rupsycho/development/).

## Citation

If you use R.U.Psycho, please cite the paper (see also [CITATION.cff](https://github.com/julianschelb/rupsycho/blob/main/CITATION.cff)):

```bibtex
@misc{schelb2025rupsycho,
  title         = {R.U.Psycho? Robust Unified Psychometric Testing of Language Models},
  author        = {Schelb, Julian and Borin, Orr and Garcia, David and Spitz, Andreas},
  year          = {2025},
  eprint        = {2503.10229},
  archivePrefix = {arXiv},
  url           = {https://arxiv.org/abs/2503.10229}
}
```

## License

[MIT](https://github.com/julianschelb/rupsycho/blob/main/LICENSE)

