Metadata-Version: 2.4
Name: popfidelity
Version: 0.1.0
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Science/Research
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Rust
Classifier: Topic :: Scientific/Engineering
Classifier: Topic :: Sociology
Requires-Dist: numpy>=1.26,<3
Requires-Dist: pandas>=2.2,<4
Requires-Dist: pyyaml>=6,<7
Requires-Dist: openai>=3,<4 ; extra == 'api'
Requires-Dist: anthropic>=1,<2 ; extra == 'api'
Requires-Dist: google-genai>=2,<3 ; extra == 'api'
Requires-Dist: python-dotenv>=1,<2 ; extra == 'api'
Requires-Dist: torch>=2.4 ; extra == 'local'
Requires-Dist: transformers>=5,<6 ; extra == 'local'
Requires-Dist: accelerate>=1 ; extra == 'local'
Requires-Dist: pyarrow>=17 ; extra == 'parquet'
Requires-Dist: matplotlib>=3.8,<4 ; extra == 'plot'
Provides-Extra: api
Provides-Extra: local
Provides-Extra: parquet
Provides-Extra: plot
License-File: LICENSE
License-File: NOTICE
Summary: Population fidelity and distributional alignment measures for LLM-simulated survey populations
Keywords: llm,survey,silicon sampling,population fidelity,evaluation
Author-email: "Neemias B. da Silva" <neemiasbsilva@gmail.com>
License-Expression: Apache-2.0
Requires-Python: >=3.11
Description-Content-Type: text/markdown; charset=UTF-8; variant=GFM
Project-URL: Homepage, https://neemiasbsilva.github.io/popfidelity/
Project-URL: Issues, https://github.com/neemiasbsilva/popfidelity/issues
Project-URL: Paper, https://arxiv.org/abs/2609.36253
Project-URL: Repository, https://github.com/neemiasbsilva/popfidelity

# popfidelity

![Python](https://img.shields.io/badge/Python-3.11%2B-3776AB?logo=python&logoColor=white)
![R](https://img.shields.io/badge/R-4.2%2B-276DC3?logo=r&logoColor=white)
![Rust](https://img.shields.io/badge/Rust_core-1.85%2B-B7410E?logo=rust&logoColor=white)
![uv](https://img.shields.io/badge/uv-locked-DE5FE9?logo=uv&logoColor=white)
![Ruff](https://img.shields.io/badge/Ruff-passing-D7FF64?logo=ruff&logoColor=black)
![mypy](https://img.shields.io/badge/mypy-strict-2A6DB2)
![tests](https://img.shields.io/badge/tests-165-4c1)
![License](https://img.shields.io/badge/License-Apache_2.0-blue)
[![arXiv](https://img.shields.io/badge/arXiv-2609.36253-B31B1B)](https://arxiv.org/abs/2609.36253)

A language model that answers survey questions can match the average answer of a
population and still erase the differences between the people in it.
**popfidelity** measures how well a model's answers represent a human population,
group by group. One Rust core gives a Python package, an R package and a command
line the same numbers.

## What it measures

The **Population Fidelity Score** (PFS; da Silva et al., 2026) compares the survey
and the model on cells, the subpopulations a survey reports, and checks three
conditions:

| Component | Question | Score |
| --- | --- | --- |
| Accuracy | Is each cell's answer distribution close to the survey's? | one minus the mean normalised earth mover's distance |
| Adaptability | Do the model's cells differ from one another as much as the survey's do? | `min(A, 1/A)` of the ratio of median pairwise distances |
| Structure | Do the cells that differ in the survey also differ in the model? | Spearman correlation of the pairwise distances, clipped at zero |

PFS is their geometric mean, so a model fails as soon as one condition fails. It
is reported for the pooled population and every demographic subgroup, for paired
comparisons of a tuned and a base model, and across replicate runs and questions,
with the paper's nine sensitivity specifications.

Beside PFS sits a registry of 126 measures from the literature on synthetic survey
samples, each with the work it comes from and the work that applied it to language
models; 66 ship in this release:

| Family | In 0.1 | Examples |
| --- | --- | --- |
| Distances | 23 | nEMD, Wasserstein, total variation, Kullback-Leibler, Jensen-Shannon, Hellinger, Kolmogorov-Smirnov, energy distance, MMD |
| Alignment scores | 11 | OpinionQA representativeness, steerability and consistency; GlobalOpinionQA similarity; Meister; SimBench; SubPOP with its noise floor; WorldValuesBench curve |
| PFS and its parts | 10 | components, centre alignment, binding component, variants |
| Dispersion | 8 | SD ratio, normalised variance, gap inflation, entropy and collision ratios, support and modal collapse, stereotyping |
| Reliability | 7 | ICC, Krippendorff's alpha, Fleiss' kappa, Cronbach's alpha, Spearman-Brown, noise-to-signal |
| Structure, inference, artefacts | 7 | Mantel test, quality bands, cell bootstrap, permutation nulls, pooled t intervals, mismatch rate |

The [measure catalogue](https://neemiasbsilva.github.io/popfidelity/measures/)
lists all 126, including those planned for later releases, and 17 benchmarks.

## Install

```bash
pip install popfidelity
```

```r
remotes::install_github("neemiasbsilva/popfidelity", subdir = "r")
```

```bash
cargo add popfidelity
```

PyPI has wheels for Linux, macOS and Windows on Python 3.11 and newer, so pip needs
no Rust toolchain there. The R package compiles the Rust core, so it needs Cargo and
rustc 1.85 or newer, as does a Python install from source. The Python extras `api`
(OpenAI, Anthropic, Google clients), `local` (transformers) and `parquet` add
optional backends and formats.

## Quickstart

Python:

```python
import popfidelity as pf
from popfidelity.examples import toy
from popfidelity.io import read_cells

cells = toy()
scores = pf.score(cells["survey"], cells["M2"], pf.Variant(round_pairs=10))
print(scores["pfs"], scores["binding_component"])

questions = pf.read_questions("examples/ces/questions.yaml")
survey = read_cells("examples/ces/data/survey_cells.csv", questions)["ideo5"]
results = "examples/ces/results/ollama-qwen3-vl-2b/model_cells.csv"
model = read_cells(results, questions, mode="ntp")["ideo5"]
facets = pf.read_facets("examples/ces/facets.yaml").from_labels(survey.labels)
groups = pf.score_groups(survey, model, facets)
```

R:

```r
library(popfidelity)
toy <- example_toy()
pfs_score(toy$survey, toy$M2)$pfs

questions <- read_questions("examples/ces/questions.yaml")
survey <- read_cells("examples/ces/data/survey_cells.csv", questions)$ideo5
results <- "examples/ces/results/ollama-qwen3-vl-2b/model_cells.csv"
model <- read_cells(results, questions, mode = "ntp")$ideo5
groups <- pfs_score_groups(survey, model, read_facets("examples/ces/facets.yaml"))
```

Command line:

```bash
popfidelity elicit --config examples/ces/elicit/ollama.yaml --limit 8
popfidelity aggregate records --config examples/ces/elicit/ollama.yaml --out model_cells.csv
popfidelity score --survey examples/ces/data/survey_cells.csv --model model_cells.csv \
    --questions examples/ces/questions.yaml --facets examples/ces/facets.yaml --out scores
popfidelity report --groups scores/groups.csv
```

## Asking models

`popfidelity elicit` interviews a model once per respondent and question, with
the respondent's demographics as prior turns, and stores resumable JSON-lines
records. It reads next-token probabilities over the answer options where the
endpoint returns them and samples full answers otherwise.

| Backend | Endpoints |
| --- | --- |
| `openai_compatible` | OpenAI, vLLM, llama.cpp, Hugging Face router and TGI, OpenRouter, Together, Fireworks, DeepSeek, Gemini's OpenAI endpoint, or any compatible URL |
| `ollama` | Ollama's native API with raw prompts |
| `anthropic`, `google` | Claude and Gemini, full answers |
| `transformers` | a local Hugging Face model, exact next-token probabilities |

## Examples

| Example | Survey | Model answers |
| --- | --- | --- |
| [toy](https://github.com/neemiasbsilva/popfidelity/tree/main/examples/toy) | the paper's toy cells | models M1 to M3 |
| [ces](https://github.com/neemiasbsilva/popfidelity/tree/main/examples/ces) | Cooperative Election Study 2024, 96 cells | qwen3-vl:2b on Ollama, and configurations for six other backends |
| [twin2k](https://github.com/neemiasbsilva/popfidelity/tree/main/examples/twin2k) | Twin-2K-500 wave 4, 24 cells | GPT-4.1-mini digital twins and a human retest ceiling |
| [globalopinionqa](https://github.com/neemiasbsilva/popfidelity/tree/main/examples/globalopinionqa) | GlobalOpinionQA, 90 countries, fetched at run time | qwen3-vl:2b on Ollama, one country per interview |
| [machine-bias](https://github.com/neemiasbsilva/popfidelity/tree/main/examples/machine-bias) | World Values Survey cells | the paper's archived and fine-tuned models |

On the CES, the local model is close to every cell on average but does not order
the cells as the survey does, so structure binds and PFS stays at 0.30 or below.
On Twin-2K-500, people answering the same items again reach PFS 0.95 and the
digital twins 0.58 to 0.66.

## Reproducibility

- Python and R give the same tables: `./run.sh parity` runs both on shared inputs
  and compares 19 tables at 1e-12.
- The library reproduces the paper: `./run.sh crosscheck` recomputes its 9,548
  subgroup rows and 5,952 paired rows with differences of at most 4.4e-16.
- Bootstrap and permutation draws are seeded one by one, so results do not depend
  on the number of threads.
- Every reference in `docs/references.bib` is checked against Crossref and arXiv
  by `scripts/check_references.py`.

## Development

```bash
uv sync --locked --group dev --extra api --extra parquet
./run.sh test     # Rust and Python tests
./run.sh lint     # rustfmt, clippy, ruff, mypy strict, style check
./run.sh r        # R CMD check --as-cran
./run.sh parity   # Python against R
```

See [CONTRIBUTING.md](https://github.com/neemiasbsilva/popfidelity/blob/main/CONTRIBUTING.md)
for the conventions.

## Citation

```bibtex
TODO - Demo Paper

@misc{dasilva2026population,
  author = {da Silva, Neemias B. and Lukk, Martin and Sutani, Ali and Moturu, Abhishek and
            Yang, Harris and Silver, Daniel and Ratto, Matt and Silva, Thiago H.},
  title = {Population Fidelity: Evaluating Population Representativeness in {LLMs}},
  year = {2026},
  eprint = {2609.36253},
  archivePrefix = {arXiv},
  primaryClass = {cs.CL}
}
```

Please also cite the works behind the other measures you report; the registry
names them. The software citation is in
[CITATION.cff](https://github.com/neemiasbsilva/popfidelity/blob/main/CITATION.cff).

## Licence

Apache License 2.0; see
[LICENSE](https://github.com/neemiasbsilva/popfidelity/blob/main/LICENSE) and
[NOTICE](https://github.com/neemiasbsilva/popfidelity/blob/main/NOTICE).

