Metadata-Version: 2.4
Name: inspect_evals
Version: 0.22.0
Summary: Collection of large language model evaluations
Author: UK AI Security Institute
License-Expression: MIT
Project-URL: Source Code, https://github.com/UKGovernmentBEIS/inspect_evals
Project-URL: Issue Tracker, https://github.com/UKGovernmentBEIS/inspect_evals/issues
Project-URL: Documentation, https://ukgovernmentbeis.github.io/inspect_evals/
Project-URL: Release Notes, https://github.com/UKGovernmentBEIS/inspect_evals/releases/
Project-URL: Changelog, https://github.com/UKGovernmentBEIS/inspect_evals/blob/main/CHANGELOG.md
Classifier: Development Status :: 4 - Beta
Classifier: Environment :: Console
Classifier: Intended Audience :: Science/Research
Classifier: Intended Audience :: Developers
Classifier: Natural Language :: English
Classifier: Programming Language :: Python :: 3
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Typing :: Typed
Classifier: Operating System :: OS Independent
Requires-Python: >=3.11
Description-Content-Type: text/markdown
License-File: LICENSE
License-File: NOTICE
Requires-Dist: backoff>=2.2.0
Requires-Dist: datasets>=4.8.5
Requires-Dist: huggingface_hub>=1.2.0
Requires-Dist: hf_xet
Requires-Dist: httpx>=0.27.0
Requires-Dist: inspect_ai>=0.3.261
Requires-Dist: jinja2
Requires-Dist: numpy>=1.26.0
Requires-Dist: pillow>=11.3.0
Requires-Dist: pydantic>=2.10.0
Requires-Dist: pyyaml>=5.1.0
Requires-Dist: requests>=2.32.0
Requires-Dist: tiktoken>=0.11.0
Requires-Dist: toml>=0.10.2
Provides-Extra: register-tooling
Requires-Dist: pypdf>=5.1.0; extra == "register-tooling"
Requires-Dist: PyGithub>=2.0; extra == "register-tooling"
Provides-Extra: gdown
Requires-Dist: gdown>=6; extra == "gdown"
Provides-Extra: b3
Requires-Dist: openai; extra == "b3"
Requires-Dist: rouge_score; extra == "b3"
Requires-Dist: tenacity; extra == "b3"
Requires-Dist: click; extra == "b3"
Requires-Dist: python-dotenv; extra == "b3"
Provides-Extra: agentdojo
Requires-Dist: pydantic[email]; extra == "agentdojo"
Requires-Dist: deepdiff; extra == "agentdojo"
Provides-Extra: bfcl
Requires-Dist: mpmath; extra == "bfcl"
Provides-Extra: bfcl-v4
Requires-Dist: html2text>=2024.2.26; extra == "bfcl-v4"
Requires-Dist: google-search-results>=2.4.2; extra == "bfcl-v4"
Requires-Dist: beautifulsoup4>=4.12.0; extra == "bfcl-v4"
Requires-Dist: overrides>=7.7.0; extra == "bfcl-v4"
Requires-Dist: rank-bm25>=0.2.2; extra == "bfcl-v4"
Requires-Dist: sentence-transformers>=3.0.0; extra == "bfcl-v4"
Requires-Dist: faiss-cpu>=1.8.0; extra == "bfcl-v4"
Provides-Extra: swe-bench
Requires-Dist: swebench>=3.0.15; extra == "swe-bench"
Requires-Dist: docker; extra == "swe-bench"
Requires-Dist: jsonlines; extra == "swe-bench"
Provides-Extra: swe-lancer
Requires-Dist: docker; extra == "swe-lancer"
Requires-Dist: types-docker; extra == "swe-lancer"
Provides-Extra: math
Requires-Dist: sympy; extra == "math"
Requires-Dist: antlr4-python3-runtime~=4.11.0; extra == "math"
Provides-Extra: worldsense
Requires-Dist: pandas; extra == "worldsense"
Provides-Extra: mind2web
Requires-Dist: beautifulsoup4; extra == "mind2web"
Requires-Dist: types-beautifulsoup4; extra == "mind2web"
Requires-Dist: lxml; extra == "mind2web"
Requires-Dist: lxml-stubs; extra == "mind2web"
Provides-Extra: sevenllm
Requires-Dist: jieba==0.42.1; extra == "sevenllm"
Requires-Dist: sentence_transformers>=5.1.1; extra == "sevenllm"
Requires-Dist: rouge==1.0.1; extra == "sevenllm"
Requires-Dist: tf-keras; python_version < "3.13" and extra == "sevenllm"
Provides-Extra: scicode
Requires-Dist: inspect_evals[gdown]; extra == "scicode"
Requires-Dist: h5py; extra == "scicode"
Requires-Dist: scipy; extra == "scicode"
Requires-Dist: sympy; extra == "scicode"
Provides-Extra: ifeval
Requires-Dist: instruction_following_eval; extra == "ifeval"
Requires-Dist: langdetect; extra == "ifeval"
Provides-Extra: anima
Requires-Dist: matplotlib; extra == "anima"
Provides-Extra: medqa
Requires-Dist: bioc; extra == "medqa"
Provides-Extra: niah
Requires-Dist: pandas; extra == "niah"
Provides-Extra: core-bench
Requires-Dist: scipy; extra == "core-bench"
Provides-Extra: healthbench
Requires-Dist: scikit-learn; extra == "healthbench"
Provides-Extra: personality
Requires-Dist: huggingface-hub; extra == "personality"
Provides-Extra: sciknoweval
Requires-Dist: nltk; extra == "sciknoweval"
Requires-Dist: rouge_score; extra == "sciknoweval"
Requires-Dist: rdkit; extra == "sciknoweval"
Requires-Dist: rdchiral; extra == "sciknoweval"
Requires-Dist: inspect_evals[gdown]; extra == "sciknoweval"
Requires-Dist: gensim; python_version < "3.13" and extra == "sciknoweval"
Requires-Dist: scipy; extra == "sciknoweval"
Provides-Extra: gdm-stealth
Requires-Dist: tabulate; extra == "gdm-stealth"
Requires-Dist: scipy; extra == "gdm-stealth"
Requires-Dist: immutabledict; extra == "gdm-stealth"
Requires-Dist: pandas; extra == "gdm-stealth"
Requires-Dist: python-dateutil; extra == "gdm-stealth"
Provides-Extra: cybench
Requires-Dist: inspect-cyber==0.1.0; extra == "cybench"
Provides-Extra: cybergym
Requires-Dist: inspect-cyber>=0.1.0; extra == "cybergym"
Provides-Extra: makemesay
Requires-Dist: nltk; extra == "makemesay"
Provides-Extra: gaia
Requires-Dist: filelock; extra == "gaia"
Provides-Extra: osworld
Requires-Dist: filelock; extra == "osworld"
Provides-Extra: cti-realm
Requires-Dist: filelock; extra == "cti-realm"
Requires-Dist: httpx; extra == "cti-realm"
Requires-Dist: psutil; extra == "cti-realm"
Requires-Dist: requests; extra == "cti-realm"
Provides-Extra: vimgolf
Requires-Dist: vimgolf==0.5.1; extra == "vimgolf"
Provides-Extra: vimgolf-challenges
Requires-Dist: vimgolf==0.5.1; extra == "vimgolf-challenges"
Provides-Extra: usaco
Requires-Dist: inspect_evals[gdown]; extra == "usaco"
Provides-Extra: ifevalcode
Requires-Dist: tree-sitter; extra == "ifevalcode"
Requires-Dist: tree-sitter-cpp; extra == "ifevalcode"
Provides-Extra: gdpval
Requires-Dist: huggingface_hub[cli]; extra == "gdpval"
Provides-Extra: agentic-misalignment
Requires-Dist: bs4; extra == "agentic-misalignment"
Provides-Extra: cje
Requires-Dist: cje-eval>=0.6.0; extra == "cje"
Provides-Extra: theagentcompany
Requires-Dist: inspect_cyber; extra == "theagentcompany"
Requires-Dist: pandas; extra == "theagentcompany"
Requires-Dist: openpyxl; extra == "theagentcompany"
Requires-Dist: odfpy>=1.4.1; extra == "theagentcompany"
Requires-Dist: pypdf; extra == "theagentcompany"
Requires-Dist: python-pptx; extra == "theagentcompany"
Requires-Dist: scikit-learn; extra == "theagentcompany"
Provides-Extra: gdm-capabilities
Requires-Dist: google-genai>=1.56.0; extra == "gdm-capabilities"
Requires-Dist: rich; extra == "gdm-capabilities"
Requires-Dist: python-dateutil; extra == "gdm-capabilities"
Provides-Extra: gdm-self-proliferation
Requires-Dist: rich; extra == "gdm-self-proliferation"
Provides-Extra: scbench
Requires-Dist: inspect-swe; extra == "scbench"
Requires-Dist: pip; extra == "scbench"
Provides-Extra: paperbench
Requires-Dist: drain3; extra == "paperbench"
Provides-Extra: cyberseceval-4
Requires-Dist: pdf2image; extra == "cyberseceval-4"
Requires-Dist: platformdirs; extra == "cyberseceval-4"
Requires-Dist: pypdf; extra == "cyberseceval-4"
Requires-Dist: pyyaml; extra == "cyberseceval-4"
Requires-Dist: sacrebleu; extra == "cyberseceval-4"
Requires-Dist: semgrep>1.68; extra == "cyberseceval-4"
Dynamic: license-file

<!-- markdownlint-configure-file { "no-inline-html": { "allowed_elements": ["h1", "img", "details", "summary"] } } -->

<h1>
  <img src="docs/images/inspect-evals-wordmark.svg#gh-light-mode-only" alt="Inspect Evals" width="320">
  <img src="docs/images/inspect-evals-wordmark-dark.svg#gh-dark-mode-only" alt="Inspect Evals" width="320">
</h1>

A library of evaluations built using [Inspect AI](https://inspect.aisi.org.uk/).

## [Explore evaluations and read the docs →](https://ukgovernmentbeis.github.io/inspect_evals/)

## Quick start

1. **[Run an eval](https://ukgovernmentbeis.github.io/inspect_evals/).** Choose an evaluation in the docs and follow the Usage section on its page. For a first run from this repository, see [Getting started](#getting-started) below.
2. **[Add an eval to the Inspect Evals Register](https://ukgovernmentbeis.github.io/inspect_evals/register/).** Share your evaluation by adding a listing that points to your implementation and documentation.

> [!IMPORTANT]
> We've updated our contribution policy to only accept PRs from [pre-approved contributors](APPROVED_CONTRIBUTORS.md); if you identify an issue with an evaluation, [raise an issue](https://github.com/UKGovernmentBEIS/inspect_evals/issues/new/choose) and upload `.eval` logs to our [log uploader](https://logfile-upload.generality.workers.dev/) which demonstrate the problem. We will review the logs to better understand the severity and cause of the issue.

## About the project

Inspect Evals is maintained by **[Generality Labs](https://generality.org/)**, a London-based nonprofit which builds tools for evaluating risks and mitigations in frontier AI. Inspect Evals was founded in 2024, with contributions from the [UK AI Security Institute](https://www.aisi.gov.uk/), [Arcadia Impact](https://www.arcadiaimpact.org/), and the [Vector Institute](https://vectorinstitute.ai/).

For inquiries, suggestions, or expressions of interest, [get in touch](https://docs.google.com/forms/d/e/1FAIpQLSeOT_nSXvc_GZSo3uRqFlZlgGEGmOAh7bm4yFuB34ZzZjxk_g/viewform?usp=dialog). See the [project maintainers](MAINTAINERS.md).

## Getting started

<details>
<summary>Installation, running evaluations, and hardware requirements</summary>

The recommended version of Python for Inspect Evals is 3.11 or 3.12. You should be able to run all evals on these versions and also develop the codebase without any issues. You can install and pin a specific Python version by running:

```bash
uv python pin 3.11
```

As for Python 3.13, you should be able to run all evals except `sciknoweval` (its dependency is `gensim` which currently does not support 3.13+). Development should work under 3.13, however it's relatively untested — if you run into issues, let us know.

When it comes Python 3.14, at the time of writing this, many packages have yet to release versions for 3.14, so it's unsupported. The major one used by some Inspect Evals is `torch`. If you find running `uv sync` succeeding on 3.14, let us know and we'll remove this paragraph.

Below, you can see a workflow for a typical eval. Some of the evaluations require additional dependencies or installation steps. If your eval needs extra dependencies, instructions for installing in the README file in the eval's subdirectory.

<!-- Usage: Automatically Generated -->

## Usage

### Installation

There are two ways of using Inspect Evals, from pypi as a dependency of your own project and as a standalone checked out GitHub repository.

If you are using it from pypi, install the package and its dependencies via:

```bash
pip install inspect-evals
```

If you are using Inspect Evals in its repository, start by installing the necessary dependencies with:

```bash
uv sync
```

### Running evaluations

Now you can start evaluating models. For simplicity's sake, this section assumes you are using Inspect Evals from the standalone repo. If that's not the case and you are not using `uv` to manage dependencies in your own project, you can use the same commands with `uv run` dropped.

```bash
uv run inspect eval inspect_evals/arc_easy --model openai/gpt-5-nano
uv run inspect eval inspect_evals/arc_challenge --model openai/gpt-5-nano
```

To run multiple tasks simultaneously use `inspect eval-set`:

```bash
uv run inspect eval-set inspect_evals/arc_easy inspect_evals/arc_challenge
```

You can also import tasks as normal Python objects and run them from python:

```python
from inspect_ai import eval, eval_set
from inspect_evals.arc import arc_easy, arc_challenge
eval(arc_easy)
eval_set([arc_easy, arc_challenge], log_dir='logs-run-42')
```

After running evaluations, you can view their logs using the `inspect view` command:

```bash
uv run inspect view
```

For VS Code, you can also download [Inspect AI extension for viewing logs](https://inspect.ai-safety-institute.org.uk/log-viewer.html).

If you don't want to specify the `--model` each time you run an evaluation, create a `.env` configuration file in your working directory that defines the `INSPECT_EVAL_MODEL` environment variable along with your API key. For example:

```bash
INSPECT_EVAL_MODEL=anthropic/claude-opus-4-1-20250805
ANTHROPIC_API_KEY=<anthropic-api-key>
```

<!-- /Usage: Automatically Generated -->

Inspect supports many model providers including OpenAI, Anthropic, Google, Mistral, Azure AI, AWS Bedrock, Together AI, Groq, Hugging Face, vLLM, Ollama, and more. See the [Model Providers](https://inspect.ai-safety-institute.org.uk/models.html) documentation for additional details.

You might also be able to use a newer version of pip (25.1+) to install the project via `pip install --group dev .` or `pip install --group dev '.[swe_bench]'`. However this is not officially supported.

## Documentation

For details on building the documentation, see [the documentation guide](docs/documentation.md).

For information on running tests and CI toggles, see the Technical Contribution Guide in [CONTRIBUTING.md](CONTRIBUTING.md).

## Hardware recommendations

### Disk

We recommend having at least 35 GB of free disk space for Inspect Evals: the full installation takes about 10 GB and you'll also need some space for uv cache and datasets cache (most are small, but some take 13 GB such as MMIU).

Running some evals (e.g., CyBench, GDM capabilities evals) may require extra space beyond this because they pull Docker images. We recommend having at least 65 GB of extra space for running evals that have Dockerfiles in their file tree (though you might get away with less space) on top of the 35 GB suggestion above.

In total, you should be comfortable running evals with 100 GB of free space. If you end up running of out space while having 100+ GB of free space available, please let us know — this might be a bug.

### Cache location

Datasets and other large assets are cached under the platform cache directory (`~/.cache/inspect_evals` on Linux, `~/Library/Caches/inspect_evals` on macOS). Set `INSPECT_EVALS_CACHE_DIR` to put them somewhere else:

```bash
export INSPECT_EVALS_CACHE_DIR=/data/inspect-evals-cache
```

Use it when the default location is not writable (read-only container filesystems, images without a writable `HOME`), when the cache should live on a larger volume, or to stage assets for a machine with no network access: populate the directory on a connected machine running the same Inspect Evals version, copy it across, and point the variable at it there.

The variable is read when `inspect_evals` is first imported, so set it in the shell or in a `.env` file rather than from within Python.

For the same reason it must be an absolute path, or start with `~` for a path under your home directory. A relative path is rejected, because it would point somewhere different depending on where the eval was started from.

### RAM

The amount of memory needed for an eval varies significantly with the eval. You'll be able to run most evals with only 0.5 GB of free RAM. However, some evals with larger datasets require 2-3 GB or more. And some evals that use Docker (e.g., some GDM capabilities evals) require up to 32 GB of RAM.

## Harbor Framework Evaluations

For running evaluations from the Harbor Framework (e.g. Terminal-Bench 2.0, SWE-Bench Pro), use the [Inspect Harbor](https://github.com/meridianlabs-ai/inspect_harbor) package, which provides an interface to run Harbor tasks using Inspect AI.

</details>
