Metadata-Version: 2.5
Name: ssebench
Version: 1.1.0
Summary: Benchmark AI coding agents on real security vulnerabilities
Project-URL: Homepage, https://github.com/42-b3yond-6ug/ssebench
Project-URL: Documentation, https://github.com/42-b3yond-6ug/ssebench/tree/main/docs
Project-URL: Repository, https://github.com/42-b3yond-6ug/ssebench
Project-URL: Issues, https://github.com/42-b3yond-6ug/ssebench/issues
Author: SSEBench Team
License-Expression: Apache-2.0
License-File: LICENSE
License-File: NOTICE
Keywords: agents,benchmark,docker,llm,patch,security,vulnerability
Classifier: Development Status :: 4 - Beta
Classifier: Environment :: Console
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: Operating System :: MacOS
Classifier: Operating System :: POSIX :: Linux
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: Implementation :: CPython
Classifier: Topic :: Security
Classifier: Topic :: Software Development :: Testing
Requires-Python: >=3.12
Requires-Dist: httpx>=0.28.1
Requires-Dist: jinja2>=3.1.6
Requires-Dist: jsonschema>=4.26.0
Requires-Dist: pydantic>=2.11.9
Requires-Dist: python-dotenv>=1.1.0
Requires-Dist: pyyaml>=6.0.3
Requires-Dist: requests>=2.32.5
Provides-Extra: kubernetes
Requires-Dist: kubernetes>=31; extra == 'kubernetes'
Description-Content-Type: text/markdown

# ssebench

The command-line tool of [SSEBench](https://github.com/42-b3yond-6ug/ssebench),
a benchmark that measures how well AI coding agents fix real security
vulnerabilities.

Every task is a publicly disclosed bug in an open-source C, Go or Rust project,
paired with its upstream fix. `ssebench` builds a Docker image for a task, runs
one agent with one model on it, and grades the patch the agent leaves behind:
does the project still build, does the proof of concept stop reproducing, and
do the project's tests pass. The `pilot` dataset has 55 tasks.

## Install and run

`ssebench` needs Docker with the buildx and Compose plugins, and Python 3.12 or
newer. It runs without a clone of the repository: the wheel carries the agent
definitions, the model list, the Compose file and the pilot manifest, and the
task images are pulled from the registry.

```sh
mkdir ssebench-work && cd ssebench-work
uvx ssebench init          # writes .env with generated secrets, models/ and results/
uvx ssebench doctor        # checks Docker, disk space and .env
uvx ssebench tasks list    # the 55 tasks of the pilot dataset

# The reference agent applies the task's known fix, so it needs no API key.
uvx ssebench run --task gjson-196-bf4efcb --agent reference

# A real agent needs the key of its model provider in .env.
uvx ssebench run --task gjson-196-bf4efcb --agent claude-code --model claude-sonnet-4-6
```

Install it with `pip install ssebench` or `uv tool install ssebench` instead of
`uvx` to keep the command. Run it from the directory that `ssebench init` set
up: that directory holds `.env` and `models/`, and `results/` is written there.

The first run of a task pulls its case image and builds the tool and agent
layers on top of it, which takes a few minutes. The results are in
`results/<task>/<model>/<agent>/<run-id>/`, a directory of its own for each run.

## Versions

SSEBench components are released together under one version. `ssebench` pulls
the `runtime` image with the same version, and the SDK inside the task
container, `ssebench-sdk`, has the same version too.

## More

- [Documentation](https://github.com/42-b3yond-6ug/ssebench/tree/main/docs)
- [Source and issues](https://github.com/42-b3yond-6ug/ssebench)

Licensed under the Apache License 2.0.
