Metadata-Version: 2.5
Name: cua-speedrun
Version: 0.3.4
Summary: A benchmark and leaderboard for the speed of computer-use agents
Requires-Python: >=3.10
Requires-Dist: dbus-next>=0.2.3
Requires-Dist: fastapi>=0.115
Requires-Dist: huggingface-hub>=0.35.0
Requires-Dist: itsdangerous>=2.2
Requires-Dist: jinja2>=3.1
Requires-Dist: jsonschema>=4.20
Requires-Dist: modal>=1.5.5
Requires-Dist: numpy>=1.26
Requires-Dist: paramiko>=3.4
Requires-Dist: pillow>=10.0
Requires-Dist: pycryptodome>=3.20
Requires-Dist: python-dotenv
Requires-Dist: python-multipart>=0.0.9
Requires-Dist: pyyaml>=6.0
Requires-Dist: requests>=2.31
Requires-Dist: rich>=13.0
Requires-Dist: sqlalchemy>=2.0
Requires-Dist: uv==0.10.2
Requires-Dist: uvicorn[standard]>=0.30
Provides-Extra: all
Requires-Dist: beautifulsoup4; extra == 'all'
Requires-Dist: borb<3; extra == 'all'
Requires-Dist: chardet; extra == 'all'
Requires-Dist: cssselect; extra == 'all'
Requires-Dist: easyocr; extra == 'all'
Requires-Dist: fastdtw; extra == 'all'
Requires-Dist: formulas; extra == 'all'
Requires-Dist: gdown; extra == 'all'
Requires-Dist: gymnasium~=0.28.1; extra == 'all'
Requires-Dist: imagehash; extra == 'all'
Requires-Dist: librosa; extra == 'all'
Requires-Dist: lxml; extra == 'all'
Requires-Dist: mutagen; extra == 'all'
Requires-Dist: numpy; extra == 'all'
Requires-Dist: odfpy; extra == 'all'
Requires-Dist: opencv-python-headless; extra == 'all'
Requires-Dist: openpyxl; extra == 'all'
Requires-Dist: pandas; extra == 'all'
Requires-Dist: paramiko; extra == 'all'
Requires-Dist: pdfplumber; extra == 'all'
Requires-Dist: pillow; extra == 'all'
Requires-Dist: playwright; extra == 'all'
Requires-Dist: pyacoustid; extra == 'all'
Requires-Dist: pydrive; extra == 'all'
Requires-Dist: pygame; extra == 'all'
Requires-Dist: pymupdf; extra == 'all'
Requires-Dist: pypdf; extra == 'all'
Requires-Dist: pypdf2; extra == 'all'
Requires-Dist: python-docx; extra == 'all'
Requires-Dist: python-dotenv; extra == 'all'
Requires-Dist: python-pptx; extra == 'all'
Requires-Dist: pytz; extra == 'all'
Requires-Dist: rapidfuzz; extra == 'all'
Requires-Dist: requests-toolbelt~=1.0.0; extra == 'all'
Requires-Dist: scikit-image; extra == 'all'
Requires-Dist: scipy; extra == 'all'
Requires-Dist: tldextract; extra == 'all'
Requires-Dist: xmltodict; extra == 'all'
Provides-Extra: dev
Requires-Dist: httpx<0.28,>=0.27; extra == 'dev'
Requires-Dist: pytest>=8.0; extra == 'dev'
Provides-Extra: gym
Requires-Dist: dbus-next>=0.2.3; extra == 'gym'
Requires-Dist: jsonschema>=4.20; extra == 'gym'
Requires-Dist: numpy>=1.26; extra == 'gym'
Requires-Dist: paramiko>=3.4; extra == 'gym'
Requires-Dist: pillow>=10.0; extra == 'gym'
Requires-Dist: pycryptodome>=3.20; extra == 'gym'
Requires-Dist: rich>=13.0; extra == 'gym'
Provides-Extra: modal
Requires-Dist: modal>=1.5.5; extra == 'modal'
Provides-Extra: osworld
Requires-Dist: beautifulsoup4; extra == 'osworld'
Requires-Dist: borb<3; extra == 'osworld'
Requires-Dist: chardet; extra == 'osworld'
Requires-Dist: cssselect; extra == 'osworld'
Requires-Dist: easyocr; extra == 'osworld'
Requires-Dist: fastdtw; extra == 'osworld'
Requires-Dist: formulas; extra == 'osworld'
Requires-Dist: gdown; extra == 'osworld'
Requires-Dist: gymnasium~=0.28.1; extra == 'osworld'
Requires-Dist: imagehash; extra == 'osworld'
Requires-Dist: librosa; extra == 'osworld'
Requires-Dist: lxml; extra == 'osworld'
Requires-Dist: mutagen; extra == 'osworld'
Requires-Dist: numpy; extra == 'osworld'
Requires-Dist: odfpy; extra == 'osworld'
Requires-Dist: opencv-python-headless; extra == 'osworld'
Requires-Dist: openpyxl; extra == 'osworld'
Requires-Dist: pandas; extra == 'osworld'
Requires-Dist: paramiko; extra == 'osworld'
Requires-Dist: pdfplumber; extra == 'osworld'
Requires-Dist: pillow; extra == 'osworld'
Requires-Dist: playwright; extra == 'osworld'
Requires-Dist: pyacoustid; extra == 'osworld'
Requires-Dist: pydrive; extra == 'osworld'
Requires-Dist: pygame; extra == 'osworld'
Requires-Dist: pymupdf; extra == 'osworld'
Requires-Dist: pypdf; extra == 'osworld'
Requires-Dist: pypdf2; extra == 'osworld'
Requires-Dist: python-docx; extra == 'osworld'
Requires-Dist: python-dotenv; extra == 'osworld'
Requires-Dist: python-pptx; extra == 'osworld'
Requires-Dist: pytz; extra == 'osworld'
Requires-Dist: rapidfuzz; extra == 'osworld'
Requires-Dist: requests-toolbelt~=1.0.0; extra == 'osworld'
Requires-Dist: scikit-image; extra == 'osworld'
Requires-Dist: scipy; extra == 'osworld'
Requires-Dist: tldextract; extra == 'osworld'
Requires-Dist: xmltodict; extra == 'osworld'
Provides-Extra: platform
Requires-Dist: fastapi>=0.115; extra == 'platform'
Requires-Dist: itsdangerous>=2.2; extra == 'platform'
Requires-Dist: jinja2>=3.1; extra == 'platform'
Requires-Dist: python-multipart>=0.0.9; extra == 'platform'
Requires-Dist: sqlalchemy>=2.0; extra == 'platform'
Requires-Dist: uv==0.10.2; extra == 'platform'
Requires-Dist: uvicorn[standard]>=0.30; extra == 'platform'
Description-Content-Type: text/markdown

# cua-speedrun

Compare computer-use agents by performance, time, and cost on real desktop tasks.

## Quickstart

Install and configure your credentials:

```bash
pip install cua-speedrun
cua-speedrun setup
cua-speedrun benchmark --dataset osworld-50 --agent qwen3vl
```

Prebuilt desktop images are imported from Docker Hub and cached in Modal.
Models and application setup are cached on first use. Data lives in
`~/.local/share/cua-speedrun`; set `CUA_SPEEDRUN_HOME` to use another location.
Use `cua-speedrun doctor` to inspect your installation.

## Run an evaluation

Evaluations run on Modal by default, using credentials from `setup`, your
Modal CLI profile, or environment variables. For an API agent:

```bash
export ANTHROPIC_API_KEY="your-api-key"
cua-speedrun benchmark --dataset osworld-50 --agent claude
```

The command follows progress until the evaluation finishes. To launch and
inspect evaluations in a browser, run `cua-speedrun dashboard`.

- Add `--parallel-evaluations 4` to run four agent replicas in parallel.
- Bundled GPU agents select their GPU automatically; override it with `--gpu L40S`.
- Add `--no-preload` to disable environment preloading.
- To use local Linux hardware, add `--compute local --environment local`.
  Local desktops require KVM/QEMU.

Inspect results or download trajectories:

```bash
cua-speedrun evaluations
cua-speedrun status RUN_ID
cua-speedrun export RUN_ID
```

Use `cua-speedrun catalog` to list available agents and benchmarks, or
`cua-speedrun help benchmark` for more options.

## Benchmarks

| Benchmark | Tasks |
| --- | ---: |
| `cua-world-26` | 26 |
| `osworld-50` | 50 |
| `osworld2-52` | 52 |
| `my-pc-bench` | 38 |
| `cua-world-offline` | 143 |
| `osworld-offline` | 295 |
| `osworld2-offline` | 63 |

The `offline` variants contain the full offline task sets; the smaller variants
are representative subsets. MyPCBench also requires a
`MYPCBENCH_JUDGE_API_KEY` for its evaluator; see its
[setup instructions](benchmarks/my-pc-bench/README.md).

## Bring your own agent

Start from an implementation in [`agents/`](agents/). Each agent has two files:

- `init.py` prepares dependencies or starts a model server before task timing begins.
- `agent.py` receives the environment URL and task description, then interacts
  through [`Computer`](src/cua_speedrun/client.py).

Submit the folder directly:

```bash
cua-speedrun validate --agent ./my-agent
cua-speedrun benchmark --dataset osworld-50 --agent ./my-agent
```

Additional packages can be installed by `init.py`; the submission uploads
`init.py` and `agent.py`. An optional `agent.json` declares `gpu` and
`required_environment_variables`. To contribute an agent, add its folder to
`agents/`; it is discovered automatically.

## Bring your own benchmark

A benchmark is a folder with a `manifest.yaml`, task folders containing
`task.yaml`, and its environment setup and verifier. The manifest lists tasks:

```yaml
name: my-benchmark
version: "1"
tasks: [tasks/my-task]
```

Each `task.yaml` specifies `task_id`, `description`, and an `env` mapping:

```yaml
task_id: my-task
description: The task for the agent to complete.
env:
  kind: gym-anything
  env_dir: ${BENCHMARK_DIR}/environment
  task_id: my-task
```

Keep the [Gym-Anything environment](https://github.com/cmu-l3/gym-anything)
and its task setup/verifier inside the benchmark folder. Then:

```bash
cua-speedrun validate --dataset ./my-benchmark
cua-speedrun benchmark --dataset ./my-benchmark --agent ./my-agent
```

The benchmark is copied into the installation. Increase its version when
changing a registered task set. To contribute it, add the folder under
`benchmarks/` and its name to `catalog/benchmarks.yaml`; packaging is automatic.

## Repository structure

- [`agents/`](agents/) — agent implementations.
- [`benchmarks/`](benchmarks/) — task sets and benchmark definitions.
- [`src/cua_speedrun/`](src/cua_speedrun/) — execution, timing, scoring, and the dashboard.
- [`src/cua_speedrun/compute_runners/`](src/cua_speedrun/compute_runners/) — local and Slurm runners.
