Metadata-Version: 2.5
Name: robouse
Version: 0.1.0
Summary: Agent as a policy for embodied agents
Project-URL: Homepage, https://robouse.ai
Project-URL: Documentation, https://robouse.ai/docs
Project-URL: Repository, https://github.com/benchflow-ai/robouse
Project-URL: Issues, https://github.com/benchflow-ai/robouse/issues
Author-email: Xiangyi Li <xiangyi@benchflow.ai>
License-Expression: Apache-2.0
License-File: LICENSE
Keywords: benchflow,benchmark,embodied-agents,llm-agents,mujoco,robotics
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: Operating System :: MacOS
Classifier: Operating System :: POSIX :: Linux
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Python: >=3.11
Requires-Dist: numpy>=1.24
Requires-Dist: pyyaml>=6
Provides-Extra: all
Requires-Dist: gymnasium-robotics>=1.4; extra == 'all'
Requires-Dist: imageio-ffmpeg>=0.5; extra == 'all'
Requires-Dist: imageio>=2.34; extra == 'all'
Requires-Dist: metaworld>=3.1; extra == 'all'
Requires-Dist: mujoco>=3.3; extra == 'all'
Requires-Dist: packaging; extra == 'all'
Requires-Dist: pillow>=10; extra == 'all'
Requires-Dist: robosuite<1.6,>=1.5; extra == 'all'
Requires-Dist: scipy>=1.10; extra == 'all'
Provides-Extra: dexjoco
Requires-Dist: imageio-ffmpeg>=0.5; extra == 'dexjoco'
Requires-Dist: imageio>=2.34; extra == 'dexjoco'
Requires-Dist: mujoco>=3.3; extra == 'dexjoco'
Requires-Dist: pillow>=10; extra == 'dexjoco'
Requires-Dist: scipy>=1.10; extra == 'dexjoco'
Provides-Extra: gymrobotics
Requires-Dist: gymnasium-robotics>=1.4; extra == 'gymrobotics'
Requires-Dist: imageio-ffmpeg>=0.5; extra == 'gymrobotics'
Requires-Dist: imageio>=2.34; extra == 'gymrobotics'
Requires-Dist: mujoco>=3.3; extra == 'gymrobotics'
Requires-Dist: pillow>=10; extra == 'gymrobotics'
Provides-Extra: metaworld
Requires-Dist: imageio-ffmpeg>=0.5; extra == 'metaworld'
Requires-Dist: imageio>=2.34; extra == 'metaworld'
Requires-Dist: metaworld>=3.1; extra == 'metaworld'
Requires-Dist: mujoco>=3.3; extra == 'metaworld'
Requires-Dist: packaging; extra == 'metaworld'
Requires-Dist: pillow>=10; extra == 'metaworld'
Provides-Extra: robosuite
Requires-Dist: imageio-ffmpeg>=0.5; extra == 'robosuite'
Requires-Dist: imageio>=2.34; extra == 'robosuite'
Requires-Dist: mujoco>=3.3; extra == 'robosuite'
Requires-Dist: pillow>=10; extra == 'robosuite'
Requires-Dist: robosuite<1.6,>=1.5; extra == 'robosuite'
Provides-Extra: sim
Requires-Dist: imageio-ffmpeg>=0.5; extra == 'sim'
Requires-Dist: imageio>=2.34; extra == 'sim'
Requires-Dist: mujoco>=3.3; extra == 'sim'
Requires-Dist: pillow>=10; extra == 'sim'
Description-Content-Type: text/markdown

# robouse

**robo-use: agent as a policy for embodied agents.** A BenchFlow extension that turns robot manipulation and navigation benchmarks into tasks an LLM agent harness solves by driving the robot itself, one shell command at a time. Think Terminal-Bench plus ARC, for robots.

- **Tasks** are BenchFlow-native task folders adapted from existing benchmarks (Meta-World, Gymnasium-Robotics, LIBERO-style scenes, ARC-style rule inference, robo-use families, RoboHarm-style safety pairs). Every task ships a reference solution that scores 1.0.
- **Harnesses** are the agents people already use (Claude Code, Codex, mini-swe-agent, ...) with any model. No harness integration is needed beyond a shell: the robot is driven through the `robo` command.
- **The episode server is trusted.** It owns the simulator, enforces the step budget, records video and judges success from the physical state. The agent never touches simulator objects.
- **Every trial is recorded** as a BenchFlow trial directory with video, so it opens in the BenchFlow viewer.

Site: [robouse.ai](https://robouse.ai). Source: [github.com/benchflow-ai/robouse](https://github.com/benchflow-ai/robouse). Status log: [STATUS.md](https://github.com/benchflow-ai/robouse/blob/main/STATUS.md).

## Install

Python 3.11 or newer. The base package is light (numpy and PyYAML): it gives you the task loader, the `robouse` command, and the agent-facing `robo` command, which uses only the standard library. Simulators are extras:

```sh
pip install robouse                  # task loader, `robouse` and `robo` commands
pip install 'robouse[sim]'           # + MuJoCo and video recording: tabletop, ARC-style, hard, safety, vision,
                                     #   robo-use families, RoboHarm, menagerie and drone suites
pip install 'robouse[metaworld]'     # + Meta-World
pip install 'robouse[gymrobotics]'   # + Gymnasium-Robotics (Fetch, PointMaze)
pip install 'robouse[robosuite]'     # + robosuite 1.5
pip install 'robouse[all]'           # all of the above
```

LIBERO, RoboCasa, BEHAVIOR and DexJoco run in their own environments; see the suite docs linked below.

**Tasks ship with the package.** All task folders (about 5 MB of text: `task.md`, reference solution, verifier) are inside the wheel, so a task can be named by its id. `robouse tasks` lists them and `robouse tasks --path` prints where they are; copy that folder if you want to edit tasks. **Robot meshes are fetched, not shipped:** the menagerie, RoboHarm and drone suites use MuJoCo Menagerie models (about 43 MB); run `robouse fetch-assets` once to download them from the pinned upstream commit into `~/.cache/robouse/menagerie`, with a SHA-256 check per file.

## Quickstart

```sh
# list the bundled tasks (id, backend, env)
robouse tasks

# check a task with its reference solution (reward should be 1)
robouse run --task arc-gravity --harness oracle --out runs/try
robouse run --task metaworld-reach --harness oracle --out runs/try          # needs robouse[metaworld]

# let an agent be the policy (needs the `claude` or `codex` CLI and credentials; see docs/harnesses.md)
robouse run --task metaworld-push --harness claude-code --model claude-opus-5-5 --out runs/try
robouse run --task metaworld-push --harness codex --model gpt-6-astra --out runs/try

# a whole suite, 4 at a time (a task folder path works anywhere a task id does)
robouse run-many --tasks "$(robouse tasks --path)/metaworld" --harness codex --model gpt-6-astra --out runs/mw-codex --concurrency 4
```

To work on robouse itself, install a checkout in editable mode; it then uses the checkout's `tasks/` and `assets/`:

```sh
git clone https://github.com/benchflow-ai/robouse.git && cd robouse
uv venv --python 3.12 .venv
uv pip install --python .venv/bin/python -e ".[metaworld,gymrobotics]"
```

Each trial writes `runs/<job>/<task>__<harness>__<id>/` with `result.json`, the agent's raw output and trajectory, the episode trace, `verifier/reward.txt` and `artifacts/recording.mp4`.

## How it works

```
 harness (Claude Code, Codex, ...)            trusted episode server (robouse serve)
 ┌─────────────────────────────────┐   JSON    ┌────────────────────────────────────┐
 │ reads instruction.md            │  over a   │ MuJoCo simulator (Meta-World,      │
 │ runs `robo observe`, `robo act` │◄─────────►│ tabletop, Gymnasium-Robotics)      │
 │ ... `robo done`                 │  Unix     │ step + wall-clock budgets          │
 └─────────────────────────────────┘  socket   │ video + trace recording            │
                                               │ success judged from physical state │
                                               └────────────────────────────────────┘
```

The agent sees only what `robo` returns: robot and object state as numbers, and camera images on request. See [docs/robo-cli.md](https://github.com/benchflow-ai/robouse/blob/main/docs/robo-cli.md).

## Docs

- [docs/task-format.md](https://github.com/benchflow-ai/robouse/blob/main/docs/task-format.md): task folders, the `robouse:` settings block in `task.md`, observation and scoring modes
- [docs/robo-cli.md](https://github.com/benchflow-ai/robouse/blob/main/docs/robo-cli.md): the agent-facing `robo` command
- [docs/harnesses.md](https://github.com/benchflow-ai/robouse/blob/main/docs/harnesses.md): supported harnesses and models, credentials, isolation
- [docs/suites/](https://github.com/benchflow-ai/robouse/blob/main/docs/suites/): each task suite and where it comes from
- [docs/results.md](https://github.com/benchflow-ai/robouse/blob/main/docs/results.md): solve rates per harness, model and suite (generated by `scripts/results.py` from the audit)
- [docs/viewer.md](https://github.com/benchflow-ai/robouse/blob/main/docs/viewer.md), [docs/audit.md](https://github.com/benchflow-ai/robouse/blob/main/docs/audit.md): looking at and checking runs
- [docs/benchflow.md](https://github.com/benchflow-ai/robouse/blob/main/docs/benchflow.md): running robo-use tasks through BenchFlow (`bench eval run`)
- [docs/molmoact2.md](https://github.com/benchflow-ai/robouse/blob/main/docs/molmoact2.md): MolmoAct2 (a VLA) runs on a remote GPU
