Metadata-Version: 2.5
Name: nomad-harness
Version: 0.1.0.dev2
Summary: Action runtime connecting frontier models to simulated and physical embodiments.
Project-URL: Repository, https://github.com/runtime-si/Nomad_Harness_Interface
Project-URL: Issues, https://github.com/runtime-si/Nomad_Harness_Interface/issues
Author: Lambda Robotics
License-Expression: MIT
License-File: LICENSE
Keywords: agents,llm,mujoco,robosuite,robotics,simulation
Classifier: Development Status :: 2 - Pre-Alpha
Classifier: Intended Audience :: Science/Research
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Typing :: Typed
Requires-Python: >=3.10
Requires-Dist: pydantic<3,>=2.6
Requires-Dist: pyyaml>=6.0
Requires-Dist: typing-extensions>=4.12
Provides-Extra: anthropic
Requires-Dist: anthropic<1,>=0.125; extra == 'anthropic'
Provides-Extra: dev
Requires-Dist: anthropic<1,>=0.125; extra == 'dev'
Requires-Dist: google-genai<3,>=2.28; extra == 'dev'
Requires-Dist: jsonschema>=4.18; extra == 'dev'
Requires-Dist: mypy>=1.10; extra == 'dev'
Requires-Dist: openai<4,>=3.24; extra == 'dev'
Requires-Dist: pytest>=8.0; extra == 'dev'
Requires-Dist: ruff>=0.5; extra == 'dev'
Requires-Dist: types-pyyaml; extra == 'dev'
Provides-Extra: gemini
Requires-Dist: google-genai<3,>=2.28; extra == 'gemini'
Provides-Extra: mujoco
Requires-Dist: imageio[ffmpeg]; extra == 'mujoco'
Requires-Dist: mujoco<3.4,>=3.3.7; extra == 'mujoco'
Requires-Dist: numpy; extra == 'mujoco'
Requires-Dist: pyopengl<4,>=3.1; extra == 'mujoco'
Provides-Extra: openai
Requires-Dist: openai<4,>=3.24; extra == 'openai'
Provides-Extra: robosuite
Requires-Dist: imageio[ffmpeg]; extra == 'robosuite'
Requires-Dist: mujoco==3.3.7; extra == 'robosuite'
Requires-Dist: numpy; extra == 'robosuite'
Requires-Dist: pyopengl<4,>=3.1; extra == 'robosuite'
Requires-Dist: robosuite==1.5.2; extra == 'robosuite'
Description-Content-Type: text/markdown

# Nomad Harness

A Python action runtime connecting model providers to simulated embodiments through
typed observations, bounded actions, and execution feedback.

Nomad includes OpenAI, Anthropic, Gemini, and OpenAI-compatible providers (including
DeepSeek and Kimi presets), deterministic offline providers, a fake cube-lift world,
Robosuite tasks on Panda, UR5e, Kinova3, IIWA, and XArm7 robots, and any MuJoCo scene
through a plain MuJoCo adapter. Python 3.10+ is required. Seeded benchmark suites run
any of these with per-task results. This is a development release; hardware adapters,
and a CLI are planned in the [roadmap](ROADMAP.md). Anthropic has offline contract
tests; live proposal validation is pending API credits.

## Install

Install the published development version explicitly:

```bash
python -m pip install 'nomad-harness==0.1.0.dev2'
# For OpenAI and Robosuite:
python -m pip install 'nomad-harness[robosuite,openai]==0.1.0.dev2'
```

Use an exact version instead of `--pre` to avoid opting all dependencies into
pre-releases. From a source checkout, install the current code with
`python -m pip install -e '.[dev]'`, adding `robosuite`, `mujoco`, `openai`,
`anthropic`, or `gemini` for those integrations (the `openai` extra also serves
DeepSeek, Kimi, and other OpenAI-compatible servers). Robosuite uses the tested
pair `robosuite==1.5.2`, `mujoco==3.3.7`, and `PyOpenGL<4`; the plain MuJoCo
adapter (`[mujoco]`) uses `mujoco` 3.3 with the rendering and video dependencies
listed in that extra.
Provider and simulator SDKs are loaded only when their adapters need them.

## Run offline first

This complete example needs no API key, simulator, or GPU:

```python
from nomad import Agent, RunConfig
from nomad.embodiments import FakeEmbodiment, lift_script
from nomad.models import FakeModel

with FakeEmbodiment() as env:
    agent = Agent(FakeModel(lift_script()), env, RunConfig(max_decisions=30))
    result = agent.run("Lift the cube off the table.")

print(result.status, result.outcome.reason)
print("trace:", result.trace_dir)
```

The fake model follows a fixed script. For a walkthrough that also demonstrates an
invalid action and recovery, run `python examples/offline_demo.py` from the checkout.

## Run a model on Robosuite

Set `OPENAI_API_KEY` and `NOMAD_OPENAI_MODEL` to a model your account can access.
This example sends images to the API and spends API credits:

```python
import os

from nomad import Agent, RunConfig
from nomad.embodiments import Robosuite
from nomad.models import OpenAI

with Robosuite(task="Lift", robot="Panda") as env, OpenAI(
    model=os.environ["NOMAD_OPENAI_MODEL"]
) as model:
    agent = Agent(model, env, RunConfig(max_decisions=30, max_wall_time_s=600))
    result = agent.run("Lift the cube off the table.")

print(result.status, result.outcome.reason)
print("trace:", result.trace_dir)
```

`Agent` owns neither adapter. Close models and embodiments with context managers,
including when an exception occurs. An environment can serve multiple runs; each
`run()` resets it with the configured seed. Stateful models may retain their state
across runs—create a fresh `FakeModel` to replay its script.

On headless Linux, set `MUJOCO_GL=egl` or `MUJOCO_GL=osmesa` for offscreen rendering.
The checkout includes `examples/quickstart.py`, `examples/robosuite_model.py`,
`examples/robosuite_scripted.py`, `examples/robosuite_custom_task.py`,
`examples/mujoco_scripted.py`, and `examples/benchmark_scripted.py` with command-line
options.

Other providers are drop-in replacements for `OpenAI(...)`; every adapter offers the
same actions as one tool per action and validates every reply locally:

| Provider | Adapter | Extra | Key variable | Verified |
|---|---|---|---|---|
| OpenAI | `OpenAI` (Responses API) | `openai` | `OPENAI_API_KEY` | Live proposals and Lift episodes |
| Anthropic | `Anthropic` (Messages API) | `anthropic` | `ANTHROPIC_API_KEY` | Offline request and response tests; live validation pending API credits |
| Gemini | `Gemini` | `gemini` | `GEMINI_API_KEY` | Live proposals with `gemini-robotics-er-2-preview` |
| DeepSeek | `DeepSeek` | `openai` | `DEEPSEEK_API_KEY` | Request shape only; live replies, strict tools, forced calls, and images not yet tested |
| Kimi | `Kimi` | `openai` | `MOONSHOT_API_KEY` | Request shape only; live replies, strict tools, forced calls, and images not yet tested |
| Any Chat Completions server | `OpenAICompatible(base_url=...)` | `openai` | set `api_key_env` | Request shape checked against OpenAI's Chat Completions |

```bash
python examples/robosuite_model.py --provider gemini --model gemini-robotics-er-2-preview
python examples/robosuite_model.py --provider anthropic --model "$NOMAD_ANTHROPIC_MODEL"
```

`create_provider("gemini")` builds an adapter from its `NOMAD_GEMINI_MODEL` variable,
and `nomad.models.PROVIDERS` lists each provider's key and model variables. Calls made
by `Agent` runs are bounded by the run deadline and are never retried; a failed call
becomes feedback and counts toward `max_consecutive_failures`. For direct `propose()`
calls, the OpenAI-compatible adapters retry rate limits and server errors but fail at
once on quota or billing errors, while `OpenAI`, `Anthropic`, and `Gemini` use
their SDKs' retries.
Free tiers can be smaller than an episode: on 2026-10-06 our Gemini free-tier quota
ran out after about a dozen requests, and an episode needs 20–30.

Install Claude support with `pip install -e '.[anthropic]'`. The adapter reads
`ANTHROPIC_API_KEY` from the environment exported by your config; select a model
with `--model` or `NOMAD_ANTHROPIC_MODEL`.
For keys that require workspace routing, set `ANTHROPIC_WORKSPACE_ID` or pass
`workspace_id` to the adapter. You can also use it directly:

```python
from nomad.models import Anthropic

model = Anthropic(model="your-claude-model-id", max_output_tokens=4096)
```

Claude receives the same state and camera images as the other providers. The
adapter disables parallel tool calls and defaults to `tool_choice="auto"` for
compatibility across Claude models. On models that support forced tools, pass
`tool_choice="any"`. Every returned proposal is still validated locally;
truncated replies cannot dispatch actions. Thinking content and image bytes
are excluded from provider trace details.

## Tasks, robots, and cameras

| Task | Goal | Robots |
|---|---|---|
| `Lift` | Lift the cube off the table | Panda, UR5e, Kinova3, IIWA, XArm7 |
| `Stack` | Stack the red cube on the green cube | Panda, UR5e, Kinova3, IIWA, XArm7 |
| `PickPlaceCan` | Place the red can in its target compartment and release | Panda, UR5e, IIWA |
| `PickPlaceMilk` | Place the milk carton in its target compartment and release | Panda, UR5e |
| `PickPlaceBread` | Place the loaf of bread in its target compartment and release | Panda, UR5e |
| `PickPlaceCereal` | Place the cereal box in its target compartment and release | Panda, UR5e |
| `NutAssemblyRound` | Drop the round nut over the round peg | Panda, UR5e |
| `NutAssemblySquare` | Drop the square nut over the square peg | Panda |

`env.default_goal` provides each task's instruction. Call
`nomad.embodiments.robosuite.supported_pairs()` to list packaged manifests.
Grippers: Panda (its own parallel gripper), UR5e and Kinova3 (Robotiq 2F-85), IIWA
(Robotiq 2F-140), XArm7 (its own linkage gripper). Each pair has measured workspace
and controller limits; inspect its packaged manifest before changing those limits.
`robot_base` is the robot's fixed root body (`robot0_base`, at the mount) for every
robot. The XArm7 tracks targets less tightly, so its `move_ee` position tolerance
defaults to 8 mm instead of 5 mm. The Kinova3 and XArm7 do not carry the can to its
compartment reliably, the UR5e's grip lets the square nut swing and land tilted, and
Sawyer and Jaco failed basic grasps, so those pairs have no manifests.
For another pair, supply matching `task`, `robot`, and `manifest="path.yaml"`;
`draft_manifest(task=..., robot=..., path=...)` measures a scene and proposes one. Your
own tasks and benchmark variants (another robosuite environment, extra `robosuite.make`
arguments, a stricter success predicate) are added with `register_task`; see
[custom tasks](docs/custom-adapters.md#custom-robosuite-tasks-and-benchmarks) and
`examples/robosuite_custom_task.py`.

Select camera streams when creating the environment:

```python
with Robosuite(
    task="Lift",
    cameras={"front_rgb": "agentview", "wrist_rgb": "robot0_eye_in_hand"},
) as env:
    print(env.manifest().camera_names)
```

Every declared stream must have a camera mapping. Added streams default to 640×480;
set `added_stream_resolution=(width, height)` to change that. Cameras are checked
against the scene at construction. `OpenAI(cameras=("wrist_rgb",), model=...)` sends
only that stream; traces retain all streams. Video uses its own `video_camera`.

Scripted baselines use privileged object state and are not model benchmarks. Run
one with `python examples/robosuite_scripted.py --task Stack --robot UR5e`.
Each packaged pair's baseline succeeded on seeds 0–9 when it was added (PickPlaceBread
on the Panda: 9/10, a finger hits the bin wall on seed 9); this checks the adapter and
controller, not visual model performance.

## Plain MuJoCo scenes

`nomad.embodiments.mujoco` drives an arm in any MJCF model, without robosuite. Each
control tick solves damped least-squares inverse kinematics for the `move_ee` target and
commands the arm's position actuators; `set_gripper` drives one gripper actuator. The
packaged demo is a six-joint arm built from primitive shapes lifting a cube:

```python
from nomad import Agent, ObservationMode, RunConfig
from nomad.embodiments.mujoco import tabletop_lift
from nomad.models import ScriptedLiftPolicy

config = RunConfig(observation_mode=ObservationMode.PRIVILEGED_STATE, max_decisions=40)
with tabletop_lift() as env:
    result = Agent(ScriptedLiftPolicy(), env, config).run(env.default_goal)
print(result.status, result.outcome.reason)
```

`python examples/mujoco_scripted.py` runs it with a trace and video; the scripted lift
succeeded on seeds 0–19. For your own scene, name the end-effector site, arm joints and
actuators, and gripper in `MujocoConfig`, describe placement and success in a
`MujocoTask`, and write a manifest with `backend: mujoco`; see
[plain MuJoCo scenes](docs/custom-adapters.md#plain-mujoco-scenes).

## Benchmarks

`nomad.benchmark.run_benchmark` runs a suite of tasks over seeds through the normal
runtime. Each `BenchmarkTask` builds its embodiment (any backend) and may bring its own
model; otherwise the suite's model is used:

```python
from nomad import ObservationMode, RunConfig
from nomad.benchmark import BenchmarkTask, run_benchmark
from nomad.embodiments import Robosuite
from nomad.embodiments.mujoco import tabletop_lift
from nomad.models import ScriptedLiftPolicy, robosuite_baseline

tasks = [
    BenchmarkTask("lift-panda", lambda: Robosuite(task="Lift"),
                  make_model=lambda: robosuite_baseline("Lift", "Panda")),
    BenchmarkTask("mujoco-lift", tabletop_lift, make_model=ScriptedLiftPolicy),
]
config = RunConfig(observation_mode=ObservationMode.PRIVILEGED_STATE, max_decisions=60)
result = run_benchmark(tasks, run_config=config, output_dir="benchmarks")
print(result.table())
```

The run directory holds `benchmark.json` (what was run), `results.jsonl` (one record per
episode, written as each finishes), `summary.json` (per-task success rates with Wilson
95% intervals), and each episode's trace. Success means the embodiment's evaluator
confirmed the task. Episodes that fail to run (including a task without a goal) are
recorded with their error and count as failures; cancelled episodes are recorded but not
scored. Every record and summary names its observation mode, so privileged-state results
stay separate from image-only ones, and different tasks are never pooled into one score. `examples/benchmark_scripted.py`
benchmarks the scripted baselines (`--pairs all` for every packaged robosuite pair).

## Configuration

`RunConfig` controls each episode. Adapter-specific settings live in `OpenAIConfig`
and `RobosuiteConfig`; pass a config object or typed keyword overrides:

```python
from pathlib import Path
from nomad import RunConfig
from nomad.models import OpenAI, OpenAIConfig

config = RunConfig(
    seed=0,
    max_decisions=30,
    max_wall_time_s=180,
    max_provider_calls=30,
    max_total_tokens=20_000,
    history_window=8,
    trace_root=Path("runs"),
    save_frames=True,
    save_video=False,
)

with OpenAI(OpenAIConfig(model="your-model-id", timeout_s=45), max_retries=1) as model:
    print(model.model_id)
```

| Setting | Default | Meaning |
|---|---|---|
| `max_decisions` | 30 | Maximum model decisions |
| `max_wall_time_s` | 180 | Cooperative episode wall-time limit |
| `max_sim_time_s` | `None` | Optional simulation-time limit, checked between actions |
| `max_provider_calls` | `None` | Optional proposal-call limit, including failures |
| `max_total_tokens` | `None` | Stop further decisions once reported tokens reach this limit |
| `max_consecutive_failures` | 5 | Stop after repeated provider, validation, or execution failures |
| `action_freshness_s` | 60 | Reject an action planned from an observation older than this window |
| `history_window` | 8 | Feedback entries sent to the model; 0 disables feedback |
| `observation_mode` | `rgb_proprio` | Images and robot state; opt into `privileged_state` for baselines |

Limits are cooperative: arbitrary synchronous custom providers cannot be forcibly
interrupted. Nomad checks the wall deadline after inference and does not dispatch
late actions. OpenAI requests receive the remaining time as a timeout cap, with SDK
retries disabled for that bounded request. Direct `OpenAI.propose()` calls retain
the adapter's configured retry policy. Transport timeouts are not hard process-level
deadlines, and an in-flight action runs to its own bounded completion or stop request.

Token limits use all reported usage as a lower bound, including partial usage after
an error. Exact totals stay `None` if any usage is missing; unreported consumption
cannot be bounded, and one call can exceed the remaining token allowance.

Records reject unknown fields and freeze field assignment. Nested JSON dictionaries
remain editable snapshots: `env.manifest()`, adapter `.config`, and provider-visible
context are detached from the adapter/runtime state. Create a new validated config
and adapter to change active settings. Pydantic's `model_copy(update=...)` does not
validate updates; use the config constructor or `model_validate()` for new settings.

## Results, errors, and cancellation

`agent.run(goal)` returns a `RunResult`:

- `status`: `success`, `failure`, `budget_exhausted`, `cancelled`, or `error`.
- `stop_reason`: the specific budget, terminal outcome, cancellation, or error.
- `outcome`: the independent task evaluator's result and evidence.
- `metrics`: decisions, calls, tokens, action results, and elapsed time.
- `trace_dir` and `error`: artifact location and any recorded error.

Task success and run health are separate. A failed `stop()` produces run status
`error` even if `outcome.status` is `success`. A video-finalization error is recorded
in `error` without changing the task's result. Inspect `error` as well as `status`.
Model refusals, malformed proposals, and execution failures become feedback for the
next decision; repeated failures exhaust the configured failure budget. Constructor
errors (such as an invalid manifest or missing API key) raise immediately. Runtime
exceptions normally become an error result; trace I/O failures can still propagate.

Pass a `threading.Event` as `agent.run(goal, cancel=event)` to request cancellation.
It is checked between decisions and after inference. To interrupt an action already
executing, also call `env.stop("cancelled")` from the cancelling thread.
`KeyboardInterrupt` finalizes the run as cancelled and is re-raised; a stop failure
is recorded as a run error.

## Traces

Each run creates a new directory under `trace_root` containing:

- `config.json`: resolved settings, manifest fingerprint, instructions, schema, and provenance.
- `trace.jsonl`: observations, proposals, validation, execution, and outcome events.
- `metrics.json`: final status, outcome, error, and counters.
- `frames/`: captured images when enabled.
- `video.mp4`: recording when enabled and supported by the embodiment.
- `explorer.html`: a self-contained viewer for the run. Open it in a browser (no server
  needed) to scrub the video against a step timeline, plot end-effector and object
  heights, and inspect each decision's model input, output, validation, and execution.
  Disable with `save_explorer=False`.

Credential keys and recognized secret environment values are redacted. Adapter
metadata must still exclude credentials. Images are stored separately and referenced
by the trace; privileged evaluator data is excluded from the default model context.

## Extend and contribute

- [Custom provider and embodiment guide](docs/custom-adapters.md), with a runnable offline example.
- [Contracts and execution semantics](docs/contracts.md).
- [Development, testing, and releases](docs/development.md).
- [Roadmap and proposed integrations](ROADMAP.md).

MIT license; see [LICENSE](LICENSE).
