Nenyax: a two-layer protocol on top of OpenEnv
Many environment formats come in at the top, many trainers connect at the bottom, and one narrow contract sits in the middle. OpenEnv stays the standard; Nenyax adds the adapters that route every other format into it, plus the audit that makes the result trustworthy.
Environment sources · any formatport in ↓ port out ↑
OpenEnv Meta · HF (native)
Verifiers Prime Intellect
Harbor Laude · Terminal-Bench
NeMo Gym NVIDIA
ORS / OpenReward General Reasoning
Gymnasium · PettingZoo Farama
dm_env · OpenSpiel Google DeepMind
MuJoCo Playground Google DeepMind
AndroidWorld Google
BrowserGym ServiceNow
Reasoning Gym · GEM · SkyRL-Gym
HUD MCP-based
nenyax build codebases · industries
↓the router reads each source's format and picks a driver
Layer 2 · adapters and routerone driver per format, grouped by how the environment is driven
Step mode
The trainer drives each step with reset / step, or with a tool call per turn.
OpenEnvGymnasiumdm_envOpenSpielPettingZooBrowserGymORSGEM
Rollout mode
The environment runs the whole multi-turn loop against a model endpoint and returns a scored trajectory.
VerifiersNeMo GymReasoning Gym
Task mode
A container holds the task. An agent harness runs inside it, and tests grade the final state.
HarborTerminal-BenchSWE-bench-styleAndroidWorld
Router: manifest.source.format → load the driver → negotiate capabilities → expose the Layer 1 contract. A driver never pretends to support something it can't: a task-mode environment reports branchable:false, and the trainer adapts.
↓every driver exposes the same contract
Layer 1 · Nenyax core contract (OpenEnv plus proposed RFCs)the narrow waist
Manifestidentity, version, licence, format, mode
Capabilitiesresettable, branchable, parallel, risk, real-world
ActionsMCP tools (OpenEnv RFC 003)
Judgmentverifier, rubric or external reward; versioned, with confidence
Trajectoryobservations, actions, tokens and logprobs, judgment
open→sample→act / observe→judge→close
· rollout and task modes run this loop internally and return the same trajectory
↓trainer bindings
Trainers · any backend
TRLverlprime-rlSkyRLNeMo RLUnslothTinkerOpenRLHFARTslime
Audit gates (every environment, ported or built)
- Builds and boots cleanly; the network is isolated
- The reference solution passes and deliberately broken solutions fail
- A red-team agent attempts reward hacks
- Results are the same when re-run
- Difficulty is calibrated: a baseline model passes between about 3% and 80%
- No contamination or leakage, such as a git history that reveals the answer
- → a standard quality report ships with the environment
Builder skills (for any coding agent)
- Installed as an MCP server or skill in Claude Code, Codex, Optimus, OpenCode and others
build from a repo, an OpenAPI spec, an app or a domain's documents
- Tools for containerizing, mocking, snapshotting and mining tasks
- The agent proposes; the audit gates accept or reject
- One codebase can yield several environments: coding, operator, support, SRE, data, security
How each format maps onto the core contract
| Format | Mode | Native shape | What the driver does |
| OpenEnv | step | reset/step over HTTP/WebSocket, MCP tools | Passes through, adds the manifest and audit |
| Verifiers | rollout | dataset + harness + rubric (v1 Taskset/Harness/Runtime) | Points its harness at the trainer's model endpoint through a proxy that records tokens and logprobs; maps rubric → judgment |
| Harbor | task | task.toml + Dockerfile + tests | Runs the container, captures the agent's calls through a proxy, maps test results → judgment; branchable:false |
| NeMo Gym | rollout | resource, agent and model servers | Registers the trainer as the model server; maps resource-server verification → judgment |
| ORS | step (tools) | HTTP/SSE, actions are tools, episode rewards | Mostly a direct mapping; rewards → judgment |
| Gymnasium / dm_env | step | numeric or text reset/step | Wraps it as an OpenEnv server; converts spaces to schemas |
| PettingZoo / OpenSpiel | step, multi-agent | agents take turns or act at the same time | Exposes a seat for each agent; participants:n |
Example: a Harbor task, trained in TRL
1 nenyax import harbor://terminal-bench/fix-git
2 The router detects task mode and loads the Harbor driver
3 The audit runs: boots, isolation, test mutation, hack attempts → report
4 The environment is exposed as OpenEnv plus a manifest (branchable:false)
5 TRL's environment_factory opens sessions; the proxy records tokens
6 Tests → judgment → reward → policy update. Can export back to Verifiers or Harbor.
The format list reflects my research as of September 2026. I haven't checked the Google entries (dm_env, OpenSpiel, MuJoCo Playground, AndroidWorld) for current activity. The mapping column is a design proposal and hasn't been implemented.