Nenyax: a two-layer protocol on top of OpenEnv

Many environment formats come in at the top, many trainers connect at the bottom, and one narrow contract sits in the middle. OpenEnv stays the standard; Nenyax adds the adapters that route every other format into it, plus the audit that makes the result trustworthy.

Environment sources · any formatport in ↓   port out ↑
OpenEnv Meta · HF (native) Verifiers Prime Intellect Harbor Laude · Terminal-Bench NeMo Gym NVIDIA ORS / OpenReward General Reasoning Gymnasium · PettingZoo Farama dm_env · OpenSpiel Google DeepMind MuJoCo Playground Google DeepMind AndroidWorld Google BrowserGym ServiceNow Reasoning Gym · GEM · SkyRL-Gym HUD MCP-based nenyax build codebases · industries
↓the router reads each source's format and picks a driver
Layer 2 · adapters and routerone driver per format, grouped by how the environment is driven

Step mode

The trainer drives each step with reset / step, or with a tool call per turn.

OpenEnvGymnasiumdm_envOpenSpielPettingZooBrowserGymORSGEM

Rollout mode

The environment runs the whole multi-turn loop against a model endpoint and returns a scored trajectory.

VerifiersNeMo GymReasoning Gym

Task mode

A container holds the task. An agent harness runs inside it, and tests grade the final state.

HarborTerminal-BenchSWE-bench-styleAndroidWorld
Router: manifest.source.format → load the driver → negotiate capabilities → expose the Layer 1 contract. A driver never pretends to support something it can't: a task-mode environment reports branchable:false, and the trainer adapts.
↓every driver exposes the same contract
Layer 1 · Nenyax core contract (OpenEnv plus proposed RFCs)the narrow waist
Manifestidentity, version, licence, format, mode
Capabilitiesresettable, branchable, parallel, risk, real-world
ActionsMCP tools (OpenEnv RFC 003)
Judgmentverifier, rubric or external reward; versioned, with confidence
Trajectoryobservations, actions, tokens and logprobs, judgment
open→sample→act / observe→judge→close  · rollout and task modes run this loop internally and return the same trajectory
↓trainer bindings
Trainers · any backend
TRLverlprime-rlSkyRLNeMo RLUnslothTinkerOpenRLHFARTslime

Audit gates (every environment, ported or built)

  • Builds and boots cleanly; the network is isolated
  • The reference solution passes and deliberately broken solutions fail
  • A red-team agent attempts reward hacks
  • Results are the same when re-run
  • Difficulty is calibrated: a baseline model passes between about 3% and 80%
  • No contamination or leakage, such as a git history that reveals the answer
  • → a standard quality report ships with the environment

Builder skills (for any coding agent)

  • Installed as an MCP server or skill in Claude Code, Codex, Optimus, OpenCode and others
  • build from a repo, an OpenAPI spec, an app or a domain's documents
  • Tools for containerizing, mocking, snapshotting and mining tasks
  • The agent proposes; the audit gates accept or reject
  • One codebase can yield several environments: coding, operator, support, SRE, data, security

How each format maps onto the core contract

FormatModeNative shapeWhat the driver does
OpenEnvstepreset/step over HTTP/WebSocket, MCP toolsPasses through, adds the manifest and audit
Verifiersrolloutdataset + harness + rubric (v1 Taskset/Harness/Runtime)Points its harness at the trainer's model endpoint through a proxy that records tokens and logprobs; maps rubric → judgment
Harbortasktask.toml + Dockerfile + testsRuns the container, captures the agent's calls through a proxy, maps test results → judgment; branchable:false
NeMo Gymrolloutresource, agent and model serversRegisters the trainer as the model server; maps resource-server verification → judgment
ORSstep (tools)HTTP/SSE, actions are tools, episode rewardsMostly a direct mapping; rewards → judgment
Gymnasium / dm_envstepnumeric or text reset/stepWraps it as an OpenEnv server; converts spaces to schemas
PettingZoo / OpenSpielstep, multi-agentagents take turns or act at the same timeExposes a seat for each agent; participants:n

Example: a Harbor task, trained in TRL

1 nenyax import harbor://terminal-bench/fix-git
2 The router detects task mode and loads the Harbor driver
3 The audit runs: boots, isolation, test mutation, hack attempts → report
4 The environment is exposed as OpenEnv plus a manifest (branchable:false)
5 TRL's environment_factory opens sessions; the proxy records tokens
6 Tests → judgment → reward → policy update. Can export back to Verifiers or Harbor.

The format list reflects my research as of September 2026. I haven't checked the Google entries (dm_env, OpenSpiel, MuJoCo Playground, AndroidWorld) for current activity. The mapping column is a design proposal and hasn't been implemented.