Metadata-Version: 2.5
Name: rl-envdoctor
Version: 0.1.0
Summary: Adversarial QA for LLM-RL environments: find the reward an empty answer earns.
Project-URL: Homepage, https://github.com/really-notabot/envdoctor
Project-URL: Repository, https://github.com/really-notabot/envdoctor
Project-URL: Issues, https://github.com/really-notabot/envdoctor/issues
Author: Aswin Jose
License-Expression: MIT
License-File: LICENSE
Keywords: ci,evaluation,fuzzing,llm,quality-assurance,reinforcement-learning,reward-hacking,rl-environments,testing,verifiers
Classifier: Development Status :: 3 - Alpha
Classifier: Environment :: Console
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Topic :: Software Development :: Quality Assurance
Classifier: Topic :: Software Development :: Testing
Classifier: Typing :: Typed
Requires-Python: >=3.10
Provides-Extra: all
Requires-Dist: inspect-ai>=0.3; extra == 'all'
Requires-Dist: verifiers>=0.3; extra == 'all'
Provides-Extra: dev
Requires-Dist: hypothesis>=6; extra == 'dev'
Requires-Dist: inspect-ai>=0.3; extra == 'dev'
Requires-Dist: pytest-cov; extra == 'dev'
Requires-Dist: pytest>=8; extra == 'dev'
Requires-Dist: ruff>=0.5; extra == 'dev'
Provides-Extra: inspect
Requires-Dist: inspect-ai>=0.3; extra == 'inspect'
Provides-Extra: verifiers
Requires-Dist: verifiers>=0.3; extra == 'verifiers'
Description-Content-Type: text/markdown

# envdoctor

Adversarial QA for LLM-RL environments. It tries to get paid without doing the task.

An RL environment is a grader, and graders have bugs. The expensive ones are quiet:
a format reward that pays for empty `<answer></answer>` tags, a substring match that
accepts every option at once, a reward function that throws on an empty completion and
gets silently scored 0.0 by the framework. You find out after the policy has spent a
week learning the bug instead of the task.

envdoctor attacks the grader directly with policies that never look at the question, and
grades the results against what a *correct* answer actually earns.

**No model, no API key, no credits.** A scan is a few hundred local function calls.

```bash
pip install rl-envdoctor
envdoctor scan envdoctor.demo::gameable
```

The distribution is `rl-envdoctor`; the command and the import package are `envdoctor`.

```
envdoctor  gameable-mcq
  8 cases x 26 policies
  reward ceiling 1.000 via oracle_answer_tag, payout threshold 0.500

  policy                     kind      mean   max    pay
  -------------------------  --------  -----  -----  ----
  oracle_adaptive            baseline  1.000  1.000  100%
  oracle_corrupted_adaptive  probe     1.000  1.000  100%
  mc_shotgun                 exploit   0.700  0.700  100%
  markdown_skeleton          exploit   0.350  0.700  50%
  copy_prompt                exploit   0.300  0.300  0%
  empty_answer_tag           exploit   0.300  0.300  0%
  empty                      exploit   0.000  0.000  0%
  ...

!! [critical] `mc_shotgun` earns 0.700 mean reward (70% of ceiling)
     Emits every multiple-choice letter, so substring matching always hits.
     repro: envdoctor scan envdoctor.demo::gameable --policy mc_shotgun --n 5 --show-trials

!  [high] `empty_answer_tag` earns 0.300 mean reward (30% of ceiling)
     Emits empty <answer></answer> tags.

!! [critical] Grader accepts corrupted answers (100% of full reward)
     The grader is matching on shape, not on the answer.
```

Exit code is 1 when anything at or above `--fail-on` (default `high`) is found, so
`envdoctor scan` works as a CI gate with no wrapper.

## Does this find anything real?

We scanned 149 environments sampled from the 1,512 public environments on the Prime
Intellect Environments Hub, then had an independent agent re-derive every finding from
the environment's source and try to refute it.

| | count | of |
| --- | --- | --- |
| scanned model-free | 87 | 149 sampled |
| confirmed reward defect | 28 | 87 scanned (32%) |
| a task-blind policy earns at least what a correct completion earns | 16 | 87 scanned (18%) |
| refuted as an envdoctor artifact | 13 | 41 flagged (32%) |

Confirmed root causes: format rewards paid for scaffolding (12), substring and fuzzy
matching (5), constraints that an empty string satisfies vacuously (4), a reward function
that never reads the completion (1), a reward function that raises and is silently scored
0.0 (1).

Three of them, in the authors' own code:

- `bhogan94/q-programming-language` keys its test-case lookup on the raw prompt string,
  which stops matching once prompts are formatted as chat, so a fallback pays a flat 0.5
  to any non-empty output. The constant `"42"` and the reference Q solution both score
  0.500 and the Q interpreter never runs.
- Six `maziyar/OpenMed_*` environments print a worked example ending in `\boxed{B}` in
  their own system prompt. Echoing the prompt is parsed as the answer `B` and collects
  full credit on every row whose gold answer is B.
- `maziyar/OpenMed_MedKnowledge` delivers its answer column as an int, so both correctness
  reward functions raise `AttributeError` on every rollout. The framework swallows that as
  0.0, leaving only content-blind format rewards: an echoed prompt scores 0.275 while a
  correct `\boxed{C}` scores 0.150.

The full write-up, including the 13 cases where this tool was wrong and why, is in
[audit/AUDIT.md](https://github.com/really-notabot/envdoctor/blob/main/audit/AUDIT.md). Per-environment records are in [audit/results/](https://github.com/really-notabot/envdoctor/tree/main/audit/results).

## How well does it work?

Detection quality is a claim, so it is measured. `envdoctor selftest` scores the checks
against a zoo of environments whose defects are known by construction: every defect class
the Hub audit confirmed, plus the healthy shapes most likely to be flagged by mistake.

```
$ envdoctor selftest
ok   pays_for_empty            empty_ties_correct     detected
ok   worked_example_in_prompt  prompt_leaks_answer    detected
ok   dense_but_correct         healthy                quiet
...
defects caught 15/15 (recall 100%), false positives 0/8
```

The healthy half is the half that matters. Building it caught a real false positive: on a
partial-credit grader, an empty `<answer></answer>` tag collects 0.44 of the ceiling from
string similarity alone, which is the metric's floor rather than a defect. Dense graders
now require a policy to take half of what a correct answer earns before it is reported.

## What it looks for

| check | what it means |
| --- | --- |
| `empty_ties_correct` | Nothing outscores a blank page. The strongest thing this tool says. |
| `completion_independent_reward` | The reward never reads the completion, so no policy differs from any other. |
| `exploit_pays_out` | A policy that ignores the task earned reward. |
| `grader_accepts_wrong` | A wrong answer, built by perturbing the right one, still scores. |
| `prompt_leaks_answer` | The prompt hands the answer over, so echoing it is a strategy. |
| `judge_pays_for_tone` | Confident phrasing beats hedged phrasing at the same length and no content. |
| `reward_saturation` | Every completion gets the same reward, so there is no gradient. |
| `nondeterministic_reward` | The same completion scored twice gave two answers. |
| `grader_crash` | The reward function raised. Frameworks swallow this and return 0.0. |
| `invalid_reward` | NaN, infinity, or a reward outside [0, 1]. |
| `slow_grader` | Grading alone is slow enough to pace a training run. |
| `baseline_sanity` | A correct answer scored nothing, so nothing else here is calibrated. |
| `empty_scan` | No cases were graded, so a clean result would mean nothing was tested. |

`envdoctor checks` lists them; `envdoctor policies` lists the attacks.

## Calibration, and why the tool says "uncalibrated"

"The empty string scored 0.3" means nothing on its own. Every scan replays the dataset's
own answer to find the reward ceiling, and reports exploits as a share of it.

Environments state their answer format in the prompt: `<reversed_text>` tags, `\boxed{}`,
`#### 42`. envdoctor reads that instruction out of the prompt and wraps the ground truth
the same way, which is how it calibrates environments whose format it has never seen.

If a correct answer still scores nothing, the run says so rather than reporting a clean
bill of health. That environment probably needs a real rollout, with tools, multiple turns
or code execution, and its grader has not been tested by this scan.

## Scanning a `verifiers` environment

```bash
pip install 'rl-envdoctor[verifiers]'
envdoctor scan gsm8k --n 20
envdoctor scan my-env --env-arg num_examples=200 --env-arg use_think=true
```

Both generations work and the right one is picked automatically. A module exposing
`load_environment()` is scanned through its rubric; a module exposing a `Taskset` is
scanned through its tasks' `@vf.reward` methods, using `Task.score(trace, runtime=None)`.
Either way it is a dataset row in, a synthetic completion graded, a reward out, with no
rollout and no model.

That is why it is free, and it is also the limit. Environments scored during a rollout
come back uncalibrated, and on v1 the rewards that need a container are skipped silently
by the framework, so envdoctor records which ones it could not exercise:

```
skipped_runtime_rewards: ["executes"]
scored_model_free: []
```

An environment in that state has not really been scanned, and the report says so instead
of reporting that nothing paid out.

## Scanning a whole collection

One scan tells you about one environment. A sweep is what turns this into a claim about
a collection, which is how the Hub audit was run:

```bash
envdoctor sweep env_a env_b env_c --json sweep.json
envdoctor sweep --from-file environments.txt --fail-on critical
```

```
target                         worst     findings  ceiling  calibrated  top exploit
-----------------------------  --------  --------  -------  ----------  ----------------
envdoctor.demo::gameable  critical  6         1.0      yes         mc_shotgun 0.700
envdoctor.demo::healthy   clean     0         1.0      yes         -

1 of 2 failed at high or worse
```

A target that fails to load is recorded and the sweep continues, because one broken
environment should not cost you the other forty nine.

## Scanning an Inspect AI task

```bash
pip install 'rl-envdoctor[inspect]'
envdoctor scan mypkg.evals::my_task
```

Inspect scorers have the same failure modes as RL graders, because they are the same
thing under a different name. `includes()` accepts the target anywhere in the output, so
a completion that lists every option satisfies it without answering.

Model-graded scorers are skipped rather than scored, because grading with a model costs
money and is not reproducible. The report names the ones it skipped instead of returning
a zero that looks like a strict grader.

## Scanning anything else

Wrap the grader and scan it in your own test suite:

```python
from envdoctor import Case, FunctionTarget, run_checks, scan

target = FunctionTarget(my_reward_fn, [Case(id="0", prompt=..., answer=...), ...])
report = run_checks(scan(target))
assert not [f for f in report.findings if f.severity == "critical"]
```

Or point the CLI at it: `envdoctor scan mypkg.envs::target`.

## In CI

```yaml
- uses: really-notabot/envdoctor@v0
  with:
    target: my-env
    fail-on: high

- uses: github/codeql-action/upload-sarif@v3
  if: always()
  with:
    sarif_file: envdoctor.sarif
```

The SARIF upload puts each finding in the Security tab and on the pull request diff,
with a stable fingerprint so a finding is tracked across runs rather than reopened every
time the wording changes.

An environment with findings already in it should not fail on day one, or the gate gets
switched off. Record what you know about and fail only on what is new:

```bash
envdoctor scan my-env --json envdoctor-baseline.json   # once, then commit it
envdoctor scan my-env --baseline envdoctor-baseline.json
```

```
baseline: 0 new, 6 known, 1 no longer found
  fixed, drop from the baseline: exploit_pays_out:copy_prompt
```

## In your test suite

Installing the package registers a pytest plugin, so gameability can be an ordinary
assertion next to your other tests:

```python
from envdoctor.pytest_plugin import assert_not_gameable

def test_my_env_is_not_gameable(envdoctor_scan):
    assert_not_gameable(envdoctor_scan("mypkg.envs::target"))

def test_my_reward_function(envdoctor_scan_function):
    report = envdoctor_scan_function(my_reward_fn, my_cases)
    assert_not_gameable(report, fail_on="medium")
```

The failure message carries the repro command for the policy that scored.

## Scanning environments you do not trust

Scanning an environment runs its code. Installing the package runs its build, importing
the module runs everything at module scope, and `load_environment()` or a `Taskset`
constructor runs whatever the author wrote. envdoctor does not sandbox any of that and
cannot: calling the grader is the entire technique.

So treat `envdoctor scan some-hub-env` as equivalent to running an untrusted program,
because it is. For anything off a public hub, run it in a container or a throwaway VM,
with no credentials in the environment. The Hub audit in `audit/` ran every environment
in a subprocess with a timeout, which contains a hang or a crash; it does not contain a
hostile author, and nothing here should be read as claiming otherwise.

## Prior art

Environment quality is a known problem, mostly written up rather than tooled.
[arXiv:2606.16062](https://arxiv.org/abs/2606.16062) found that 28.5% of sampled
SWE-bench-Verified tasks accept incorrect patches; `verifiers` v1 ships a per-task
`validate()` hook for checking that a task's own verifier accepts the gold answer.
envdoctor is the adversarial counterpart, and it is meant to be run on every commit
rather than once for a paper.

## Credits

Built with [Claude Code](https://claude.com/claude-code).

## License

MIT
