Example reports

Read the evidence behind a change.

Native agent workflows with retained evidence. Follow controlled task comparisons and a real research pilot from the observed failure to a tested source-handoff fix.

Three case studies.

Native runs · small development samples · explicit limits

01 / OpenSRE

Recovered, then kept acting.

Both incidents recovered safely. One extra action after the accepted report exposed a stopping-contract mismatch.

Safe recovery · before → after
2 / 2 → 2 / 2
Calls after accepted report
1 → 0
Preview of the OpenSRE before-and-after report.

The fix uses OpenSRE’s existing termination signal in our demo binding. Conflicting task cues and an instruction-delivery gap limit attribution; no diagnosis improvement is claimed.

02 / OpenKritt

Found the bug. Checked the evidence.

A tenant-boundary finding came with a reproducible request. We required execution evidence and clarified the reporting contract.

Custom grounding · before → after
0 / 1 → 1 / 1
Severity agrees with demo policy
0 / 1 → 0 / 1
Preview of the OpenKritt before-and-after report.

This combines workflow and contract changes. The result is specific to a hinted synthetic source pair. The original-case scan took longer: 296 → 532 seconds.

03 / GPT Researcher

Found the right lead. Lost it before writing.

The planner saw sources for the expected award winner. Follow-up searches failed, and the writer received evidence about a different award.

Correct target answers · recorded pilot
2 / 3
Offline tests after source-preservation fix
37 passed

Three selected SimpleQA cases, one run each. The tests verify source handoff; they do not establish improved live answer accuracy. Rebuild the report offline from the curated capture.

Scope, evidence and trying it yourself

These are curated report summaries, with one trial per case and candidate. They do not establish general reliability or security accuracy. Raw local captures and the experimental harnesses are not distributed. The public offline examples below demonstrate separate, reproducible tasks without model calls.