Example reports

Read the evidence behind a change.

Two native agent workflows on controlled tasks. Follow the observed behavior, a change selected after reviewing the results, and what the rerun showed.

Two case studies. Four reports.

Native runs · synthetic tasks · small development samples

01 / OpenSRE

Recovered, then kept acting.

Both incidents recovered safely. One extra action after the accepted report exposed a stopping-contract mismatch.

Safe recovery · before → after
2 / 2 → 2 / 2
Calls after accepted report
1 → 0
Preview of the OpenSRE before-and-after report.

The fix uses OpenSRE’s existing termination signal in our demo binding. Conflicting task cues and an instruction-delivery gap limit attribution; no diagnosis improvement is claimed.

02 / OpenKritt

Found the bug. Checked the evidence.

A tenant-boundary finding came with a reproducible request. We required execution evidence and clarified the reporting contract.

Custom grounding · before → after
0 / 1 → 1 / 1
Severity agrees with demo policy
0 / 1 → 0 / 1
Preview of the OpenKritt before-and-after report.

This combines workflow and contract changes. The result is specific to a hinted synthetic source pair. The original-case scan took longer: 296 → 532 seconds.

Scope, evidence and trying it yourself

These are curated report summaries, with one trial per case and candidate. They do not establish general reliability or security accuracy. Raw local captures and the experimental harnesses are not distributed. The public offline examples below demonstrate separate, reproducible tasks without model calls.