Baseline history
Fixed benchmarks and recent traces are separate populations. Trace scores do not establish controlled improvement between commits.
Settings and connections
Changes apply to this checkout. Credential fields accept environment variable names.
Audit workspace
Loading local audits…
Your first audit
Find failures.
Test improvements.
Start with your existing agent and one concern. Audit its code and available traces, or bring a known failure straight into a measured fix.
ag:auditAudit your agentag:fixImprove your agentAudit can inspect local changes or existing evals. Fix creates or repairs evals and prepares delivery. Run the workflow in your coding agent, then select Refresh.
Fix runs
No fix runs yet.
Invoke ag:fix with a problem, trace or eval request. Measured comparisons and reviewed patches will appear here.
ag:fixEvaluation preparation
Start with useful checks.
Use ag:audit to assess existing evals, or ag:fix to create or improve them. Baseline and deliberately incorrect cases will appear here.
ag:audit · ag:fixAudit overview
Review progress
Issues and recommendations
Ordered by severity, affected traces, and confidence.
Findings awaiting grouping
Scope and limits
Fix run
Result for review
Guide this run
Controls enabledStrategy and candidate decisions apply here. Proposals, directives and continuation wait for your coding agent.
Control history
Queued work waits for your coding agent. Applied controls have taken effect; candidate verification is tracked separately.
Work in progress
Experiment tree
Objective comparison
Each point is one candidate. Select a point or tree entry to inspect its evidence. Every objective remains separate.
Candidate comparisons
Measured values and changes from the baseline. The frontier retains eligible tradeoffs; it does not imply one winner.
