OPEN SOURCE / RESEARCH PREVIEW __EVALARC_VERSION__
Look past
the score.
A tool returned an error. The write already happened.
A retry added the same note twice.
Review an agent change before you accept it. Compare failed checks, follow the recorded actions, and keep the evidence behind your decision.
Find the regression ↘Interactive recorded example · no account, install or model key
Run your first local review ↗Partial score. Duplicated write.
Scripted negative control · public seed 17
No model call or live customer data.
QWEN3.8 / ACTUAL MODEL GENERATIONS
Review a model upgrade, case by case.
Qwen3-8B BF16 and Qwen3.8-27B FP8. Eight support-planning cases, three fresh generations each. Inspect every original answer, native pytest check and recorded difference.
Selected public development cases; plans are not executed. Model size and quantization differ.
48 original generations · offline comparison · reproducible checks
Watch the 30-second walkthrough
Four annotated views of saved scripted Docker controls. No audio or new agent run. Transcript and recording method ↗
COMPARE TWO RECORDED REVISIONS
Better score. New failure.
Two closure checks improve.
A previously passing note check fails.
These recorded policies use the same task, seeds, grader and runtime. Select a changed case to compare the observed outcomes.
Loading the recorded comparison…
After
Final ticket state
Reproduce this comparison without running candidates
evalarc compare evalarc-evidence-explorer/comparison/baseline.json \
evalarc-evidence-explorer/comparison/current.json \
--output runs/comparison-001
# Exit 1: a check regressed, despite the higher score.Download the records using the first-review guide and use a fresh output directory. The CLI validates report consistency and matching execution conditions before comparison. It does not authenticate the producer of the reports.
NATIVE STRANDS EVALS REPORTS
Inspect the state rules.
Recheck the same saved states with two native Strands evaluators. The eight-row mean improves from 75% to 87.5%; the notes rule still regresses.
Two equally weighted state rules, separate from EvalArc’s five-dimension score. Saved scripted controls; no new agent or AWS run.
All eight checks · expected and observed state · offline download
FOLLOW THE EVIDENCE
One score never tells the whole story.
Switch the control. Inspect the checks.
Follow each recorded state change.
Loading recorded audit evidence…
Action & result
State changes in this step
Initial and final state
The coding grader checks observable responses and restart behavior. This record contains the failure message and a transcript fingerprint; it does not include the agent's coding process or a full response transcript.
Reproduction record & fingerprints
DECLARE YOUR ACCEPTANCE RULES
Same score. Different gate.
One frozen policy.
Two explicit acceptance rules.
Both support jobs score 93.75% and resolve 0/2 attempts. A deliberately permissive gate accepts the result; requiring every notes check to pass rejects it. The task outcome stays the same.
Loading the recorded suite…
PERMISSIVE / NO REQUIRED DIMENSION
Partial progress allowed
Inspect both attempts ↗PROTECTED / EVERY NOTES CHECK REQUIRED
Notes must be correct
Inspect both attempts ↗Gate acceptance is a configured decision, separate from full task resolution. The permissive gate intentionally accepts an unresolved policy. No score is averaged across coding and support jobs.
Shared candidate, gate decisions & provenance
Recorded Docker suite: three jobs, five attempts, 31 case executions. JUnit distinguishes a rejected gate from an environment error; a hosted CI importer was not exercised. This is scripted development evidence.
Prefer a filterable table or Python? Open the Hugging Face Casebook ↗ and compare gate_accepted with fully_resolved. The original records accompany every row.
REPEAT ONE FROZEN CANDIDATE
One run is not the whole record.
Same cases. Fresh state.
Every attempt stays visible.
Three Docker attempts per scripted control. The reference passes every time; the duplicate-write policy keeps its 93.75% score and fails acceptance every time. Switch controls and open any attempt's full evidence.
Loading the recorded attempts…
Frozen candidate & execution conditions
These v0.4 recordings show fixed public cases, with no observed check variation in either control. They are not stochastic model trials or an estimate of reliability on unseen tasks. The v0.3 comparison below retains its original records and grading fingerprint.
THREE TASK PACKS / DECLARED FAULTS
A perfect score.
How much coverage remains?
A fault detected by one case loses its coverage if that case is removed or weakened. Open each dependency to inspect the recorded checks, status and seed. Margins count distinct case IDs, not repeated runs.
These summaries derive from the original scripted audit records. They are not new agent runs or a guarantee against unseen faults.
YOUR NEXT REVIEW
Bring a change
you need to trust.
Start with two EvalArc evaluations, or prepare a saved AgentCore export using the documented input contract. Review locally, then share a minimal finding.
Export-to-review walkthrough ↗
Bounded import format; no automatic cloud collection. The scored trace controls are synthetic. Live AgentCore evaluation has not been validated by this demo.
FIRST-USE FEEDBACK
Where did the review help?
Tell us what you were checking, where you got stuck, and whether the result changed a decision. A failed setup is useful feedback too.
Share a first-use report ↗Use a minimal redacted example. No production traces, private prompts or credentials are needed.
SAVED TRACE REVIEW
Zero, skipped, or missing?
A zero can be a valid judgment. A skipped evaluator needs context. A required skill may never have loaded. Inspect each against a versioned golden case, with recording identity and explicit acceptance rules.
The five controls are synthetic. The separate MCP record contains a real local load and no evaluator scores. Review your own saved AgentCore Evaluate export offline with evalarc trace-import.
FIXED TRACE, REPEATED JUDGMENTS
Same trace. Same verdict?
Keep the execution fixed. Inspect three saved judgments per target: a changing score, a flipped acceptance decision, and unavailable assessments are different findings.
| Control | Judge run 1 | Judge run 2 | Judge run 3 |
|---|
Agreement is not accuracy: three rejections still reject the case. Agent reruns remain separate in the execution repeatability report.
RECORDED MODEL EVIDENCE
A skill loaded. Did the task pass?
27 real Qwen3-8B trials compare no skill, direct loading and MCP delivery on attributed robot recordings. Inspect every tool receipt, candidate, independent score and ATIF trajectory. Three engineering profiles, including negative results.
L40S recordings; one public development task. No skill accuracy gain or general model ranking is claimed.
ANSWER FILE VS DELIVERED PROGRAM
100% answer reward. 80% program score.
A correct answer file can accompany an incorrect program. Follow three actual Harbor container runs from recorded commands to collected files and independent acceptance.
Scripted controls with native ATIF, including partial credit and a detached answer file. No model inference or unknown-exploit claim.
SAME LENGTH · DIFFERENT GUIDANCE
The agent finished. Did its program work?
Twelve recorded Qwen3-8B attempts compare relevant guidance with unrelated prose, each delivered in a 476-token MCP payload. Inspect an initial protocol failure and a separate diagnostic follow-up, including unchanged starter programs and numerical errors.
Each cohort resolved 0 of 6 tasks. Partial scores, finish signals and task acceptance remain separate; these public development results are not pooled into an efficacy claim.
RETRIEVE · CONTINUE · VERIFY
The history loaded. Did the work advance?
Qwen3-4B continues a public Qwen3-8B program with or without agent-requested Funes MCP retrieval. Six valid retrieval results, six unchanged programs, zero fully resolved tasks. Inspect the cited history, command outputs and independent coordinate checks.
One public task and selected prior session. The earlier pre-injected-context pilot remains separate; repeated-operation counts do not measure time saved.
Continue with a fixed skill version
A separate cohort starts from a prior MCP skill load. Both conditions receive the exact historical skill through workflow-selected MCP; one also offers Funes retrieval. Six preloads and six retrieval results succeed, but all programs remain unchanged and full acceptance is 0/6.
Inspect the pinned skill handoff ↗ · Download the fixed-skill cohort
SOURCE TASK · TOOL TRACE · NATIVE VERDICT
Inspect what the grader actually established.
36 GPU attempts compare four fixed workflows on three public SWE-bench Verified tasks. 31 have assessable native reports; five remain uncertain after upstream infrastructure flags. None obtained acceptance. Follow eight nonempty patches, actual MCP preloads and every failed operation.
Six separate upstream controls expose how a 95% aggregate rule can accept an unresolved required defect. All repetitions are retained; no general skill benefit is claimed.
FINAL FILE · ACTUAL OPERATIONS
The file is correct. Inspect how it was made.
A temporary public file can be written and deleted before the final snapshot. Follow file access and actual service receipts to check the execution against its authorization contract.
Experiment method and results
32 authored native controls cover file operations and service requests. A separate 12-attempt Qwen3-8B pilot on L40S has no fully accepted task; one attempt has incomplete HTTP evidence. Inspect all outcomes and the fixed source records. These results do not establish a skill composition effect.
HAND OFF EVIDENCE
Download the evidence.
Check every gate.
Recompute the original suite configuration, plan, five attempts, acceptance decisions and JUnit without executing a candidate.
Download suite evidence (ZIP) · CLI on PyPI · Install and verify locally ↗
# Unzip the downloaded archive first.
evalarc verify suite-evidence --json
# Require all configured gates to accept:
evalarc verify suite-evidence \
--json --require-acceptedThis example is consistent (exit 0), but one gate rejects it (exit 1 with --require-accepted). Two jobs are accepted; one is fully resolved. No grader rerun or producer authentication is implied.
FROM THE BROWSER TO YOUR TERMINAL
Make the grader
earn your trust.
Run the known-good reference and the declared faulty controls. Keep the outcomes, seeds, runtime limits and fingerprints together.
Installation & execution guide ↗git clone https://github.com/noteflowai/evalarc.git
cd evalarc
python3 -m venv .venv
. .venv/bin/activate
python -m pip install -e .
evalarc audit --task support-routing \
--backend local --trust-local \
--output runs/supportLocal mode runs with your user privileges. Use the documented Docker backend for candidate isolation.
What this evidence establishes
The saved audits detect __FAULT_COUNT__ declared faults across __TASK_PACK_COUNT__ task packs. These are public development tasks and scripted policies. They do not establish frontier-model performance, coverage of arbitrary reward hacks, a human time horizon, or RL training gains. The browser replays saved reports; it does not execute submissions.