Exact predictions
2/15
Improved from 0/15 before the source-first evidence pass.
Prediction Engine v2 precision benchmark
Vaner prepared locally on the RTX 5090 with qwen3.5:35b.
Claude Sonnet generated the user-facing answers and Claude Opus judged them.
The run confirms the evidence hygiene fix worked: no supporting-only packages.
It also shows the next problem clearly: when Vaner chooses the wrong primary source file,
the warm answer can get worse.
2/15
Improved from 0/15 before the source-first evidence pass.
5/15
Dropped from 6/15; the stricter gate discarded broad docs-only context.
9/15
Source/config/test evidence was present in 60% of turns.
1 / 1 / 1
One Vaner win, one tie, one cold-start win.
The old run used Claude Sonnet as judge; this run used Claude Opus, so compare directionally, not as a perfectly controlled A/B.
| Miss reason | Count | Meaning |
|---|---|---|
| No primary evidence | 6/15 | No source/config/test target was ready. |
| Wrong primary target | 4/15 | Vaner found source files, but not the mechanism the user asked for. |
| Tangential primary evidence | 3/15 | Source evidence was nearby but incomplete. |
| Supporting-only misses | 0/15 | The hygiene/readiness fix removed the old docs/generated/data-only failure. |
Local 35B exploration, serial concurrency, 90s cycle cap.
Warm lookup stayed interactive; preparation cost is background work.
Vaner won 1/3 sampled answer comparisons under Opus judging.