Exact predictions
2/15
Meets the current exact-hit gate and matches the prior precision run.
Exact component targeting benchmark
Vaner prepared locally with qwen3.5:35b. Claude Sonnet generated
cold and warm answers; Claude Opus judged prediction relevance and answer quality.
The run improved related coverage from 5/15 to 8/15
and mean relevance from 0.31 to 0.36 versus the prior
precision run, while preserving 2/15 exact hits.
2/15
Meets the current exact-hit gate and matches the prior precision run.
8/15
Up from 5/15; now above the 45% promotion threshold.
0.36
Just over the 0.35 gate; prior precision run was 0.31.
2 / 0 / 1
Vaner wins / ties / cold wins under Claude Opus judging.
Baseline: v2_precision_local5090_sonnet_opus_15turn. Current: v2_exact_targeting_local5090_sonnet_opus_15turn.
| Criterion | Status | Result |
|---|---|---|
| Related rate >= 45% | Pass | 53% |
| Exact rate >= 10% | Pass | 13% |
| Mean score >= 0.35 | Pass | 0.36 |
| Weight drift <= 0.20 | Pass | 0.0432 |
| Source drift <= 0.60 | Pass | 0.2282 |
| Numeric policy drift <= 0.80 | Fail | Above threshold |
Local 35B background preparation per turn, down from 62.2s on the prior precision run.
Higher than the previous 342ms, but still outside answer generation and dominated by local preparation strategy.
No absolute local paths found in the new benchmark artifact or this canvas.