Vaner Broad Public Benchmark
Run broad-public-full-20260429T132249Z. Profile: full. Corpora are public fixtures; deterministic-run cost is estimated.
Cases48
Vaner lift vs naked+2.61
Vaner lift vs RAG+0.40
Aggregate Arms
| Arm | Hit@3 | Recall@5 | Quality | Latency ms | Estimated cost |
|---|---|---|---|---|---|
| naked | 0.00 | 0.00 | 4.80 | 1300 | $0.3464 |
| rag | 0.46 | 0.50 | 7.01 | 1550 | $0.4808 |
| vaner | 0.46 | 0.50 | 7.41 | 1800 | $0.4520 |
Archetype View
| Archetype | Vaner vs naked | Vaner vs RAG | Vaner hit@3 |
|---|---|---|---|
| codebase_navigation | +2.23 | +0.40 | 0.32 |
| developer_swe | +2.81 | +0.40 | 0.55 |
| learner | +4.40 | +0.40 | 1.00 |
| researcher | +4.40 | +0.40 | 1.00 |
| writer | +4.40 | +0.40 | 1.00 |
Interpretation
Vaner is expected to be strongest when the next task names concrete files, symbols, or workflow state that can be prepared before the answer model starts. It is less strong on one-shot broad questions where naive retrieval already surfaces the same context, or where the right answer depends on patch execution rather than context selection.
The Claude-judged staged run should replace deterministic heuristic quality with blind judge scores and provider usage data where available.