Vaner Broad Public Benchmark

Run broad-public-full-20260429T132249Z. Profile: full. Corpora are public fixtures; deterministic-run cost is estimated.

Cases48
Vaner lift vs naked+2.61
Vaner lift vs RAG+0.40

Aggregate Arms

ArmHit@3Recall@5QualityLatency msEstimated cost
naked0.000.004.801300$0.3464
rag0.460.507.011550$0.4808
vaner0.460.507.411800$0.4520

Archetype View

ArchetypeVaner vs nakedVaner vs RAGVaner hit@3
codebase_navigation+2.23+0.400.32
developer_swe+2.81+0.400.55
learner+4.40+0.401.00
researcher+4.40+0.401.00
writer+4.40+0.401.00

Interpretation

Vaner is expected to be strongest when the next task names concrete files, symbols, or workflow state that can be prepared before the answer model starts. It is less strong on one-shot broad questions where naive retrieval already surfaces the same context, or where the right answer depends on patch execution rather than context selection.

The Claude-judged staged run should replace deterministic heuristic quality with blind judge scores and provider usage data where available.