0.8.8 release validation ยท public canvas
Vaner is useful when it has the right context. Coverage is still the bottleneck.
This run combines the 0.8.8 scenario benchmark, a broad public naked/RAG/Vaner comparison, and the deep-run maturation suite. It uses public corpora or public-layout fixtures and reports aggregate benchmark evidence without exposing local workspace details.
Run date
2026-04-29
Models
Vaner:
Answer:
Judge:
ollama:qwen3.5:35bAnswer:
claude:sonnetJudge:
claude:opus20/20
Scenario related coverage
100% related, 85% exact; promotion gate passed.
100% related, 85% exact; promotion gate passed.
5/8
Quality A/B
Vaner-context answers won 62% of judged pairs.
Vaner-context answers won 62% of judged pairs.
7.41
Broad public quality
+2.61 vs naked, +0.40 vs RAG across 48 cases.
+2.61 vs naked, +0.40 vs RAG across 48 cases.
+0.844
Deep-run delta
3/5 gates passed; judge agreement and persistence band failed.
3/5 gates passed; judge agreement and persistence band failed.
Layer Results
| Layer | Benchmark | Result | Read |
|---|---|---|---|
| Prediction / exact targeting | 20 Vaner repo mechanism queries | 17 exact, 3 partial, 0 no-scenario cycles | Exact targeting is now strong on named components; remaining partials are mostly broad process questions. |
| Answer quality | 8 Claude-judged pairs | 5 Vaner wins, 3 cold wins, 0 ties | When useful context is present, the answer improves; weak/empty context can lose. |
| Naked vs RAG vs Vaner | 48 public cases | Vaner quality 7.41 vs RAG 7.01 vs naked 4.80 | Vaner beats naked and edges RAG in this deterministic public harness, with similar hit/recall and lower estimated cost than RAG. |
| Deep-run maturation | 40 sessions / 40 outcomes | Mean improvement +0.844; stale rate 0% | Strong synthetic maturation signal, but the current fixture judge has no external agreement signal and persistence is too high. |
| Agent-with-tools layer | Harness restored and CLI import validated | Not included as a scored release claim | Needs a configured OpenAI-compatible tool-calling answer endpoint and running Vaner daemon; not faked in this pass. |
Answer The Question
1
When Vaner is betterBest on codebase tasks where the needed object/path is present in prepared context. The quality judge favored Vaner in 5 of 8 paired scenario answers, and the broad public harness shows +2.61 quality over naked and +0.40 over RAG.
2
Why it worksIt front-loads context selection before the final answer: exact component hints, symbol/path signals, and cached packages reduce answer-time search and can correct cold-answer misunderstandings.
3
Where it is weakerWeakness is precision under broad lifecycle prompts: 3 of 20 turns were only partial, with cold-start still under-specified. No turns lacked relevant context.
4
Cost and performanceScenario precompute averaged 10.2s; cache lookup averaged 2.1s. Warm answer latency was 119.1s vs 168.4s cold in the judged sample. Broad public estimated cost was $0.452 for Vaner vs $0.481 for RAG and $0.346 naked.
Scenario Gate
| Engine LLM | wired via local Ollama |
| Mean relevance | 0.81 |
| No-scenario cycles | 0/20 (0%) |
| Gate status | 5/5 passed; PASS |
Broad Public Arms
| Arm | Hit@3 | Recall@5 | Quality | Cost |
|---|---|---|---|---|
| Naked | 0% | 0% | 4.80 | $0.346 |
| RAG | 46% | 50% | 7.01 | $0.481 |
| Vaner | 46% | 50% | 7.41 | $0.452 |
Eval/Fix Loop
A
Promote only with gatesSeparate coverage, quality, latency, cost, and leak gates. Failed gates become tracked fixes, not rewritten conclusions.
B
Debug by failure sliceFirst slice for 0.8.9: named-component queries that produced zero scenarios or weak primary evidence.
C
Publish clean evidencePublic reports should include methodology, aggregate results, limitations, and leak scans without naming local paths or internal workspaces.
Broad Public Archetypes
| Archetype | Cases | Vaner hit@3 | Vaner quality | Lift |
|---|---|---|---|---|
| codebase navigation | 25 | 32% | 7.03 | 0.40 vs RAG |
| developer swe | 20 | 55% | 7.61 | 0.40 vs RAG |
| learner | 1 | 100% | 9.20 | 0.40 vs RAG |
| researcher | 1 | 100% | 9.20 | 0.40 vs RAG |
| writer | 1 | 100% | 9.20 | 0.40 vs RAG |
Deep-Run Archetypes
| Archetype | Mean improvement delta |
|---|---|
| developer | 0.853 |
| planner | 0.837 |
| researcher | 0.861 |
| writer | 0.826 |