0.8.8 release validation ยท public canvas

Vaner is useful when it has the right context. Coverage is still the bottleneck.

This run combines the 0.8.8 scenario benchmark, a broad public naked/RAG/Vaner comparison, and the deep-run maturation suite. It uses public corpora or public-layout fixtures and reports aggregate benchmark evidence without exposing local workspace details.

Run date
2026-04-29
Models
Vaner: ollama:qwen3.5:35b
Answer: claude:sonnet
Judge: claude:opus
20/20
Scenario related coverage
100% related, 85% exact; promotion gate passed.
5/8
Quality A/B
Vaner-context answers won 62% of judged pairs.
7.41
Broad public quality
+2.61 vs naked, +0.40 vs RAG across 48 cases.
+0.844
Deep-run delta
3/5 gates passed; judge agreement and persistence band failed.

Layer Results

LayerBenchmarkResultRead
Prediction / exact targeting20 Vaner repo mechanism queries17 exact, 3 partial, 0 no-scenario cyclesExact targeting is now strong on named components; remaining partials are mostly broad process questions.
Answer quality8 Claude-judged pairs5 Vaner wins, 3 cold wins, 0 tiesWhen useful context is present, the answer improves; weak/empty context can lose.
Naked vs RAG vs Vaner48 public casesVaner quality 7.41 vs RAG 7.01 vs naked 4.80Vaner beats naked and edges RAG in this deterministic public harness, with similar hit/recall and lower estimated cost than RAG.
Deep-run maturation40 sessions / 40 outcomesMean improvement +0.844; stale rate 0%Strong synthetic maturation signal, but the current fixture judge has no external agreement signal and persistence is too high.
Agent-with-tools layerHarness restored and CLI import validatedNot included as a scored release claimNeeds a configured OpenAI-compatible tool-calling answer endpoint and running Vaner daemon; not faked in this pass.

Answer The Question

1
When Vaner is betterBest on codebase tasks where the needed object/path is present in prepared context. The quality judge favored Vaner in 5 of 8 paired scenario answers, and the broad public harness shows +2.61 quality over naked and +0.40 over RAG.
2
Why it worksIt front-loads context selection before the final answer: exact component hints, symbol/path signals, and cached packages reduce answer-time search and can correct cold-answer misunderstandings.
3
Where it is weakerWeakness is precision under broad lifecycle prompts: 3 of 20 turns were only partial, with cold-start still under-specified. No turns lacked relevant context.
4
Cost and performanceScenario precompute averaged 10.2s; cache lookup averaged 2.1s. Warm answer latency was 119.1s vs 168.4s cold in the judged sample. Broad public estimated cost was $0.452 for Vaner vs $0.481 for RAG and $0.346 naked.

Scenario Gate

Engine LLMwired via local Ollama
Mean relevance0.81
No-scenario cycles0/20 (0%)
Gate status5/5 passed; PASS

Broad Public Arms

ArmHit@3Recall@5QualityCost
Naked0%0%4.80$0.346
RAG46%50%7.01$0.481
Vaner46%50%7.41$0.452

Eval/Fix Loop

A
Promote only with gatesSeparate coverage, quality, latency, cost, and leak gates. Failed gates become tracked fixes, not rewritten conclusions.
B
Debug by failure sliceFirst slice for 0.8.9: named-component queries that produced zero scenarios or weak primary evidence.
C
Publish clean evidencePublic reports should include methodology, aggregate results, limitations, and leak scans without naming local paths or internal workspaces.

Broad Public Archetypes

ArchetypeCasesVaner hit@3Vaner qualityLift
codebase navigation2532%7.030.40 vs RAG
developer swe2055%7.610.40 vs RAG
learner1100%9.200.40 vs RAG
researcher1100%9.200.40 vs RAG
writer1100%9.200.40 vs RAG

Deep-Run Archetypes

ArchetypeMean improvement delta
developer0.853
planner0.837
researcher0.861
writer0.826