0.8.8 release validation ยท public canvas

Exact targeting is strong. Broad lifecycle prompts are the remaining bottleneck.

This run combines the 0.8.8 scenario benchmark, a broad public naked/RAG/Vaner comparison, and the deep-run maturation suite. It uses public corpora or public-layout fixtures and reports aggregate benchmark evidence without exposing local workspace details.

Run date
2026-04-30
Models
Vaner: openai:Qwen/Qwen3.6-35B-A3B-FP8
Answer: claude:sonnet
Judge: claude:opus
20/20
Scenario related coverage
100% related, 90% exact; promotion gate passed.
6/8
Quality A/B
Vaner-context answers won 75% of judged pairs.
8.13
Broad public quality
+3.33 vs naked, +0.43 vs RAG across 125 cases.
+0.844
Deep-run delta
effective pass; raw 3/5 gates passed with scaffold-only non-blocking gates

Layer Results

LayerBenchmarkResultRead
Prediction / exact targeting20 Vaner repo mechanism queries18 exact, 2 partial, 0 no-scenario cyclesExact targeting is now strong on named components; remaining partials are mostly broad process questions.
Answer quality8 Claude-judged pairs6 Vaner wins, 2 cold wins, 0 tiesWhen useful context is present, the answer improves; weak/empty context can lose.
Naked vs RAG vs Vaner125 public casesVaner quality 8.13 vs RAG 7.70 vs naked 4.80Vaner beats naked and edges RAG in this deterministic public harness, with similar hit/recall and lower estimated cost than RAG.
Deep-run maturation40 sessions / 40 outcomesMean improvement +0.844; stale rate 0%Strong reference-ceiling maturation signal; raw anti-self-judging gates remain visible and are non-blocking only in this scaffold mode.
Agent-with-tools layerHarness restored and CLI import validatedNot included as a scored release claimNeeds a configured OpenAI-compatible tool-calling answer endpoint and running Vaner daemon; not faked in this pass.

Answer The Question

1
When Vaner is betterBest on codebase tasks where the needed object/path is present in prepared context. The quality judge favored Vaner in 6 of 8 paired scenario answers, and the broad public harness shows +3.33 quality over naked and +0.43 over RAG.
2
Why it worksIt front-loads context selection before the final answer: exact component hints, symbol/path signals, and cached packages reduce answer-time search and can correct cold-answer misunderstandings.
3
Where it is weakerWeakness is precision under broad lifecycle prompts: 2 of 20 turns were only partial, with cold-start still under-specified. No turns lacked relevant context.
4
Cost and performanceScenario precompute averaged 71.4s; cache lookup averaged 2.2s. Warm answer latency was 63.9s vs 171.4s cold in the judged sample. Broad public estimated cost was $1.155 for Vaner vs $1.230 for RAG and $0.883 naked.

Scenario Gate

Engine LLMwired via openai:Qwen/Qwen3.6-35B-A3B-FP8
Mean relevance0.86
No-scenario cycles0/20 (0%)
Gate status5/5 passed; PASS

Broad Public Arms

ArmHit@3Recall@5QualityCost
Naked0%0%4.80$0.883
RAG66%70%7.70$1.230
Vaner67%71%8.13$1.155

Eval/Fix Loop

A
Promote only with gatesSeparate coverage, quality, latency, cost, and leak gates. Failed gates become tracked fixes, not rewritten conclusions.
B
Debug by failure sliceFirst slice for 0.8.9: named-component queries that produced zero scenarios or weak primary evidence.
C
Publish clean evidencePublic reports should include methodology, aggregate results, limitations, and leak scans without naming local paths or internal workspaces.

Broad Public Archetypes

ArchetypeCasesVaner hit@3Vaner qualityLift
codebase navigation2532%7.030.40 vs RAG
developer swe2560%7.810.42 vs RAG
learner2568%8.210.40 vs RAG
researcher2580%8.520.54 vs RAG
writer2596%9.060.40 vs RAG

Deep-Run Archetypes

ArchetypeMean improvement delta
developer0.853
planner0.837
researcher0.861
writer0.826