0.8.8 release validation ยท public canvas
Exact targeting is strong. Broad lifecycle prompts are the remaining bottleneck.
This run combines the 0.8.8 scenario benchmark, a broad public naked/RAG/Vaner comparison, and the deep-run maturation suite. It uses public corpora or public-layout fixtures and reports aggregate benchmark evidence without exposing local workspace details.
Run date
2026-04-30
Models
Vaner:
Answer:
Judge:
openai:Qwen/Qwen3.6-35B-A3B-FP8Answer:
claude:sonnetJudge:
claude:opus20/20
Scenario related coverage
100% related, 90% exact; promotion gate passed.
100% related, 90% exact; promotion gate passed.
6/8
Quality A/B
Vaner-context answers won 75% of judged pairs.
Vaner-context answers won 75% of judged pairs.
8.13
Broad public quality
+3.33 vs naked, +0.43 vs RAG across 125 cases.
+3.33 vs naked, +0.43 vs RAG across 125 cases.
+0.844
Deep-run delta
effective pass; raw 3/5 gates passed with scaffold-only non-blocking gates
effective pass; raw 3/5 gates passed with scaffold-only non-blocking gates
Layer Results
| Layer | Benchmark | Result | Read |
|---|---|---|---|
| Prediction / exact targeting | 20 Vaner repo mechanism queries | 18 exact, 2 partial, 0 no-scenario cycles | Exact targeting is now strong on named components; remaining partials are mostly broad process questions. |
| Answer quality | 8 Claude-judged pairs | 6 Vaner wins, 2 cold wins, 0 ties | When useful context is present, the answer improves; weak/empty context can lose. |
| Naked vs RAG vs Vaner | 125 public cases | Vaner quality 8.13 vs RAG 7.70 vs naked 4.80 | Vaner beats naked and edges RAG in this deterministic public harness, with similar hit/recall and lower estimated cost than RAG. |
| Deep-run maturation | 40 sessions / 40 outcomes | Mean improvement +0.844; stale rate 0% | Strong reference-ceiling maturation signal; raw anti-self-judging gates remain visible and are non-blocking only in this scaffold mode. |
| Agent-with-tools layer | Harness restored and CLI import validated | Not included as a scored release claim | Needs a configured OpenAI-compatible tool-calling answer endpoint and running Vaner daemon; not faked in this pass. |
Answer The Question
1
When Vaner is betterBest on codebase tasks where the needed object/path is present in prepared context. The quality judge favored Vaner in 6 of 8 paired scenario answers, and the broad public harness shows +3.33 quality over naked and +0.43 over RAG.
2
Why it worksIt front-loads context selection before the final answer: exact component hints, symbol/path signals, and cached packages reduce answer-time search and can correct cold-answer misunderstandings.
3
Where it is weakerWeakness is precision under broad lifecycle prompts: 2 of 20 turns were only partial, with cold-start still under-specified. No turns lacked relevant context.
4
Cost and performanceScenario precompute averaged 71.4s; cache lookup averaged 2.2s. Warm answer latency was 63.9s vs 171.4s cold in the judged sample. Broad public estimated cost was $1.155 for Vaner vs $1.230 for RAG and $0.883 naked.
Scenario Gate
| Engine LLM | wired via openai:Qwen/Qwen3.6-35B-A3B-FP8 |
| Mean relevance | 0.86 |
| No-scenario cycles | 0/20 (0%) |
| Gate status | 5/5 passed; PASS |
Broad Public Arms
| Arm | Hit@3 | Recall@5 | Quality | Cost |
|---|---|---|---|---|
| Naked | 0% | 0% | 4.80 | $0.883 |
| RAG | 66% | 70% | 7.70 | $1.230 |
| Vaner | 67% | 71% | 8.13 | $1.155 |
Eval/Fix Loop
A
Promote only with gatesSeparate coverage, quality, latency, cost, and leak gates. Failed gates become tracked fixes, not rewritten conclusions.
B
Debug by failure sliceFirst slice for 0.8.9: named-component queries that produced zero scenarios or weak primary evidence.
C
Publish clean evidencePublic reports should include methodology, aggregate results, limitations, and leak scans without naming local paths or internal workspaces.
Broad Public Archetypes
| Archetype | Cases | Vaner hit@3 | Vaner quality | Lift |
|---|---|---|---|---|
| codebase navigation | 25 | 32% | 7.03 | 0.40 vs RAG |
| developer swe | 25 | 60% | 7.81 | 0.42 vs RAG |
| learner | 25 | 68% | 8.21 | 0.40 vs RAG |
| researcher | 25 | 80% | 8.52 | 0.54 vs RAG |
| writer | 25 | 96% | 9.06 | 0.40 vs RAG |
Deep-Run Archetypes
| Archetype | Mean improvement delta |
|---|---|
| developer | 0.853 |
| planner | 0.837 |
| researcher | 0.861 |
| writer | 0.826 |