Vaner full-scale benchmark

Predictive context is back above the release bar.

Spark 122B validation moved from a failed broad-core run to 100% exact scenario coverage across short, medium, and long idle windows, with all PR checks green after fixes.

36/36

Exact predictions across 12 turns x 3 idle windows.

0.95

Mean judge score across the 300s long-window run.

28.9

Mean scenarios explored at the 300s budget.

0

Partials or misses in the final thematic sweep.

Time-Budget Sweep

Idle window Human context Exact Score Scenarios
30s Short read 100% 0.9542 14.7
90s Reading response 100% 0.9542 15.9
300s Testing code 100% 0.9500 28.9

What Changed

1
Thematic core seeding Frontier, cache, reward, scorer, reasoner, MCP, LLM, store, daemon, and broker areas are explicitly represented.
2
Better source ranking Prompt terms now split code-style identifiers and favor relevant source paths over incidental docs/tests.
3
Structured model wiring String model specs now use structured clients, token clamps, and reasoning-off behavior reliably.
4
Cockpit restored The current cockpit renders with the newer layout and passes build, unit, and accessibility checks.
Before fixes

0/12 exact

Initial Spark 122B sweep failed promotion: predictions were broad, related context was weak, and the harness exposed reasoning/token handling issues.

After fixes

12/12 exact at every budget

Final thematic sweep passed 30s, 90s, and 300s gates with 100% exact prediction and no partials.