36/36
Exact predictions across 12 turns x 3 idle windows.
Vaner full-scale benchmark
Spark 122B validation moved from a failed broad-core run to 100% exact scenario coverage across short, medium, and long idle windows, with all PR checks green after fixes.
36/36
Exact predictions across 12 turns x 3 idle windows.
0.95
Mean judge score across the 300s long-window run.
28.9
Mean scenarios explored at the 300s budget.
0
Partials or misses in the final thematic sweep.
| Idle window | Human context | Exact | Score | Scenarios |
|---|---|---|---|---|
| 30s | Short read | 100% | 0.9542 | 14.7 |
| 90s | Reading response | 100% | 0.9542 | 15.9 |
| 300s | Testing code | 100% | 0.9500 | 28.9 |
0/12 exact
Initial Spark 122B sweep failed promotion: predictions were broad, related context was weak, and the harness exposed reasoning/token handling issues.
12/12 exact at every budget
Final thematic sweep passed 30s, 90s, and 300s gates with 100% exact prediction and no partials.