Owledge compared 3 completed benchmark runs: 3/3 Owledge profiles passed, privacy failures prevented=3, stale failures prevented=3, average pollution reduction=88.36%, average tokens/correct reduction=83.54%.
| Model | Baseline | Owledge | Pollution reduction | Privacy prevented | Stale prevented | Token reduction | Pass rate | tokens/sec |
|---|---|---|---|---|---|---|---|---|
| gemma4:latest | Fail | Pass | 88.36% | 1 | 1 | 87.15% | 0.8333 | 74.5858 |
| qwen3.5:4b | Fail | Pass | 88.36% | 1 | 1 | 77.66% | 0.8333 | 37.6034 |
| glm-5.1:cloud | Fail | Pass | 88.36% | 1 | 1 | 85.8% | 0.8333 | 89.4848 |
Lower privacy failures, stale failures, Context pollution, and Tokens per correct answer are better. Higher reduction/prevention is better in these charts.
Illustrative API prices per 1M tokens. Verify current provider pricing before using these numbers for budgets. Sources checked: Anthropic Claude pricing (docs.anthropic.com/en/docs/about-claude/pricing), Google Gemini API pricing (ai.google.dev/gemini-api/docs/pricing), and OpenAI API pricing (platform.openai.com/docs/pricing).
This table applies provider API prices to the measured baseline and Owledge token counts. It estimates cost pressure avoided by cleaner context; it is not a full project ROI calculation.
| Provider | Model | Input $/1M | Output $/1M | Baseline cost | Owledge cost | Estimated savings | Savings |
|---|---|---|---|---|---|---|---|
| Anthropic | Claude Opus 4.8 | $5.0 | $25.0 | $0.489345 | $0.396005 | $0.09334 | 19.07% |
| Anthropic | Claude Sonnet 4.6 | $3.0 | $15.0 | $0.293607 | $0.237603 | $0.056004 | 19.07% |
| Anthropic | Claude Haiku 4.5 | $1.0 | $5.0 | $0.097869 | $0.079201 | $0.018668 | 19.07% |
| Gemini 3 Pro | $2.0 | $12.0 | $0.218954 | $0.185236 | $0.033718 | 15.4% | |
| Gemini 2.5 Pro | $1.25 | $10.0 | $0.165866 | $0.149315 | $0.016551 | 9.98% | |
| Gemini 2.5 Flash | $0.3 | $2.5 | $0.040969 | $0.037177 | $0.003792 | 9.26% | |
| OpenAI | gpt-5.5 | $5.0 | $30.0 | $0.547385 | $0.46309 | $0.084295 | 15.4% |
| OpenAI | gpt-5.5-pro | $30.0 | $180.0 | $3.28431 | $2.77854 | $0.50577 | 15.4% |
| OpenAI | gpt-5.4 | $2.5 | $15.0 | $0.273693 | $0.231545 | $0.042148 | 15.4% |
| Model | Scenario | Baseline | Owledge |
|---|---|---|---|
| gemma4:latest | needle | Warn | Pass |
| gemma4:latest | multi-hop | Warn | Pass |
| gemma4:latest | stale-conflict | Warn | Pass |
| gemma4:latest | privacy-trap | Fail | Pass |
| gemma4:latest | distractor-heavy | Warn | Warn |
| gemma4:latest | handoff-resume | Warn | Pass |
| qwen3.5:4b | needle | Warn | Pass |
| qwen3.5:4b | multi-hop | Warn | Pass |
| qwen3.5:4b | stale-conflict | Warn | Pass |
| qwen3.5:4b | privacy-trap | Fail | Pass |
| qwen3.5:4b | distractor-heavy | Warn | Warn |
| qwen3.5:4b | handoff-resume | Warn | Pass |
| glm-5.1:cloud | needle | Warn | Pass |
| glm-5.1:cloud | multi-hop | Warn | Pass |
| glm-5.1:cloud | stale-conflict | Warn | Pass |
| glm-5.1:cloud | privacy-trap | Fail | Pass |
| glm-5.1:cloud | distractor-heavy | Warn | Warn |
| glm-5.1:cloud | handoff-resume | Warn | Pass |
| Baseline | Shows what happens when retrieval over-selects noisy, stale, or private context. |
|---|---|
| Owledge | Shows the product behavior under test: cleaner selected context before model inference. |
| Token reduction | Estimates cost pressure avoided by cleaner context, not total project ROI. |
| tokens/sec | Runtime throughput for the tested model and environment, not an Owledge quality score. |
| Oracle | Ground-truth reference ceiling from the fixture generator. |
Inputs are completed Benchmark Kit reports; this command does not run models. Oracle is ground-truth reference, not a model or product claim. API prices are illustrative snapshots and must be verified against provider pricing before budgeting. Small scale is release proof for v0.7.0.
.owledge\exports\benchmark-kit-nemotron-nano-cloud\latest.json