Owledge Benchmark Comparison Report

Executive Verdict
Pass

Owledge compared 3 completed benchmark runs: 3/3 Owledge profiles passed, privacy failures prevented=3, stale failures prevented=3, average pollution reduction=88.36%, average tokens/correct reduction=83.54%.

Creator Pull Quote
Owledge makes the context safer and cheaper before the model sees it.
Models compared
3
Privacy failures prevented
3
Stale failures prevented
3
Avg pollution reduction
88.36%
Avg tokens/correct reduction
83.54%

Model Matrix

ModelBaselineOwledgePollution reductionPrivacy preventedStale preventedToken reductionPass ratetokens/sec
gemma4:latestFailPass88.36%1187.15%0.833374.5858
qwen3.5:4bFailPass88.36%1177.66%0.833337.6034
glm-5.1:cloudFailPass88.36%1185.8%0.833389.4848

Before vs Owledge Charts

Lower privacy failures, stale failures, Context pollution, and Tokens per correct answer are better. Higher reduction/prevention is better in these charts.

Before vs Owledge Comparison Higher reduction/prevention is better. tokens/sec is shown in the model matrix. gemma4:latest - Pollution reduction % 88.36 gemma4:latest - Token reduction % 87.15 gemma4:latest - Owledge tokens/sec 74.5858 gemma4:latest - Privacy prevented 1.0 gemma4:latest - Stale prevented 1.0 qwen3.5:4b - Pollution reduction % 88.36 qwen3.5:4b - Token reduction % 77.66 qwen3.5:4b - Owledge tokens/sec 37.6034 qwen3.5:4b - Privacy prevented 1.0 qwen3.5:4b - Stale prevented 1.0 glm-5.1:cloud - Pollution reduction % 88.36 glm-5.1:cloud - Token reduction % 85.8 glm-5.1:cloud - Owledge tokens/sec 89.4848 glm-5.1:cloud - Privacy prevented 1.0 glm-5.1:cloud - Stale prevented 1.0

Estimated API Cost Impact

Illustrative API prices per 1M tokens. Verify current provider pricing before using these numbers for budgets. Sources checked: Anthropic Claude pricing (docs.anthropic.com/en/docs/about-claude/pricing), Google Gemini API pricing (ai.google.dev/gemini-api/docs/pricing), and OpenAI API pricing (platform.openai.com/docs/pricing).

This table applies provider API prices to the measured baseline and Owledge token counts. It estimates cost pressure avoided by cleaner context; it is not a full project ROI calculation.

ProviderModelInput $/1MOutput $/1MBaseline costOwledge costEstimated savingsSavings
AnthropicClaude Opus 4.8$5.0$25.0$0.489345$0.396005$0.0933419.07%
AnthropicClaude Sonnet 4.6$3.0$15.0$0.293607$0.237603$0.05600419.07%
AnthropicClaude Haiku 4.5$1.0$5.0$0.097869$0.079201$0.01866819.07%
GoogleGemini 3 Pro$2.0$12.0$0.218954$0.185236$0.03371815.4%
GoogleGemini 2.5 Pro$1.25$10.0$0.165866$0.149315$0.0165519.98%
GoogleGemini 2.5 Flash$0.3$2.5$0.040969$0.037177$0.0037929.26%
OpenAIgpt-5.5$5.0$30.0$0.547385$0.46309$0.08429515.4%
OpenAIgpt-5.5-pro$30.0$180.0$3.28431$2.77854$0.5057715.4%
OpenAIgpt-5.4$2.5$15.0$0.273693$0.231545$0.04214815.4%

Scenario Heatmap

ModelScenarioBaselineOwledge
gemma4:latestneedleWarnPass
gemma4:latestmulti-hopWarnPass
gemma4:lateststale-conflictWarnPass
gemma4:latestprivacy-trapFailPass
gemma4:latestdistractor-heavyWarnWarn
gemma4:latesthandoff-resumeWarnPass
qwen3.5:4bneedleWarnPass
qwen3.5:4bmulti-hopWarnPass
qwen3.5:4bstale-conflictWarnPass
qwen3.5:4bprivacy-trapFailPass
qwen3.5:4bdistractor-heavyWarnWarn
qwen3.5:4bhandoff-resumeWarnPass
glm-5.1:cloudneedleWarnPass
glm-5.1:cloudmulti-hopWarnPass
glm-5.1:cloudstale-conflictWarnPass
glm-5.1:cloudprivacy-trapFailPass
glm-5.1:clouddistractor-heavyWarnWarn
glm-5.1:cloudhandoff-resumeWarnPass

How To Read This Report

BaselineShows what happens when retrieval over-selects noisy, stale, or private context.
OwledgeShows the product behavior under test: cleaner selected context before model inference.
Token reductionEstimates cost pressure avoided by cleaner context, not total project ROI.
tokens/secRuntime throughput for the tested model and environment, not an Owledge quality score.
OracleGround-truth reference ceiling from the fixture generator.

Caveats

Inputs are completed Benchmark Kit reports; this command does not run models. Oracle is ground-truth reference, not a model or product claim. API prices are illustrative snapshots and must be verified against provider pricing before budgeting. Small scale is release proof for v0.7.0.

Skipped Inputs