Live web retrieval benchmark · 28 August 2026

Search,
measured.

Kestrel’s three-provider fanout delivered nearly the same answer quality as native Web Search—with lower median latency and substantially fewer tokens.

8 research tasks 3 trials per task 24 trials per arm Live, uncached retrieval

The headline

Comparable quality. Leaner runs.

Native Kestrel

End-to-end latency

Retrieval plus answer generation, in seconds

Total model tokens

Input + output + reasoning tokens

Transparency

Take the numbers with you.

Every chart above is computed in your browser from the 48 embedded trial records below.

View all trial-level data
ArmTaskTrialLatencyTokensScoreResultFailure stage

Matched setup

Both arms received the same eight prompts and ran three times per task. Kestrel fanout used DuckDuckGo, Bing, and Yahoo concurrently.

Human-readable scoring

Quality uses semantic review, not literal keyword checks. A pass requires ≥6/8 plus minimum correctness and grounding thresholds.

Read with care

This is a small live-web benchmark, not a universal claim. Results can vary as providers, sources, and models change. The partial fallback run is excluded.