End-to-end latency
Retrieval plus answer generation, in seconds
Live web retrieval benchmark · 28 August 2026
Kestrel’s three-provider fanout delivered nearly the same answer quality as native Web Search—with lower median latency and substantially fewer tokens.
The headline
Retrieval plus answer generation, in seconds
Input + output + reasoning tokens
Transparency
Every chart above is computed in your browser from the 48 embedded trial records below.
| Arm | Task | Trial | Latency | Tokens | Score | Result | Failure stage |
|---|
Both arms received the same eight prompts and ran three times per task. Kestrel fanout used DuckDuckGo, Bing, and Yahoo concurrently.
Quality uses semantic review, not literal keyword checks. A pass requires ≥6/8 plus minimum correctness and grounding thresholds.
This is a small live-web benchmark, not a universal claim. Results can vary as providers, sources, and models change. The partial fallback run is excluded.