mcp-gauntlet leaderboard

Each MCP server is run through the gauntlet: a live LLM agent attempts generated tasks using only the server's tools, alongside schema, description, security, reliability, and robustness checks. Grade is the weighted overall; a critical security finding caps it.

#ServerGradeScoreTask successSecurityToolsScanned
1everythingA99.9100132026-07-26
2filesystemA99.4100142026-07-26
3good-fixtureA99.410032026-07-26
4gitA98.9100122026-07-26
5memoryA96.510092026-07-26
6sqliteA94.27862026-07-26
7timeA92.67322026-07-26
8bad-fixtureC75.06752026-07-26
9malicious-fixtureC75.08042026-07-26

Agent model: gemini:gemini-flash-latest · generated 2026-07-26T02:11+00:00 · scores from a live agent are stochastic (repeated and averaged); ⚠ = static tool-poisoning in a description (caps the grade), ⚡ = injection in a live tool output (does not cap).

Scored by mcp-gauntlet 0.4.0. Scoring changes between releases, so compare scores only within the same version — each row's Scanned date is when that score was measured, not when this page was rebuilt.

Partially evaluated not comparable

These servers were not measured the same way as the ranked ones — the agent either never scored them, or stopped partway when a tool hung. Either way their score rests on a different basis: the overall is a weighted mean over the dimensions present, over however many runs completed, so a missing Agent Task Success (the heaviest dimension) or a short sample inflates the number relative to the ranked table. Each row says which applied.

#ServerGradeScoreTask successSecurityToolsScanned
fetchA100.012026-07-26
no read-only tools to test (all excluded as possibly-mutating)
sequential-thinkingA100.012026-07-26
no read-only tools to test (all excluded as possibly-mutating)

Add your badge

Listed here? Paste this into your README — it reads this board's published score and updates whenever the board is regenerated. Swap everything for your server's slug: the filename of its row link, without the .html.

[![mcp-gauntlet](https://img.shields.io/endpoint?url=https://ghalebdweikat.github.io/mcp-gauntlet/badges/everything.json)](https://ghalebdweikat.github.io/mcp-gauntlet/)

Not listed, or think a score is wrong? Open an issue on the repository — re-runs and corrections are free. How these scores are computed, what they are not, and the disclosure policy behind publishing them: see METHODOLOGY.md.