Each MCP server is run through the gauntlet: a live LLM agent attempts generated tasks using only the server's tools, alongside schema, description, security, reliability, and robustness checks. Grade is the weighted overall; a critical security finding caps it.
| # | Server | Grade | Score | Task success | Security | Tools | Scanned |
|---|---|---|---|---|---|---|---|
| 1 | everything | A | 99.9 | 100 | ✓ | 13 | 2026-07-26 |
| 2 | filesystem | A | 99.4 | 100 | ✓ | 14 | 2026-07-26 |
| 3 | good-fixture | A | 99.4 | 100 | ✓ | 3 | 2026-07-26 |
| 4 | git | A | 98.9 | 100 | ✓ | 12 | 2026-07-26 |
| 5 | memory | A | 96.5 | 100 | ✓ | 9 | 2026-07-26 |
| 6 | sqlite | A | 94.2 | 78 | ✓ | 6 | 2026-07-26 |
| 7 | time | A | 92.6 | 73 | ✓ | 2 | 2026-07-26 |
| 8 | bad-fixture | C | 75.0 | 67 | ⚠ | 5 | 2026-07-26 |
| 9 | malicious-fixture | C | 75.0 | 80 | ⚠ | 4 | 2026-07-26 |
Agent model: gemini:gemini-flash-latest · generated 2026-07-26T02:11+00:00 · scores from a live agent are stochastic (repeated and averaged); ⚠ = static tool-poisoning in a description (caps the grade), ⚡ = injection in a live tool output (does not cap).
Scored by mcp-gauntlet 0.4.0. Scoring changes between releases, so compare scores only within the same version — each row's Scanned date is when that score was measured, not when this page was rebuilt.
These servers were not measured the same way as the ranked ones — the agent either never scored them, or stopped partway when a tool hung. Either way their score rests on a different basis: the overall is a weighted mean over the dimensions present, over however many runs completed, so a missing Agent Task Success (the heaviest dimension) or a short sample inflates the number relative to the ranked table. Each row says which applied.
| # | Server | Grade | Score | Task success | Security | Tools | Scanned |
|---|---|---|---|---|---|---|---|
| — | fetch | A | 100.0 | — | ✓ | 1 | 2026-07-26 |
| no read-only tools to test (all excluded as possibly-mutating) | |||||||
| — | sequential-thinking | A | 100.0 | — | ✓ | 1 | 2026-07-26 |
| no read-only tools to test (all excluded as possibly-mutating) | |||||||
Listed here? Paste this into your README — it reads this board's published score and updates whenever the board is regenerated. Swap everything for your server's slug: the filename of its row link, without the .html.
[](https://ghalebdweikat.github.io/mcp-gauntlet/)
Not listed, or think a score is wrong? Open an issue on the repository — re-runs and corrections are free. How these scores are computed, what they are not, and the disclosure policy behind publishing them: see METHODOLOGY.md.