mcp-gauntlet leaderboard

Each MCP server is run through the gauntlet: a live LLM agent attempts generated tasks using only the server's tools, alongside schema, description, security, reliability, and robustness checks. Grade is the weighted overall; a critical security finding caps it.

#ServerGradeScoreTask successSecurityTools
1good-fixtureA99.41003
2sqliteA95.5836
3memoryA95.4839
4timeA95.2832
5everythingB89.46713
6bad-fixtureC75.0675
7gitC72.51712
8filesystemD67.3014

Agent model: gemini:gemini-flash-latest · generated 2026-07-24T20:31+00:00 · scores from a live agent are stochastic (repeated and averaged); ⚠ = static tool-poisoning in a description (caps the grade), ⚡ = injection in a live tool output (does not cap).

Partially evaluated not comparable

These servers were not measured the same way as the ranked ones — the agent either never scored them, or stopped partway when a tool hung. Either way their score rests on a different basis: the overall is a weighted mean over the dimensions present, over however many runs completed, so a missing Agent Task Success (the heaviest dimension) or a short sample inflates the number relative to the ranked table. Each row says which applied.

#ServerGradeScoreTask successSecurityTools
fetchA100.01
no read-only tools to test (all excluded as possibly-mutating)
sequential-thinkingA100.01
no read-only tools to test (all excluded as possibly-mutating)