mcp-gauntlet leaderboard

Each MCP server is run through the gauntlet: a live LLM agent attempts generated tasks using only the server's tools, alongside schema, description, security, reliability, and robustness checks. Grade is the weighted overall; a critical security finding caps it.

#ServerGradeScoreTask successSecurityTools
1sequential-thinkingA100.01
2good-fixtureA99.41003
3everythingB82.85013
4memoryB81.9429
5filesystemC79.23514
6bad-fixtureC75.0504

Agent model: gemini:gemini-flash-latest · generated 2026-07-23T23:27+00:00 · scores from a live agent are stochastic (repeated and averaged); the ⚠ flag marks tool-poisoning / injection findings.