Coder Eval

An open-source framework for evaluating and benchmarking AI coding agents and their Claude Code skills: it runs a real agent — Claude Code, Codex, or Gemini — in a sandbox against declarative YAML tasks, then scores the files and commands the agent actually produced.

The documentation has moved to coder-eval.com/docs.

UiPath/coder_eval stars forks