Eval Report: ci-post-merge

Profile: gdm-swebench-lite-v1 | Tasks: 50 | Pass rate: 100.0% | Cost: $0.0500

Task IDBandScorePassedCost
python-security-fix-easy-001easy0.740$0.0010
typescript-security-fix-easy-001easy0.740$0.0010
python-bugfix-easy-001easy0.740$0.0010
typescript-bugfix-easy-001easy0.740$0.0010
python-performance-easy-001easy0.740$0.0010
typescript-performance-easy-001easy0.740$0.0010
python-test-writing-easy-001easy0.740$0.0010
typescript-test-writing-easy-001easy0.740$0.0010
python-multi-file-easy-001easy0.740$0.0010
typescript-multi-file-easy-001easy0.740$0.0010
python-refactor-easy-001easy0.740$0.0010
typescript-refactor-easy-001easy0.740$0.0010
python-config-easy-001easy0.740$0.0010
typescript-config-easy-001easy0.740$0.0010
python-recovery-easy-001easy0.740$0.0010
typescript-recovery-easy-001easy0.740$0.0010
python-dependency-easy-001easy0.740$0.0010
typescript-dependency-easy-001easy0.740$0.0010
python-explain-easy-001easy0.740$0.0010
typescript-explain-easy-001easy0.740$0.0010
python-security-fix-medium-001medium0.740$0.0010
shell-security-fix-medium-001medium0.740$0.0010
python-bugfix-medium-001medium0.740$0.0010
shell-bugfix-medium-001medium0.740$0.0010
python-performance-medium-001medium0.740$0.0010
shell-performance-medium-001medium0.740$0.0010
python-test-writing-medium-001medium0.740$0.0010
shell-test-writing-medium-001medium0.740$0.0010
python-multi-file-medium-001medium0.740$0.0010
shell-multi-file-medium-001medium0.740$0.0010
python-refactor-medium-001medium0.740$0.0010
shell-refactor-medium-001medium0.740$0.0010
python-config-medium-001medium0.740$0.0010
shell-config-medium-001medium0.740$0.0010
python-recovery-medium-001medium0.740$0.0010
shell-recovery-medium-001medium0.740$0.0010
python-dependency-medium-001medium0.740$0.0010
shell-dependency-medium-001medium0.740$0.0010
python-explain-medium-001medium0.740$0.0010
shell-explain-medium-001medium0.740$0.0010
python-security-fix-easy-001easy0.740$0.0010
typescript-security-fix-easy-001easy0.740$0.0010
python-bugfix-easy-001easy0.740$0.0010
typescript-bugfix-easy-001easy0.740$0.0010
python-performance-easy-001easy0.740$0.0010
typescript-performance-easy-001easy0.740$0.0010
python-test-writing-easy-001easy0.740$0.0010
typescript-test-writing-easy-001easy0.740$0.0010
python-multi-file-easy-001easy0.740$0.0010
typescript-multi-file-easy-001easy0.740$0.0010

Leaderboard Snapshot

Latest run: 7409445c-1bda-4983-be63-03565757cb2c | Latest model: coder | Latest score: 0.740 | Recorded at: 2026-04-27T16:30:25.719665+00:00

Recent Trend

Run IDModelGit SHAScoreCreated
7409445c-1bda-4983-be63-03565757cb2ccoder4669773b4fbe9d507f1396f38777a1b36998faf30.7402026-04-27T16:30:25.719665+00:00
0a5d44de-0876-4372-a678-a8c585d25089coder4669773b4fbe9d507f1396f38777a1b36998faf30.7402026-04-27T16:30:25.651372+00:00
fde6188e-b7a8-452d-8085-19a463b27b51coder4669773b4fbe9d507f1396f38777a1b36998faf30.7402026-04-27T16:30:25.586912+00:00
3df08868-d322-46a5-9004-6270efc62af1coder4669773b4fbe9d507f1396f38777a1b36998faf30.7402026-04-27T16:30:25.517481+00:00
053ee29d-4d91-4048-840a-19545bc1b366coder4669773b4fbe9d507f1396f38777a1b36998faf30.7402026-04-27T16:30:25.460509+00:00

Failure Breakdown

TaxonomyFailuresBar
wrong-logic750
##############################################################################################################################################################################################################################################################################################################################################################################################################################################################################################################################################################################################################################################################################################################################################################################

Recent Failures

Run IDTask IDTaxonomyScoreCostCreated
7409445c-1bda-4983-be63-03565757cb2ctypescript-multi-file-easy-001wrong-logic0.740$0.00102026-04-27T16:30:25.719665+00:00
0a5d44de-0876-4372-a678-a8c585d25089python-multi-file-easy-001wrong-logic0.740$0.00102026-04-27T16:30:25.651372+00:00
fde6188e-b7a8-452d-8085-19a463b27b51typescript-test-writing-easy-001wrong-logic0.740$0.00102026-04-27T16:30:25.586912+00:00
3df08868-d322-46a5-9004-6270efc62af1python-test-writing-easy-001wrong-logic0.740$0.00102026-04-27T16:30:25.517481+00:00
053ee29d-4d91-4048-840a-19545bc1b366typescript-performance-easy-001wrong-logic0.740$0.00102026-04-27T16:30:25.460509+00:00
526fc3dc-54b3-4fd5-ab44-99cad0eefd90python-performance-easy-001wrong-logic0.740$0.00102026-04-27T16:30:25.378674+00:00
5d243663-91c0-454c-88ef-a1dd67f5e2c4typescript-bugfix-easy-001wrong-logic0.740$0.00102026-04-27T16:30:25.304097+00:00
ce2f3aa6-e4e5-4c0e-82a0-3b8703c7ae0bpython-bugfix-easy-001wrong-logic0.740$0.00102026-04-27T16:30:25.222842+00:00
cdda712f-149e-4cf3-801a-a78e85a67dfctypescript-security-fix-easy-001wrong-logic0.740$0.00102026-04-27T16:30:25.141231+00:00
97a67790-0ec1-4e06-8728-62f25143453fpython-security-fix-easy-001wrong-logic0.740$0.00102026-04-27T16:30:25.080419+00:00

Cost Frontier

pass_rate vs cost_usd (Pareto frontier marked with *)
* [####################] 100.0% @ $0.0054  (coder)