An experiment in better code review

One pull request.
Two ways to judge it.

An agent prepares the environment and runs the tests. Then an LLM and Jev review the same evidence, side by side.

Public repositories · Up to 30 changed files · Live runs use your configured model accounts and sandbox provider.

Ready when you are.
Choose a PR to start a real review.

Evaluation timings exclude shared preparation and the final writer. Judgments are model assessments, not a correctness benchmark. Token usage is reported; dollar costs require provider pricing.

Activity & evidence

No activity yet