← All articles

ReviewBench: an open benchmark for automated code review

Imani Castellanos · 1 min read

A benchmark for review agents built on real pull requests, multi-source ground truth and production-aligned metrics.

Measuring review quality is harder than measuring code generation. A good review is often the comment that was not made.

What is in the set

  • 4,200 pull requests from public repositories, with consent.

  • Ground truth from maintainers, CI outcomes and later fixes.

  • Metrics that reward precision over volume.

from reviewbench import load, score
suite = load("v1")
print(score(my_agent, suite, metric="precision@3"))

A reviewer that leaves forty comments has not done forty times the work.

0 comments

  • Be the first to comment.