ReviewBench: an open benchmark for automated code review

A benchmark for review agents built on real pull requests, multi-source ground truth and production-aligned metrics.
Measuring review quality is harder than measuring code generation. A good review is often the comment that was not made.
What is in the set
4,200 pull requests from public repositories, with consent.
Ground truth from maintainers, CI outcomes and later fixes.
Metrics that reward precision over volume.
from reviewbench import load, score
suite = load("v1")
print(score(my_agent, suite, metric="precision@3"))A reviewer that leaves forty comments has not done forty times the work.