Cerebe / Benchmark

How we measured 79%.

On Martian's public Code Review Bench, Cerebe caught 79% of known issues; the best of five other AI code reviewers caught 63%. This page says exactly what that sentence means, how the run was done, and what we do not claim.

The bench

Martian's Code Review Bench is a public set of real pull requests, each with a list of known issues a reviewer should find. The public offline set has 50 pull requests. A review tool is scored on how many of the known issues its posted comments match.

Our run

We ran Cerebe on the public 50-PR set in September 2026 and scored it beside Martian's published reviews from five other AI code reviewers. Cerebe's panel was three critics from competing model vendors, reviewing in isolation as they do in production. Every tool's comments, ours included, were scored by the same judge, so the numbers are comparable.

What we report

Recall: the share of known issues a tool caught. Cerebe caught 79%; the best of the five other tools caught 63%. On the subset of serious known issues, Cerebe caught 82% against 68%.

The caveat

Cerebe also comments more than those tools. The bench scores precision as well as recall, and more comments means more unmatched ones; cutting that volume is current work.

What we do not claim

Not a ranking, and not a headline score. The bench's headline score combines precision and recall; we report recall, with the caveat above beside it. Not "the best AI code review". Not a measurement of your repository: run the App on it and read the verdicts.

Reproduce the public part

Martian publishes the dataset and the scoring code on the bench site. Our own harness and per-finding data are not published.

The five other tools

Named as fact, not ranked.

  • GitHub Copilot
  • Qodo
  • CodeRabbit
  • Greptile
  • Cursor Bugbot

The five tools whose published reviews we scored. Each was scored by the same judge on the same 50 pull requests.

Measure it on your own pull requests.

Install the GitHub App on one repository. Every pull request gets a verdict, with the evidence behind every finding.