Reader-supported. We may earn a commission from links, at no cost to you. How scoring works →
AI Code Review, Read Off a Public Benchmark Whose Ranking Moves With a Slider
3 tools measured · Vouch Score data collected 15 August 2026 · highest composite first
AI code review means a bot that reads the diff on a pull request and writes comments on it, the way a colleague would, before a human approves the merge. What separates one from another is not whether it finds problems (they all find some) but the ratio between what it surfaces and what a reviewer then waves through. That ratio has a name on the one public board that measures these tools by a single method: precision against recall, combined into an F1 score. The board carries a slider that trades the two, and moving it re-orders the table, which is why every figure on this page is printed with the setting and the date it was read at. Vendors announce in their own blogs that they lead that board; on the reading saved here, at the default setting, one claimant held the top row while others sat further down it. Beside the benchmark sit three things a benchmark cannot tell you: what the plan costs and whether it bills per developer or per team, what the contract says when a suggestion the bot wrote turns into a claim, and what customers report once the trial is over. Each card below converts those separately and prints the source and the day it was read.
Top score: CodeRabbit. It holds the highest Vouch Score composite on this page, 73.3/100, measured 15 August 2026. The label is computed from the cards, not chosen by us, and it moves the day another tool measures higher. Read what each card measured before you take it as advice for your own work.
1CCodeRabbittop score · highest composite73.3/100 · average
Pull-request review on GitHub, priced per seat
AI code-review bot that comments on GitHub pull requests, billed per developer seat.
CodeRabbit is a GitHub-native pull-request review bot; its card reads 73.3/100 on
evidence dated 15 August 2026. Commercial terms, 92.5, is the strongest cell:
indemnity runs both ways and output is assigned to the customer. Capability, 74.5,
is where its row sits in a public board's own field: highest recall at the default
setting, eighth on precision. No free tier reaches private repositories.
Visit CodeRabbit ↗Read the full review →Pro $24/seat/mo billed annually ($30 month-to-month) · Pro Plus $48/$60 · Security add-on $40 · no free tier for private repositories (coderabbit.ai/pricing, read 15 August 2026)
AI code review priced by a team-wide credit pack rather than per seat
AI code-review platform, formerly branded CodiumAI, priced through a pooled team-wide credit pack rather than a per-seat rate.
Qodo is an AI code-review platform, formerly branded CodiumAI; its card reads 68.8/100 on evidence dated 16 August 2026. Usability, 85.5, is the strongest cell, built from a 4.5-of-5 G2 rating on 102 reviews. Its entry price is a $30 credit pack pooled across the whole team, not a per-seat rate. The indemnity runs one way, and capability, 64.3, rests on a single benchmark board.
Visit Qodo ↗Read the full review →Pro Team $30 / 2,500 credits (~18 reviews a month) · $60 / 5,000 (~36) · $240 / 20,000 (~144), at a published $.012 per credit pooled across the team · monthly billing only, no annual self-serve option · Enterprise quoted, above 30 users · no permanent free tier; a 14-day unlimited trial and an apply-to-qualify open-source programme (qodo.ai/pricing, read in a rendered browser 16 August 2026)
Pull-request review where the bill follows merge rate, not headcount
AI code-review bot for pull requests, sold per seat with a credit meter that counts reviews rather than people.
Greptile is an AI code-review bot for pull requests. It scores 66.2/100 on our card, an
average band with four of four dimensions measured. Precision is its strength: 75.3% on a
public board read 15 August 2026, against 48.1% recall. Against it, a $30 seat that meters
reviews at $1 each past fifty, a fourteen-review crowd sample, and an agreement that
permits AI training on de-identified data with an opt-out.
Visit Greptile ↗Read the full review →Pro $30/seat/month, monthly billing only (the page carries no annual view) · 50 credits included per seat, 1 credit per standard review and 3 per TREX review, then $1 per extra credit · Starter free for 1 active developer with unlimited repositories and 50 credits/month · 14-day trial · Enterprise quoted · free for qualified MIT/Apache non-commercial projects and 50% off for pre-Series A companies under $2M revenue (greptile.com/pricing, read in a rendered browser 18 August 2026)
Outbound links may be affiliate links and can earn us a commission: they never touch a score, and the order on this page is the composite order, computed at build time from the cards themselves.
The score cards in full
What each card measured and what it did not: every number dated, sourced and reproducible.
Pro $24/seat/mo billed annually ($30 month-to-month) · Pro Plus $48/$60 · Security add-on $40 · no free tier for private repositories (coderabbit.ai/pricing, read 15 August 2026)
Pro Team $30 / 2,500 credits (~18 reviews a month) · $60 / 5,000 (~36) · $240 / 20,000 (~144), at a published $.012 per credit pooled across the team · monthly billing only, no annual self-serve option · Enterprise quoted, above 30 users · no permanent free tier; a 14-day unlimited trial and an apply-to-qualify open-source programme (qodo.ai/pricing, read in a rendered browser 16 August 2026)
Pro $30/seat/month, monthly billing only (the page carries no annual view) · 50 credits included per seat, 1 credit per standard review and 3 per TREX review, then $1 per extra credit · Starter free for 1 active developer with unlimited repositories and 50 credits/month · 14-day trial · Enterprise quoted · free for qualified MIT/Apache non-commercial projects and 50% off for pre-Series A companies under $2M revenue (greptile.com/pricing, read in a rendered browser 18 August 2026)
Read the dimensions, not the composite
Two decisions sit ahead of the price, and the second one is the expensive one to get wrong. First, where you want to land between precision and recall: a bot tuned to catch more also raises more that a reviewer has to dismiss, and that trade-off decides whether the tool blocks a merge or only annotates it. A team that already ignores its linter will ignore a noisier reviewer too. Second, the billing unit, because it changes the answer as a team grows rather than at signup: a per-seat price scales with headcount, while a pooled credit pack costs the same for five people as for twenty-five and runs out on volume instead. Work out which of those two your growth looks like before comparing the headline numbers, because they are not the same kind of number. Then read the indemnity, and read it for scope rather than for direction. A clause can run toward the customer and still carve out the thing you needed it for; one contract in this category names the tool's OWN generated output among the things the customer indemnifies the vendor against, which on a product whose entire job is to generate suggestions is worth knowing before it matters. What a card here does not answer is whether the bot is right about YOUR code. Every figure on this page converts somebody else's repositories, somebody else's opinions, or a document. No run against your codebase sits behind any number, and none is claimed.
Where these numbers come from
Capability is sourced from public output-quality arenas (blind pairwise votes) or from a published accuracy study where no arena covers the category, usability from review-crowd aggregates weighted by sample size, value from verified pricing, and commercial terms from clause positions read off the vendor’s own legal documents on a stated date. Where none of those exists for a tool, the row carries a named criterion instead: a different measurement, taken from saved sources under its own rubric, and the card prints that criterion’s name and the question it answers in place of the axis heading, so the row is never read as the axis it could not fill. Full detail: methodology.
The order on this page is computed from the composites at build time and nothing on it was typed by hand. What is worth saying in words is what the numbers do NOT mean.
Coverage is not depth. A card can carry a value in every dimension and still rest on thinner evidence than one beside it. CodeRabbit's capability cell converts two independent benchmarks and Qodo's converts one, and Qodo's own card says so. Greptile's crowd cell rests on fourteen reviews that the platform itself labels incentivized and invited, which is the thinnest consumer evidence on this page and is marked as such where the figure appears. A cell filled twice and a cell filled once print the same way.
Matching value scores are arithmetic, not agreement. The billing structures behind them do not map onto each other: a per-seat rate, a pooled credit pack, and a seat rate with a review meter behind it are three different questions wearing one number, and this category's value basis had to be rebased once already because of it.
And a rank is a reading of one view. The public board these capability cells convert reports precision and recall separately as well as the F1 that combines them, and the ordering changes with the weighting: the card seventh on F1 holds the second-highest precision on that board. Read the row, not the position.
Questions buyers actually ask
How accurate is AI code review?+
There is one public board that measures these tools by a single method against real pull requests, and it reports precision, recall and an F1 score that combines them. Read on 15 August 2026 at its default balanced setting, the eleven tools on it ran from 42.9% to 62.4% F1, with the middle of the field at 59.2%. That range is the useful part: no tool on the board is close to catching everything, and the spread between the best and the median is about three points. The board also moves (a slider re-weights precision against recall and re-orders the table, and the window resets monthly) so an accuracy figure quoted without the setting and the date is not a measurement.
Is there a free AI code review tool for private repositories?+
Free tiers here differ in kind rather than in size, so read the free-tier row on each card instead of the headline price. Greptile publishes a standing free plan (one active developer, unlimited repositories, fifty credits a month) and a credit is one standard review. Elsewhere on this page a permanent free plan is absent and a time-limited trial stands in for it. Open-source maintainers are a separate case: an apply-to-qualify programme exists in this category and is not the same thing as a free plan.
Does AI code review bill per developer or per team?+
Both, and increasingly at the same time. One structure prices per seat, so the bill follows headcount. Another sells a credit pack pooled across a whole team at one price, so a small team and a large one pay the same and the limit arrives as volume rather than as an extra invoice line. And a seat rate can carry a meter behind it: a fixed allowance of reviews included, each further review billed on its own. That last shape is the one a headline price hides, because the bill follows pull requests rather than people. The value cell on each card converts the entry price on the vendor's own page, in the billing view stated beside the figure.
What does the contract say if the bot's suggestion causes a problem?+
It varies more than the price does, and it is read off each vendor's own agreement rather than summarised. The commercial-terms cell on each card converts clause positions: which way the indemnity runs, who owns code the bot suggested and a developer accepted, and what a refund actually requires. Read the scope as well as the direction: a clause pointing the customer's way can still exclude the case you care about.
Which AI code review tool should I pick, and why does this page not crown one?+
Read the order, which is computed from the composites when the site is built rather than typed by anyone, and then read what each card actually measured before treating it as advice for your own repository. What this page will not do is stage a contest between two products or declare a winner in prose, because a sentence about who leads stops moving the moment either card is rescored, and both were rescored this month. The pick that survives that problem is the one you make from the rows: the benchmark row if noise is your failure mode, the value row if headcount is, the contract row if a client contract is.