Reader-supported. We may earn a commission from links, at no cost to you. How scoring works →

CodeRabbit

code review tools

AI code-review bot that comments on GitHub pull requests, billed per developer seat.

Visit CodeRabbit

Pro $24/seat/mo billed annually ($30 month-to-month) · Pro Plus $48/$60 · Security add-on $40 · no free tier for private repositories (coderabbit.ai/pricing, read 15 August 2026) · affiliate links never affect the score

Vouch Score v11 · data collected 15 August 2026
73.3/100
average
  • Capability74.5
  • Usability & control83.7
  • Value50
  • Commercial terms92.5
Sources: . Recomputable from the committed inputs; see methodology.

By Minel Gunesoglu, founder of Vouch. The work behind this page was documentary: I read the Martian leaderboard in a rendered browser on 15 August 2026 and kept its bytes, opened CodeRabbit's pricing page in both billing views, read the Terms of Service and the refund article in the vendor's knowledge base, pulled the two changelog entries where the vendor names the noise problem itself, and corrected a G2 figure our own records had wrong in both directions. No repository of mine has been reviewed by CodeRabbit, and no benchmark run of mine sits under any number below. Published 15 August 2026, re-checked monthly.

Three companies selling code-review bots have each published a post announcing that they lead the same public leaderboard, and CodeRabbit is one of the three. That board was opened in a browser on 15 August 2026 and its bytes were saved. At the default balanced setting the table put cubic.dev's row first, CodeRabbit's fifth and Greptile's seventh, out of eleven rows. The board is not a fixed ranking: a control on the left rail, labelled "How important is it to avoid noise vs being thorough?", trades precision against recall and re-orders the table as it moves, and the window resets every month. Each of those three announcements can be true of a particular view on a particular date. None of them says which view, and none of them carries a date you can hold the table to.

What that table shows about CodeRabbit specifically is one shape rather than two findings. At beta 1, the balanced default, on 15 August 2026, CodeRabbit's row reads 59.3% F1, 55.0% recall and 64.3% precision. The recall figure is the highest in that field and the precision figure sits eighth once you sort the column. Those are not a strength and a separate weakness. They are the same behaviour seen from two sides: a reviewer that surfaces more of what is there also surfaces more of what a developer then waves through. The most repeated complaint in this product's public record says exactly that in plainer words, and so, in its own changelog, does the vendor.

CodeRabbit is a pull-request review bot from CodeRabbit, Inc. It installs against a repository, reads the diff on each pull request, and posts line-level comments and suggested changes in the review thread where the team already argues. Everything below describes it on GitHub, because GitHub is the platform every document read for this review names; third-party listings mention other hosts and none of those was confirmed here. Two identity notes before any number is worth reading. Search results on both G2 and Capterra surface SalesRabbit, an unrelated field-sales product, beside this one, so every figure on this page comes from a URL confirmed to carry CodeRabbit's own product path. And the vendor's GitHub Marketplace listing counted over three hundred thousand installs when it was read on 15 August 2026 (the counter is live and moved by roughly two hundred inside that working day, so it is stated as a scale rather than a figure), which is a fact about distribution and about nothing else. It says how many places the bot is installed. It says nothing about how good the reviews are, and one section below turns on how much of that install base the vendor's own refund policy does not reach.

How Accurate Is CodeRabbit? A Public Benchmark Whose Ranking Moves With a Slider

Every cell on CodeRabbit's card carries a value, and every source behind them was read on 15 August 2026. That is unusually clean for a card here, and it makes the interesting question a different one: not which row is missing, but what each row is actually a reading of. One converts a leaderboard position. One converts a review-platform average. One converts a price list. One converts a contract. Only the first two sample anyone's experience, and neither of them samples yours.

DimensionReadingWhat that reading converts
Capability74.5Martian Code Review Bench, online tracker, F1 at the board's default beta of 1: CodeRabbit's 59.3% across 879 scored pull requests inside a 4,299-PR "Last month" window, placed against the eleven rows the board published beside it
Usability83.7G2's profile for CodeRabbit at 4.4 out of 5 across 97 reviews, read in a rendered browser after three other methods were refused
Value50.0coderabbit.ai/pricing, Pro at $30 per seat per month month-to-month — the basis this card scores, because the sibling tool in this category publishes no annual option to compare against; $24 on the annual view the page loads by default
Commercial terms92.5Three clause positions in CodeRabbit's own Terms of Service and knowledge-base refund article; a fourth the documents never address

The first row needs one sentence of explanation, because 74.5 is not a number the board prints anywhere. F1 on real pull requests has no working ceiling at 100: across all eleven rows read that day the field ran from 42.9% to 62.4%, with its middle at 59.2%, so treating a raw percentage as a score out of a hundred would hand the top third of the scale to ground no reviewer on that board occupies. What the cell converts is therefore the position of the row inside the field the board itself published, which is how the arena-scored cards elsewhere on this site have always worked. CodeRabbit's 59.3% sits a tenth of a point above that middle, and 74.5 is what that position comes to.

Composite: 73.3 out of 100, or 3.7 on the five-point scale. The word printed next to it is a label for that total and says nothing about how much evidence went into producing it, on this card the answer happens to be all four cells, which is not true of every card on this site. The four are combined by multiplication rather than addition, so the lone reading at 50.0 pulls down harder than the 92.5 pushes up: add the same four numbers and average them and you get 75.2, while the published figure is 73.3, nearly two points lower. That gap is the whole argument for reading the rows instead of the total. A buyer who cares about what a contract commits the vendor to and a buyer who cares about how much noise lands in a review thread are looking at opposite ends of this card, and the middle number serves neither of them well.

One consumer figure deliberately scores nothing here. CodeRabbit's Trustpilot profile showed 2.4 out of 5 across seven reviews on 15 August 2026. Seven is below the sample floor this desk requires before a rating is converted into a cell at all, so that figure contributes nothing to the 73.3 above and is carried further down this page as dated prose instead. It is stated plainly rather than quietly dropped, because a reader who meets a low rating in the same article as a low-70s composite is entitled to know which one the arithmetic used. What goes into a card here, and what is deliberately kept out of one, applies to every tool measured on this site.

What no row above measures is whether CodeRabbit is right about your code. A leaderboard position converts other people's pull requests; a star average converts other people's opinions of a workflow; a price and a contract convert documents. A controlled run against a repository under fixed conditions would answer the question all three of those approximate, and no such run stands behind any figure on this page.

CodeRabbit's Martian Benchmark Score: Highest Recall of the Field, Eighth of Eleven on Precision

Martian's Code Review Bench is the one public, cross-vendor measurement this desk could open that scores CodeRabbit and its named rivals by a single method, which is why the capability cell converts it rather than any vendor's own study. The board runs two things. An offline benchmark works against a fixed set of hard-to-find bugs. The online tracker, which this card reads, works against real pull requests where these bots participated, and in its own words it "analyzed thousands of real pull requests across GitHub where AI code review bots actively participated", collecting each review timeline and then asking whether "the bot's suggestions lead to real code changes?" Each pull request is scored on precision, recall and F1.

Three things about that method decide how much a single number off it is worth. It measures suggestions that turned into code changes, which is a proxy for usefulness rather than a test of correctness, since a developer can accept a bad suggestion and refuse a good one. It publishes a per-tool sample size, which is what makes it usable as a scored dimension at all: CodeRabbit's row rests on 879 scored pull requests, more than any other row on that board carried when it was read on 15 August 2026, inside a header the board itself labels "Across 4,299 scored PRs" for the period selector's "Last month" setting. And it moves. The board's own evolution chart shows rows changing places week to week, the window rolls monthly, and the capture behind the figures here carries an expiry of 15 September 2026 for exactly that reason.

The parameter matters as much as the date. The left rail carries a slider with a manual F-score input, and it was sitting at 1, the balanced default weighting precision and recall equally, when the capture was taken. Change that setting and the order changes, because the two quantities trade against each other. So the honest form of every figure in this section is the figure plus beta plus date: at beta 1, on 15 August 2026, CodeRabbit's row reads F1 59.3%, recall 55.0% and precision 64.3%, in fifth position on a field of eleven. Sort the recall column on that same reading and CodeRabbit's row is at the top of it. Sort the precision column and the same row sits eighth of the eleven.

That is where the three competing announcements come from. The board's header calls the table an "Overall online leaderboard balancing noise and thoroughness", and balancing is doing real work in that sentence: a table whose order depends on where a reader sets a slider does not have a permanent winner to name, and nothing on the board read that day declared one. It publishes an order under a setting. A vendor post quoting a position from it without naming the setting, the date and the window has not quoted a measurement, it has quoted a moment, and that applies to CodeRabbit's own post as much as to anyone else's.

A second public benchmark exists, and reading it is the most useful thing anyone can do with the Martian row, because it disagrees violently about the level while agreeing exactly about the shape. Entelligence AI (a company selling a competing review bot, and one that ranks itself first in its own table), published a comparison of eight reviewers against 67 real production bugs drawn from Cal.com, Sentry, Discourse, Keycloak and Grafana, judged by an LLM against a golden set of known fixes. Read on 15 August 2026, its aggregate table gives CodeRabbit 33.0% F1, 49.2% recall and 24.8% precision, on 65 of the 67 pull requests by its own footnote. Sort that recall column and CodeRabbit is first of eight. Sort the precision column and it is seventh of eight.

Hold that beside the Martian reading: highest recall of eleven, eighth on precision. Two boards that share no data, no judge and no tool field agree on which way this product leans, and land 40 points apart on how precise it is: 64.3% against 24.8%. Neither number is the answer to "how accurate is CodeRabbit", and the gap between them is not a scandal, it is what happens when two operators pick different bugs and different judges. What survives both is the trade itself. A buyer can act on a reviewer that reliably surfaces more and is reliably noisier; nobody can act on a single precision percentage lifted from whichever page loaded first.

What CodeRabbit's Comments Actually Are, by an Open-Source Maintainer Who Counted Them

The most rigorous account of what lands in a review thread was not produced by a vendor. A maintainer on the Lychee project ran CodeRabbit against his own repository, tagged every comment it left over thirty days, and published the tally: "15% were useless 13% were wrong assumptions 21% were nitpicking, 13% were thoughtful, 35% were quality improvements and 3% of those were security/critical findings" (lycheeorg.dev, 13 September 2025). The punctuation is his, the first two figures run on without commas in the post itself, and adding them would be tidying a quotation.

That sentence is quoted rather than tabulated on purpose, and its last clause looks at first like an unresolved ambiguity: "3% of those" reads as though the security findings might sit inside the 35%, in which case the five figures before it come to 97 and three per cent of the whole is unaccounted for. Reading further in the same post settles it, and settles it twice. His taxonomy section defines the category outright ("Security/Critical findings: This is a special sub-category of Quality Improvements. They are not counted towards the previous") and his own arithmetic in the same paragraph as the summary only works on that reading: he writes that "28% of the findings were not much interesting, but also that 72% of the findings were relevant, of which a bit less than 3 over 4 (51/72 =~ 71%) brought actual value". Fifteen and thirteen make his 28. Twenty-one, thirteen, thirty-five and three make his 72. Thirteen, thirty-five and three make his 51. Six buckets, outside rather than inside, summing to exactly 100. What the count establishes, on its own terms and from a maintainer with no stake in the answer, is that a little over a quarter of the comments landed in his two lowest-value categories while just over a third were quality improvements.

The same maintainer, in the same write-up, tells other maintainers to adopt it: "If you are an opensource maintainer, I highly recommend you give it a try." Both halves belong in one passage, because a page that prints only the percentages has selected its evidence as surely as one that prints only the recommendation. A tool whose output is a third useful and a quarter noise is a good trade for a maintainer reviewing volunteer contributions alone, and a bad trade for a team of six who all have to read the thread.

Two dated records from other people describe the same shape independently. WooCommerce's maintainers opened a public GitHub issue on 16 June 2025 asking for CodeRabbit's reviews to be configured more defensively, on the grounds that "a lot of false positives" meant contributors were ignoring them (issue #58887). A G2 reviewer, writing on 28 July 2026 under the title "Powerful AI PR reviews, but expect tuning, noise, and CI/CD slowdowns", describes alert fatigue and minor stylistic nitpicks escalated as major issues, and adds two operational specifics worth more than the adjective: "on larger or more complex pull requests, CodeRabbit can sometimes take upwards of 20 minutes to complete its analysis and post its comments", and the review is not deterministic, since "you can resolve all six of CodeRabbit's initial PR comments, but when the webhook triggers the second review, the AI suddenly finds five new issues it completely ignored the first time".

Then there is the vendor's own record, which is stronger evidence than any number of reviewers agreeing. CodeRabbit's changelog carries an entry dated 12 February 2026 introducing an auto-pause: "To avoid noisy feedback, CodeRabbit now automatically pauses incremental reviews after 5 reviewed commits on a pull request (default)." A second entry, dated 2 July 2026, records that "CodeRabbit now offers a Quiet review profile alongside the existing Chill and Assertive profiles" (docs.coderabbit.ai/changelog, both read 15 August 2026). A company naming the problem in its own product log, in its own words, and shipping a default against it twice in five months is the clearest confirmation available that the complaint describes something real. What those two entries do not establish is that the noise is now solved, or that it was worse before. They are dated shipments, not measurements, and this page reports them as such rather than turning them into a tuning walkthrough.

One caution about where that G2 review turns up again, because a second sighting of a rating is not a second rating. CodeRabbit's AWS Marketplace listing hosts no reviews of its own: every entry on it is labelled "Review provided by G2", and its headline figure, 4.4 across 86 ratings when read on 15 August 2026, is the same G2 aggregate seen through a smaller window. A number that reappears on a second site under syndication is one opinion displayed twice, not two sources agreeing, and nothing on that listing enters this card as an independent input.

CodeRabbit Pricing: $24-$30 a Seat, and No Free Tier Once a Repository Turns Private

CodeRabbit's pricing page was opened on 15 August 2026 and both billing states were toggled by hand, because the annual view is what the page loads by default and the monthly figures render only behind the switch. A capture of one state would have recorded half the answer.

PlanMonthlyAnnual, per monthWhat the tier is
Pro$30 per seat$24 per seat, billed annuallyThe entry paid tier
Pro Plus$60 per seat$48 per seat, billed annuallyMarked "Recommended" on the page
CodeRabbit Security$40 per seat$40 per seatAdd-on; the toggle does not change it
EnterpriseQuotedQuoted"Talk to us" in both views

Every figure there was read on 15 August 2026, and that date carries more weight on this row than on any other, because a pricing page is edited without notice and nothing on this page would know. The "Save 20%" badge is exact rather than approximate, which is worth stating because it is not always: 24 divided by 30 is 0.800 and 48 divided by 60 is 0.800, so the advertised discount and the arithmetic agree. A separate Slack agent is billed by usage at fifty cents per agent minute, with no billing-cycle variant. The value cell of 50.0 converts the $24 annual-equivalent seat price, which is the basis this category is declared on; it is a reading of what the vendor charges, not a verdict on what the tool is worth to you.

The header on both views reads "All plans include a 14-day free trial", and that trial is the whole of the free path for most buyers. The pricing page shows no free tier at all. A free offer does exist, and it appears only on the GitHub Marketplace listing rather than on the vendor's own pricing page: "Open Source — Pro Plus plan is free for Open Source projects. $0". So the top paid tier is given away, on the condition that the repository is public. A team whose repository is private has fourteen days and then a card. That condition is why this card records no free tier, while cards elsewhere on this site use that phrase to mean anyone may use the tool for nothing.

The same page publishes a rate limit in its own comparison table, and it is a number worth reading before a rollout rather than after: "Rate limits: PR reviews per developer per hour" runs 5 on Pro, 10 on Pro Plus and 12 on Enterprise, each carrying the footnote "* Subject to Fair Usage Policy". Other ceilings in the same table run 5, 15 and 20 MCP connections, and 1, 10 and 20 repositories of multi-repo analysis. Five reviews per developer per hour is generous for a team merging a few pull requests a day and tight for anyone pushing rapid iterations, which is the same population the auto-pause default was aimed at. A 200-file-an-hour cap also circulates for this product and appears nowhere on the pricing page, which is not a contradiction once you ask which plan it describes. It belongs to the open-source plan, and the Lychee maintainer documents it directly: "Number of files reviewed per hour: 200 Number of reviews: 3 back to back reviews followed by 2 reviews per hour." The vendor's table covers the paid tiers, at 5, 10 and 12. Different plans, not competing numbers, and worth knowing before a team on the free public-repository path benchmarks its throughput against a paid tier's ceiling. A separate "four reviews an hour" in third-party trackers matches neither and is not published here.

The $150 Per-Seat Surprise: What Two CodeRabbit Trustpilot Reviews Reported in 2025

Two separate accounts on CodeRabbit's Trustpilot profile (Louis Huort, posting a first-ever review, and Joao Henriques, posting from an account with eleven), describe the identical mechanism on the same date, 14 April 2025, and land on the identical figure. Both report that a plan they understood as a flat monthly subscription began charging per additional collaborator the moment more than one contributor was connected to the repository, "$30 per person" in the first account and "$30 per head" in the second, and both say the total they were charged came to $150. The first adds that they found out "the hard way" and that the charge was not refunded "even though I caught it within the hour"; the second writes that "even if you notice this the next hour, they will NOT refund you whatsoever".

Two accounts converging on the same trigger, the same rate and the same total is a stronger signal than either would be alone. What this page does not claim is independence: both are listed in France and both posted on 14 April 2025, which is a coincidence the profile shows and does not explain. So it is reported as what it is, two dated customer accounts on a public profile, not a finding of ours about how anyone's billing works. What can be checked today is the disclosure. The GitHub Marketplace listing prices the plan explicitly as $30 per seat per month, and the vendor's own pricing page prices every tier per seat rather than per repository or per organisation, so the per-seat basis is stated where a buyer meets it now. Those two reviews are dated more than a year before the pages read for this review, and what they describe is how clearly that basis surfaced at the moment a collaborator was added, not a structure the current listing leaves out.

A buyer's practical takeaway is short and does not require anyone's characterisation to be adopted. Per-seat billing plus automatic seat assignment on collaborator connection means the bill is a function of who touches the repository, not of what was agreed at signup. Before a rollout, find out which control adds a seat, and who on the team can trigger it.

CodeRabbit's Contract: Two-Way Indemnification, and Who Owns Code It Suggested

The commercial-terms cell is the highest on this card at 92.5, and unlike the other three it does not depend on any platform serving a page that day. It reads clause positions off documents anyone can open. CodeRabbit's Terms of Service, carrying its own stamp of 20 April 2026 and read 15 August 2026, answers three of these four positions and is silent on the fourth.

The indemnity runs in both directions, which is uncommon enough among the contracts read for this site to be the single most valuable thing in the document. Section 10.1 has the customer indemnify the vendor in the usual way. Section 10.2 runs the other way: "CodeRabbit agrees to indemnify, defend, and hold Customer and its officers, directors, employees, agents and representatives harmless, including costs, liabilities and legal fees, from any Claim made by any third party against Customer alleging that the Services infringe or misappropriate any patent, copyright, or trade secret of such third party." The carve-out attached to it is the universal combination clause, covering use of the Services with other services, hardware, data or business processes not provided by CodeRabbit, or use contrary to the Agreement. It removes no tier and no class of the customer's own content. The remedy ladder is stated too: replace or modify the Services, obtain a licence, substitute an equivalent, or "terminate this Agreement and refund any prepaid, unused fees", and the section closes by declaring itself the customer's exclusive remedy for an infringement claim.

Here is the limit, and this page states it rather than letting a strong cell speak for itself. That clause indemnifies the customer against a claim that the Services infringe. It does not, on its face, indemnify the customer for code they ship after accepting a suggestion the bot made. Those are different exposures, and for a tool whose whole purpose is proposing changes a developer accepts with one action, the second is the one an engineering lead is actually worried about. The document does not answer that question in either direction. An unanswered question is not a denial, and it is not a promise either.

Output rights are the second position and they are unusually clean. Section 6 reads: "CodeRabbit hereby assigns to Customer all of its rights, title and interest (if any) in and to the Output." An assignment is a stronger instrument than a licence, and it is conditioned on "compliance with this Agreement, including but not limited to paying all fees when due." Two qualifications ride with it and neither converts an assignment into something weaker. The "(if any)" hedge acknowledges that whatever rights exist in machine-generated output are unsettled, which is a question no vendor's contract can settle alone. And the same section reserves the vendor's ability to use Aggregated Data and Output to "provide, maintain, protect and improve the Services or operate its business", while noting that "The Services may provide the same or similar Output to others, and CodeRabbit's assignment to Customer in the preceding sentence does not apply to any outputs resulting from other users' use of the Services."

The fourth position, whether a free tier may be used commercially, is not addressed anywhere in the document. A full read finds no free-tier, free-plan or open-source clause at all, and the pricing page shows no free tier either; the open-source offer lives on the Marketplace listing. That absence is recorded as an absence and excluded from the average rather than guessed at in either direction. One further term belongs in a buyer's reading even though it scores nothing: the Terms carry an arbitration notice stating that "you agree that disputes arising under these Terms will be resolved by binding, individual arbitration", with a jury-trial and class-action waiver printed in capitals.

CodeRabbit's Refund Policy: Discretionary, With a 24-Hour Window and a Marketplace Carve-Out

The Terms alone would have left the refund position unreadable. Section 5.2 gives one contractual refund right, on the customer's termination for cause, and section 5.4 says termination otherwise does not entitle the customer to a refund. That is neither a cooling-off window nor a flat no. The missing document is a knowledge-base article, carrying its own stamp of 18 December 2025, which is where the vendor publishes its discretion and its criteria: "Generally, CodeRabbit subscriptions are not refundable. That said, refunds may be issued at CodeRabbit's discretion, as outlined in our Terms of Service. We review every refund request and will help depending on the circumstances."

The article then publishes both lists, which is more disclosure than most vendors in this corpus offer. A refund is more likely when "You cancel the subscription within 24 hours of the initial signup", when usage was blocked by an issue determined to be caused by the product rather than customer-side setup, and when the requester is an authorised Admin or Billing Admin. It is commonly declined when "Cancellation occurred more than 24 hours after purchase", when "The subscription was actively used during the billing period", or when the request does not come from an authorised person. One line deserves to be read twice by anyone who assumes the two acts are the same: "Cancellation does not automatically trigger a refund."

Now the part neither the pricing page nor the Terms mentions, and which changes who this policy is for. Its closing paragraph reads: "If your subscription was purchased through a third-party marketplace (e.g., GitHub Marketplace), all billing and refund requests must be handled directly through that marketplace. CodeRabbit does not have the ability to issue refunds for marketplace purchases." Set that beside the three hundred thousand-plus installs on the vendor's own Marketplace listing. For the purchase route the product page pushes hardest, this refund policy is not the governing document at all, and the buyer's rights live in GitHub's terms rather than CodeRabbit's. That appears in the last paragraph of one help-centre article and nowhere else this desk read.

One tension sits between the written policy and a dated customer account, and this page states both and stops. The policy names cancellation within 24 hours of initial signup as one of the few circumstances where a refund is more likely. One of the Trustpilot reviewers above reports catching a $150 per-seat charge within the hour and being refused. The two may not be the same clock: the published window is framed around initial signup, and a later per-seat charge on an existing subscription is a different event. It is the plainest place in this record where a document and a customer's account point in different directions, and resolving it from outside would take facts neither source supplies.

What CodeRabbit's Customers Report: a Security Disclosure Answered Fast, No Replies on an Unclaimed Review Profile

The usability cell of 83.7 converts the largest review aggregate available for this product. G2's profile for CodeRabbit showed 4.4 out of 5 across 97 reviews when it was read in a rendered browser on 15 August 2026, after curl, a fetch tool and a reader proxy were each refused by the site's protection layer. Four point four stars out of five is 88 when the same rating is expressed out of a hundred, and the cell prints 83.7 instead. The difference is the sample doing its work: ninety-seven people is enough to establish a direction and not enough to fix a precise level, so the reading is held back toward the middle of the range rather than taken at face value. Read the 83.7 as clearly liked, on a sample small enough that the second digit is not worth arguing about.

Three caveats travel with that figure wherever it appears. The profile is unclaimed: the page carries a "Re-claim Profile / Unlock Access" prompt, meaning the vendor does not currently curate the listing. The top review is footed "Incentivized · Source: G2 invite", so it was solicited, which does not make it false but belongs beside any number drawn from the page. And this reading corrects our own earlier record, which had 4.8 out of 5 on 26 reviews taken from a search-engine summary that the community pass honestly flagged as unread. Both halves of that were wrong, in the rating and in the sample, and a cell built on the snippet would have been wrong in both directions at once.

Where the praise concentrates is exactly where a thorough, imprecise reviewer fits the job best, and that is a finding rather than a balancing gesture. The maintainer who counted something over a quarter of the output as noise recommends it to other maintainers anyway. A Product Hunt reviewer values a reviewer that "never tires of nitpicks" precisely because most of what he now reviews was written by coding agents rather than people; that review carries only a relative age stamp on the page as read on 15 August 2026, so no calendar date is asserted for it here. A separate G2-syndicated review on the AWS Marketplace listing, dated 12 August 2026 and written by a solo developer, describes code review as the step that "either gets skipped entirely or done by me reviewing my own PR an hour after writing it", and then names the same twenty-minute latency the critical review does, which is the clearest sign in this record that the complaint and the recommendation are about one behaviour rather than two camps. Read forward, the Product Hunt line is the most interesting argument in the whole record: the trait that reads as alert fatigue on a team of humans reads as coverage when the pull requests arrive machine-written.

That same 28 July 2026 review is also worth reading past its 3.5-out-of-5 star count, because a mid rating with specifics is more useful to a rollout decision than a five with none. Alongside the 20-minute waits and the non-deterministic second pass quoted earlier, it flags a scope limit that matters in any service-oriented codebase: "If your PR relies on internal packages, shared libraries, or microservices housed in completely different repositories, CodeRabbit often loses the plot." The same reviewer writes that there have been "widespread reports of 'dark patterns' in their billing systems", a characterisation this page quotes and does not adopt: the review names no source for those reports, nothing in the documents read here speaks to a missing cancellation control, and this desk did not verify it. What can be stated in our own voice is narrower and checkable, and it sits in the two sections either side of this one.

CodeRabbit's August 2025 Security Disclosure, and How It Compared to Vendors Who Stayed Silent

A code-review bot is granted read access to a customer's source by design, so a remote-code-execution finding in the vendor's own pipeline is a fact an engineering lead weighing a rollout is entitled to have. On 20 August 2025, updated five days later, Endor Labs published a writeup of research credited to Kudelski Security describing a flaw in CodeRabbit's review pipeline, summarised in the document's own words as having "uncovered an RCE flaw in CodeRabbit exposing 1M+ repos. Here's what happened, how it was fixed, and key lessons for secure AI apps." The chain ran through an unsandboxed linting step processing untrusted pull-request input, reaching environment variables and, from there, write access across the connected repository estate.

Two boundaries around that paragraph are load-bearing. The finding is exposure: a reachable flaw, at the scale the source itself states. Nothing in the documents read here says any customer repository was accessed by anyone, and exposure and access are different claims with different evidence behind them. And the fix is part of the finding rather than a softener bolted on to it. The source's own framing is what happened and how it was fixed, and reporting the first half without the second would misreport the document.

The counterweight belongs in the same passage, because leaving it out would be evidence selection. The Hacker News thread that ran alongside the disclosure on 19 August 2025 is not a straightforward pile-on. Highly upvoted comments treat running an unsandboxed analyzer against untrusted input as a basic engineering lapse. In the same thread, one of the researchers writes that other vendors they contacted "never responded at all", and that those products remained vulnerable. A CodeRabbit team member replies repeatedly that day under a named account, identifying himself as "Howon from CodeRabbit" and observing that most security bugs get fixed with no public notice at all. A company that answers publicly, by name, within a day of a disclosure about its own pipeline is doing something the same thread says its competitors did not do.

Zero Company Replies Across CodeRabbit's Seven Trustpilot Reviews

Set that security response beside the other public channel and the picture is less symmetrical than it first looks. CodeRabbit's Trustpilot profile carried a TrustScore of 2.4 out of 5 across seven reviews on 15 August 2026, spanning 14 April 2025 to 5 July 2026, and not one of the seven shows a visible reply from the company. The profile is also unclaimed, and Trustpilot's own banner on it reads: "No history of asking for reviews — This company hasn't invited their customers, so reviews may not be representative." So the absence of replies is an absence on a listing the vendor does not hold, which is the ordinary explanation and the one this page reports. Seven reviews is a small enough sample that it scores nothing on the card above; it is also small enough that the count of replies is complete rather than sampled.

What those complaints are about is as notable as the rating. They cluster on billing and on leaving rather than on the product: the two per-seat accounts above, an annual subscription a reviewer says was charged twice because it went through a GitHub organisation (17 June 2026), a plan sold as unlimited that a reviewer says was rate limited within days (5 July 2026), and a cancellation a reviewer describes as confirmed and then billed anyway (23 April 2026). The rate-limit complaint has a documentary counterpart worth putting beside it: the vendor's own pricing table publishes a per-developer hourly review limit on each tier, so a ceiling exists in writing, and what that reviewer describes is the gap between the table and what they understood they were buying.

Two dated observations, side by side, are the whole of what this page can show. The same company answered a security researcher publicly, by name, within a day. It has replied to none of the seven billing and cancellation complaints on an unclaimed Trustpilot profile across fifteen months. Both are checkable in one click, and the second carries the caveat above, so the conclusion is the reader's to draw rather than ours to state.

CodeRabbit Alternatives: Greptile, Qodo and cubic.dev

Greptile, Qodo and cubic.dev are the names that appear beside CodeRabbit most often, and each of them carries a row on the same Martian board, which is exactly why this section names them and stops. Qodo has since been scored here on its own card, published 16 August 2026, and no ranking between it and CodeRabbit appears anywhere on this page. Greptile and cubic.dev carry no score card here as of that date, so no composite or cell for either of them appears either, and the board's own numbers for their rows are theirs to publish rather than ours to reprint beside a card they were never measured against.

What the board does show about the shape of the field is worth one paragraph, because it changes what a shortlist means. On that 15 August reading at beta 1, nine of the eleven rows sat within nine points of each other on F1, while the precision and recall columns behind those same nine each spread wider than that. The tools separate less on the balanced total than on the two quantities the total is balancing. Choosing between them is therefore choosing where you want to sit between missing things and being interrupted, and no composite answers that question, including the one at the top of this page. Both benchmarks quoted on this page were published by companies with a bot in the race, one of which puts itself at the top of its own table, and that is a reason to read the method rather than a reason to discard the numbers, the check that matters is whether a figure names its parameter, its window and its date, because one that does not cannot be re-derived by anyone, including its publisher. Every card this desk has scored in this category sits on the code-review tools hub, with what each score rests on written beside it.

Is CodeRabbit Worth It? The Verdict for a Team Deciding Whether to Roll It Out

Start from what the four cells actually settle. A public board using one method across eleven bots puts CodeRabbit's row at the top of the recall column and eighth on precision, at beta 1, on 15 August 2026, over 879 scored pull requests. A review platform puts it at 4.4 out of 5 across 97 reviews, on an unclaimed profile whose top entry is marked solicited. The entry seat costs $24 a month on the annual view and $30 month to month. And the contract defends the customer against third-party infringement claims about the Services and assigns output rights outright, which is the strongest reading on this card and the reason the composite is not lower.

Pay for it if pull requests are going unreviewed and the alternative is nobody looking. That is where every endorsement in this record comes from: a maintainer fielding volunteer contributions, a developer for whom review was the step that got skipped, a workflow where the pull requests are increasingly written by coding agents. Pay for it if a two-way indemnity and an assignment of output rights are what your own client contracts oblige you to hold, because those two clauses are unusually strong here and readable in five minutes at the links above. And pay for it knowing the trade you are buying: the row with the highest recall on that board is also the row eighth on precision, and both the maintainer's count and the most detailed G2 review describe the same consequence landing in a thread.

Do not roll it out to a team that will read every comment out of duty. Something over a quarter of the output landing in a maintainer's two lowest-value buckets is survivable for one person and expensive for six, and a review bot a team has started muting has cost its seat price twice. Do not buy it expecting a free path if your repositories are private: the free offer is the top tier given to public repositories, and everyone else has fourteen days. Do not assume the refund policy on the vendor's site governs your purchase, because if you subscribe through the Marketplace it says plainly that it does not, and the vendor's own listing counts over three hundred thousand installs through that route. And do not buy on the strength of any "#1 on Martian" post, the vendor's own included, without opening the board, reading the slider setting and checking the date.

What no number on this page settles is whether CodeRabbit is right about your code, in your language, on your architecture. The cross-repository limit the G2 reviewer describes, the 20-minute waits on large pull requests, the second pass that finds five new issues after the first six were resolved: each of those is one dated account and none of them is a measurement. Answering them would take a controlled run against a fixed repository under fixed conditions, and no such run stands behind any figure here. The reproducible part of the picture is above, with its parameters, its dates and its links attached, and the board it rests on resets monthly, so the capability row is the first thing to re-read at the next refresh.

Written and scored by Minel Gunesoglu, founder of Vouch — LinkedIn. Every source behind this page was read on 15 August 2026: the Martian online tracker at beta 1, coderabbit.ai/pricing in both billing views, the Terms of Service stamped 20 April 2026, the refund article stamped 18 December 2025, the changelog, the G2 and Trustpilot profiles and the GitHub Marketplace listing. Prices and the leaderboard row are re-checked monthly, and the leaderboard resets on its own schedule. Disclosure: no affiliate relationship exists with CodeRabbit, Inc., so every vendor link above is a plain link; where affiliate links appear elsewhere on this site they are marked, and a paid submission buys review speed only, never a listing and never a score.

Published 15 August 2026Last updated 26 August 2026

Scores and evidence on this page are re-checked monthly. Read about the person behind the scores, or find me on LinkedIn.

Licence terms, ownership and litigation status are reported here with the date they were read and a link to the source. They change, and a summary is not a clearance: nothing on this site is legal advice, and whether a particular use is safe for your work is a question for a lawyer in your own country. Verify a vendor’s current terms before you commit a deliverable to them.