GPTZero
· affiliate links never affect the score
- Capability20
- Usability & control48.6
- Purchase terms70.4
- Commercial terms61.5
By Minel Gunesoglu, founder. I ran no detection tests for this page and submitted no text to GPTZero. What I did was read three peer-reviewed studies that did run them, take the vendor's own accuracy claims off its home page, read its Terms of Use end to end, and read its pricing page in a browser with both billing views open. Every figure below is dated and traces to a saved copy of the document it came from. If you are here because a piece of your own writing was flagged, the part you want is the first section.
There is no single answer to the question in the title, and the absence of one is the finding rather than a failure to find one.
Four sources have published a number for how often GPTZero calls human writing AI-generated. The vendor's own site says 1%. A peer-reviewed comparison of eight detectors counted 52 false positives across 114 pieces of student writing collected before ChatGPT existed. Two further studies, in 2023 and 2025, report 10% and 16%. Those are not four attempts at the same measurement that happen to disagree. They are four different samples, and which one describes your situation depends entirely on whose writing you are feeding it.
A fifth measurement, the largest of them, asks a different question and answers it more kindly: tuned to a fixed error budget and then attacked, GPTZero degrades less than any other detector in its group. Both readings are below, because a page that printed only one of them would be choosing a conclusion rather than reporting the record.
Is GPTZero Accurate? What the Independent Record Actually Says
Start with the study that matters most to anyone who might be accused of something, because it is the only one built on writing that could not possibly have been generated.
GPTZero's False Positive Rate: 52 of 114 Essays Written Before ChatGPT Existed
In 2023 a team based across three universities (Toronto Mississauga in Canada, Maranatha Christian in Indonesia and Escuela Superior Politécnica del Litoral in Ecuador), published a comparison of eight AI detectors. It was posted to arXiv in July 2023 and peer-reviewed at IEEE COMPSAC in July 2024. One of its tables does something unusually clean: it runs every detector over 114 pieces of student writing collected before ChatGPT was released, so every flag is by definition a false positive.
| Detector | False positives out of 114 |
|---|---|
| CopyLeaks | 1 |
| GPT2 Detector | 2 |
| CheckForAI | 2 |
| AI Text Classifier | 6 |
| OriginalityAI | 7 |
| GLTR | 20 |
| GPTZero | 52 |
That is 45.6% of a corpus that contained no AI writing at all, and it is the highest count in the table by more than a factor of two over the next tool.
Two things have to be said immediately, and neither of them softens the number. The first is that this is 2023-era measurement of a 2023-era product, and GPTZero has shipped continuously since. The second is that a false-positive count on pre-ChatGPT student prose is a hard test on purpose: it is the case where a detector has nothing to find and every flag is an error. A tool tuned to catch AI aggressively will look worse here than on a balanced set, and that trade-off is real. Which of those two framings matters to you depends on whether you are the person running the scan or the person being scanned.
The same study measured two more things worth carrying. On accuracy over ChatGPT-generated submissions before any paraphrasing, GPTZero read 70.00% against 100.00% for both Copyleaks and Originality.ai. On a weighted view across the whole set it read 88.00%, second best of the eight.
GPTZero Accuracy in Two More Studies: 10% and 16%, on Different Writing
The picture does not get simpler with more papers, and printing only the worst one would be its own kind of dishonesty.
A 2023 paper by F. Habibzadeh, cited 97 times as of this reading, reports that GPTZero has a 10% false-positive rate and a 35% false-negative rate, on a sample of 20 pieces of text. A 2025 arXiv paper by S. Dik and colleagues reports a 16% false-positive rate, eight out of fifty human essays, contributing to an overall error rate of 10.3% across seventy-eight documents.
So the independent range runs from 10% to 45.6%, and the spread is not noise. It is the corpus. Twenty texts, fifty essays and 114 pre-ChatGPT student submissions are three different questions wearing the same words, and a detector's error rate is a property of the pairing rather than of the tool alone. What survives all three is that the figure is not small, and that no published independent measurement puts it near 1%.
GPTZero Accuracy on RAID: 66.5 at a Fixed Error Budget, and the Smallest Fall Under Attack
The largest and most recent independent measurement points the other way, and leaving it out would have made this page a selective reading.
RAID, published at ACL 2024, tests twelve detectors over 509,000 generations. Its method matters more than any single figure: rather than reporting raw accuracy, it tunes each detector to a fixed false-positive rate of 5% and then measures how much machine text it catches at that setting, so no tool can look good by simply refusing to flag anything. On non-adversarial text GPTZero scores 66.5 under that rule, which puts it fifth of the six detectors the paper singles out for its attack table. Every figure in that table is RAID's accuracy at a fixed 5% error budget and none of them is a score from this site.
Then the paper attacks the text, and the ordering changes:
| Detector | Clean | Worst result under attack |
|---|---|---|
| Originality.ai | 85.0 | 9.3 (homoglyphs, a fall of 75.7) |
| Binoculars | 79.6 | 37.7 (homoglyphs, a fall of 41.9) |
| RADAR | 70.9 | 59.3 (a fall of 11.6) |
| GPTZero | 66.5 | 61.0 (a fall of 5.5) |
| GLTR | 62.6 | 24.3 (a fall of 38.3) |
GPTZero starts second from the bottom of that group and finishes second from the top. Its worst fall across six attacks is 5.5 points, where the highest-scoring tool in the table loses 75.7 to a homoglyph substitution.
Two limits keep that from being a verdict. A detector that moves very little may be steady or may simply be less sensitive, and this table cannot separate those. And elsewhere in the same paper, under changes to how the text was generated rather than to the text itself, GPTZero does collapse: two cells of another table read 9.4 and 4.8. It has company there (Binoculars falls to 0.6, GLTR to 0.5, ZeroGPT to 0.3), which is the paper's own conclusion, that detectors in this category do not generalise across generators and decoding settings.
One figure from RAID that a careless reading would misuse: the paper also prints raw false-positive rates at naive thresholds, where GPTZero shows 0.03%. That is not evidence it rarely misfires. It is what a detector's error rate looks like at a threshold where it flags very little, and it is exactly why the paper fixes the false-positive rate at 5% before comparing anything.
None of this enters the score. The capability figure on this card stays on the 2023 study because that is the one all three cards in this category were scored from, and RAID publishes no single accuracy per detector on its main English dataset that could replace it like for like.
What GPTZero Publishes About Its Own Accuracy
The vendor's home page, read 23 August 2026, is specific. Answering its own FAQ on whether the tool is accurate, it says that "GPTZero is the most accurate AI detector, with a 99% accuracy rate when spotting AI-generated text vs. human writing". Elsewhere on the same page:
GPTZero was shown to be the most accurate AI detector in North America, detecting 95.7% of AI texts while only incorrectly predicting 1% of human texts as AI
It also reports 96.5% accuracy on mixed human-and-AI documents, and says plainly that "no AI detector can ever truly be 100% perfect." Worth noting how it frames its own evidence: the 99% figure is attributed to "Independent benchmarks, like our partners at Penn State's AI Research Lab, and our own large-scale testing". A partner's benchmark and the vendor's own testing are both legitimate sources; neither is independent in the sense the studies above are.
Set the 1% beside the 10%, the 16% and the 45.6% and the honest reading is not that somebody is wrong. It is that a false-positive rate is only meaningful with its corpus attached, and the vendor's figure comes from the vendor's own testing while the other three come from samples the vendor did not choose. This page takes no position on why they differ and makes no claim about how the vendor's testing was conducted. It reports that a buyer comparing them is comparing four different experiments.
GPTZero Reviews: 4.3 on One Platform, 2.2 on Another, the Same Week
The customer record splits the same way, and for a reason the platforms themselves publish.
G2 rates GPTZero 4.3 out of 5 across 101 reviews. Its per-review metadata carries the strings "Validated Reviewer", "Incentivized" and "G2 invite". This is a population the vendor invited.
Trustpilot rates it 2.2 across 138 reviews, on a profile the company claimed in November 2024, under a banner reading: "This company hasn't invited customers recently, so reviews may not be representative." This is a population that arrived on its own.
Since August 2026 this site does not average consumer records whose solicitation regimes disagree, because doing so hides the only thing the gap is telling you. The cell on this card is filled from the Trustpilot record, the larger of the two, and the G2 figure is reported here beside it rather than blended into it. A third number exists and belongs on the record: GPTZero's own site footer displays 4.7 across 707 reviews, which is the vendor's widget and is not independent of the vendor.
How GPTZero's 45.3 Score Was Built
| Dimension | Score | What it read |
|---|---|---|
| Capability | 20.0 | Accuracy after QuillBot paraphrasing, arXiv 2307.07411, n=10 |
| Usability & control | 48.6 | Trustpilot 2.2 across 138 reviews, 23 August 2026 |
| Purchase terms | 70.4 | Four positions at the point of purchase, 23 August 2026 |
| Commercial terms | 61.5 | Two clause positions in the Terms of Use, 23 August 2026 |
Those four rows combine into 45.3, and it is the lowest total this site currently publishes. The arithmetic is a geometric mean, which is why a single very weak row pulls the whole card down rather than being averaged away by three ordinary ones, and on this card the very weak row is capability.
Read it beside its category before you read it as a verdict on this product specifically. Copyleaks scores 69.6 and Originality.ai scores 52.6. Every detector measured here scores below the site's median, and the reason is structural rather than particular: the only capability evidence that exists for this category is academic accuracy studies, those studies test detectors against paraphrasing tools, and every detector tested degrades sharply under paraphrase. This is a category-wide finding and it belongs in the same paragraph as the number.
GPTZero Accuracy After Paraphrasing: Where the Capability Row Comes From
The capability figure comes from Table 9 of the same study, which measures each detector twice on ten ChatGPT submissions: once as generated, once after running them through QuillBot.
| Detector | Before paraphrasing | After |
|---|---|---|
| GLTR | 100.00% | 100.00% |
| GPT2 Detector | 100.00% | 60.00% |
| CopyLeaks | 100.00% | 50.00% |
| CheckForAI | 100.00% | 40.00% |
| OriginalityAI | 100.00% | 40.00% |
| AI Text Classifier | 60.00% | 20.00% |
| GPTZero | 70.00% | 20.00% |
Both published siblings on this site are scored from this same table, same arm, same ten submissions: Copyleaks at 50.0 and Originality.ai at 40.0. That makes these three cards comparable row for row rather than merely on the same scale, which is rare here and worth more than a larger, less commensurable figure would have been.
The caveats travel with it. Ten submissions is a thin sample and is stated as thin wherever the number appears. The study is from 2023 and predates every model these tools now advertise detecting. And a blind academic comparison is not a test this desk ran.
GPTZero Pricing Could Not Be Read in Dollars, So the Value Row Was Renamed
This one is a limitation of the reading, not of the product, and saying so precisely matters.
GPTZero publishes an ordinary monthly subscription price, which is exactly the quantity this site's value ladder is built for. It could not be read in that ladder's currency from here. The pricing page geolocates, this desk works from Turkey, and every figure it served was Turkish Lira. Converting lira into dollars would print a price the vendor never published, so it was not done, and the Wayback copy of 19 August 2026 does not rescue it because the plan prices are written into the page by script and the archive captured the shell.
So the row publishes a different measurement under its own name: what a buyer meets at the point of purchase. A reading through a US connection would restore the ordinary value row, and that is a refresh item on this desk rather than anything about GPTZero.
GPTZero Pricing, Free Tier and Refunds: What a Buyer Signs
Is GPTZero Free? The Free Tier, and the Trial That Charges Itself
There is a free tier and it is real: GPTZero's detector runs without payment, which is why the product is as widely used as it is.
What the paid ladder costs could only be read in lira here, so the figures are given as served rather than converted. Premium reads TRY 999 per month on monthly billing, TRY 549 per month billed annually, and TRY 384.30 per month with the promotional code BTS26. Professional reads TRY 1,949, TRY 1,049 and TRY 734.30 on the same three bases.
The toggle above those cards is labelled "Annual (Save 62%)". That 62% is only reachable with the back-to-school code, which the page itself marks as a limited-time offer; annual billing alone is 45%, a figure the same page prints in its team block. And section 6 of the Terms of Use adds the clause that most often surprises people: "If the user initiates a free trial, the account will be charged according to the user's chosen subscription at the end of the free trial."
GPTZero Refund and Cancel Subscription: Four Words in Section 7
Section 7 is short enough to quote whole:
All purchases are non-refundable.
Cancellation is self-serve and takes effect at the end of the paid term. An email address is offered for anyone unsatisfied, which is an invitation rather than a published policy, so this card scores the clause as written rather than as a discretionary refund scheme.
The Accuracy Disclaimer in GPTZero's Own Terms
Two more clauses decide what this product is contractually, and both are worth reading beside the marketing.
Section 21 disclaims warranties in capitals:
WE MAKE NO WARRANTIES OR REPRESENTATIONS ABOUT THE ACCURACY OR COMPLETENESS OF THE SITE'S CONTENT
Section 22 caps what could ever be recovered:
OUR LIABILITY TO YOU FOR ANY CAUSE WHATSOEVER AND REGARDLESS OF THE FORM OF THE ACTION, WILL AT ALL TIMES BE LIMITED TO THE AMOUNT PAID, IF ANY, BY YOU TO US DURING THE SIX (6) MONTH PERIOD PRIOR TO ANY CAUSE OF ACTION ARISING
Section 23 runs the indemnity one way, from user to vendor, with nothing running back. That is the ordinary shape of this market, found in 19 of the 25 agreements read across this site, and it scores where ordinary scores.
Put the three together and the contractual position is legible: the accuracy figure is marketing, the agreement promises nothing about accuracy, and the ceiling on any claim is what you paid in the last six months. Both halves of that are the vendor's own words, published on the same domain, and this page draws no conclusion about why they sit differently. What it does say is that anyone planning to act on a GPTZero score should read section 21 before they do.
On what you submit, section 9 is unambiguous and in the user's favour: "We do not assert any ownership over your Contributions. You retain full ownership of all of your Contributions."
GPTZero for Teachers: What a Score Can and Cannot Support
GPTZero's own pricing page states that it is "Trusted by over 10 million teachers and students" and describes the company as "Official AI detector partners of the American Federation of Teachers". Those are the vendor's claims, transcribed as such; neither is independently verified here and neither feeds any figure on this card.
The decision they bear on is narrow and important. If a score is going to start a conversation with a student, three published facts belong in that conversation before it starts: the independent false-positive range runs from 10% to 45.6% depending on the corpus, the worst of those figures came from writing that predated ChatGPT entirely, and the vendor's contract disclaims warranties about accuracy. None of that argues against using a detector. It says a number from one is evidence to weigh rather than a finding to act on, and the vendor's own agreement says the same thing in legal language.
GPTZero Alternatives: What Else Carries a Card Here
Two other detectors on this site carry scored cards built from the same study and the same four dimensions: Copyleaks and Originality.ai. The AI detectors hub orders them by composite, computed at build time rather than written by hand.
The one thing to carry across all three: on the false-positive table, Copyleaks flagged 1 of 114 and Originality.ai flagged 7, against GPTZero's 52. That is the sharpest separation between these products in any evidence this desk holds, and it is nine months older than anything else on this page. Originality.ai has its own contractual problem in exchange, documented on its card, so a low false-positive count is not the whole decision either.
Is GPTZero Worth Paying For? The Verdict
Use it as one signal among several, on writing you can already form a judgement about, where a flag prompts a question rather than a conclusion. The free tier makes that cheap to do.
Look harder before committing if you are buying it for an institution. The purchase terms are ordinary but the ladder was only readable in lira from here, all purchases are non-refundable, and a free trial charges itself at the end.
Do not act on a score alone, and this is the one place where the evidence, the vendor and the contract all agree. Independent measurements put the false-positive rate between 10% and 45.6% depending on whose writing was tested; the vendor says no detector is ever perfect; and section 21 disclaims warranties about accuracy outright.
If your deciding question is how well it performs on the specific writing you deal with, no figure here answers it, and neither does any figure the vendor publishes. Every number on this page describes somebody else's sample.
Every figure above is dated, and the methodology explains how each one was taken. The about page says who took them.
Published 23 August 2026Last updated 26 August 2026
Scores and evidence on this page are re-checked monthly. Read about the person behind the scores, or find me on LinkedIn.
Licence terms, ownership and litigation status are reported here with the date they were read and a link to the source. They change, and a summary is not a clearance: nothing on this site is legal advice, and whether a particular use is safe for your work is a question for a lawyer in your own country. Verify a vendor’s current terms before you commit a deliverable to them.