How we score AI tools
Principles
- Independence. Anchors, weights and gates are fixed before scoring and identical for every tool in a category. No score, no placing and no badge is ever sold, and no affiliate payout weights a dimension.
- Evidence over vibes. Every dimension score is tagged with how it was obtained, and a number we could not measure never becomes a guessed one. Where the usual evidence is missing, we go looking for a different kind that answers the same buyer question, and the card names which kind it used, because two evidence types do not read like-for-like across cards.
- An honest ceiling. No tool gets a perfect score, by design. Even the best-evidenced cards land in the mid-to-high range.
- No covering up weak spots. We combine dimensions with a geometric mean, so a brilliant score on one axis cannot average away a weak one on another. A poor customer record gets no special exemption either, and no override sitting outside the table: it goes into the usability row, weighted by how many people stand behind it, where a reader can point at it.
- Freshness. Every score is pinned to a collection date and the tool version measured. Aged scores are marked stale and re-collected.
The four dimensions
| Dimension | What it answers, and where the number comes from |
|---|---|
| Capability | How good is the core output? Sourced from public output-quality arenas (thousands of blind pairwise votes) or from a published third-party accuracy study where no arena exists. |
| Usability & control | How easy is it to get what you want? Aggregated from capability-focused review crowds (G2, Product Hunt, Capterra), weighted by sample size. |
| Value | What does a usable result actually cost? Verified official pricing, free-tier availability, billing terms. |
| Commercial terms | What the vendor's own documents permit, on which tier: which way the indemnity runs, who owns the output on the entry paid tier, whether the free tier may be used commercially, and the published refund posture. Weighted 0.35 / 0.25 / 0.20 / 0.20 over whichever of the four the documents answer, and blank below two. Every value is a named clause position read off a saved copy of the document, on a stated date, not our verdict on the clause, and not a claim that the vendor honours it. |
Bands are published so a number means the same thing everywhere: 90–100 exceptional · 75–89 strong · 60–74 adequate · 40–59 weak · below 40 poor.
When a dimension’s own evidence does not exist
Some tools cannot be measured on some axes, and the reason is almost never the tool. No blind-vote arena grades avatar video, deck generation, product-photo editing or instrumental music. Some products have no review sample that can honestly be attributed to them. Some are billed per second of output, which cannot be placed on a ladder calibrated in dollars per month. We work through four steps in this order, and each step may only use what the step before it produced:
- The axis’s own evidence: an arena, a review aggregate, a published price, the contract.
- Our own controlled battery, where one has run. None has; evidence grade A is empty site-wide.
- A substitute: a different evidence type answering the same buyer question. The cell keeps its name, because the column still reads like-for-like. Capability is filled this way on some cards, from a published accuracy study where no arena covers the category.
- A named criterion: a different measurement answering a different question, and then the row takes that criterion’s name on that card, not the axis’s.
Three named criteria are in use, all adopted on 19 August 2026. Delivery envelope asks what the entry paid tier lets you take away: fidelity ceiling, export formats, watermarking, whether delivery can be automated. Access terms asks what the vendor publishes about getting and keeping access: regional availability, usage limits, support channels, service status. Purchase terms asks what a buyer faces at the point of purchase: how the price is published, what activation commits you to, whether a failed generation is billed, whether a balance expires. Each has a fixed rubric of clause-descriptive values, weights renormalised over whichever positions the documents answer, and a floor of two positions before it may publish at all. Every value traces to a saved copy of the document it was read from, and the engine refuses a value outside its rubric, a position with no saved source, or a capture past its stated expiry.
A renamed row is not the axis it sits in, and reading it as one is the mistake this mechanism exists to prevent.A delivery-envelope figure says nothing about output quality: a product handing back a large, unmarked, automatable file of poor work scores above one handing back a small file of good work. So the card prints the criterion’s name and its question where the number is, rather than in a footnote, and a composite containing a renamed row is not the geometric mean of the four dimensions named above. It does not read like-for-like against a card whose rows all answer their own questions, and the pages that carry one say so. The alternative, a hole, tells a reader even less, because a hole cannot be weighed at all.
Evidence grades
Every dimension carries a grade for how we know:
- A: controlled test. A repeatable benchmark run by us. No published score currently carries an A. When one does, the page will say which battery ran and on what date.
- B: measured public data. Blind-vote arenas and review aggregates (G2, Trustpilot, Product Hunt, Capterra) with sample sizes and dates disclosed. Sources are routed to the dimension they actually measure, so billing complaints never contaminate a capability score.
- C: vendor documentation. Official pricing and legal terms, verified on a stated date.
Thin or contested evidence lowers the published confidence; a vendor’s unverified claim can never lift a score on its own.
What we do not do
The omissions matter as much as the method, so they are stated rather than implied:
- We do not claim hands-on testing we have not done. Today every published score rests on grade B and C evidence: public arena data, sample-weighted review aggregates, and vendor pricing and licence terms read as written. Where a page would need a controlled run to say something, it says the gap instead.
- We do not guess a missing number. Where a dimension’s own evidence does not exist, the row publishes a different measurement under a different name rather than an impression under the axis’s name: the four steps above. The commonest case is usability, where a review sample large enough to convert is not retrievable for many vendors, often for a reason that has nothing to do with the tool: the platforms holding it return a bot check to us. A refusal is recorded as silence, never as a zero, and it never becomes a low score by default.
- We do not keep an axis we cannot fill. A fifth dimension, reliability, was carried until 12 August 2026 and was never once measured on any tool: it needs a controlled repeat-run, and we run none. A row that can only ever read “not measured” advertises a measurement as pending when it is not, so the dimension was removed rather than left standing. No score changed: the composite has always combined only the dimensions that carry a number.
- We do not keep an axis that answers two questions at once. The fourth dimension used to average a consumer-review rating with a one-word summary of a vendor’s indemnification clause and call the result “safety & trust”. On 12 August 2026 we read the actual contracts for every tool on the site and found six of those summaries wrong, two of them live: because the field recorded our verdict about a clause rather than what the clause said, so a wrong entry looked exactly like a right one. The dimension is now commercial terms: four clause positions, each named after what the document does, each traceable to a saved copy of that document. Every affected composite moved and the before-and-after is published rather than absorbed.
- We do not sell position or score. Paid submission buys review speed and nothing else, not a listing, not a placement, not a number. Affiliate relationships are disclosed where the link appears and never weight a dimension.
- We do not write head-to-head pages. There are no “X vs Y” pages here and no head-to-head sections inside a review; the two we had published were withdrawn on 11 August 2026. A category page orders its own cards by composite and marks the highest as the top score, and that is the whole of the ordering we do: a placing is a restatement of one number, not a duel between two tools.
- The top score label is arithmetic, not an opinion. The label goes to the highest composite in that category on the date shown, computed when the page is built. Nobody types it into a file, so it moves on its own the day a tool is rescored or a higher-scoring one publishes. Two tools level at the top are shown as level and neither is marked. A category with one card is not a ranking and gets no number and no badge.
- Read the placing with the row headings beside it. A composite is the geometric mean of the rows on that card, and cards differ in what those rows measure: so a card carrying a renamed row can sit above one whose rows all answer their own questions without having measured better. That is why each row prints its own heading and, where it is renamed, the question it answers, next to the number rather than in a footnote; a card whose composite rests on fewer than three rows also prints that count beside the score. The cross-category ledger at /reviews is listed by name, never by score: those numbers answer different questions.
- We do not crown tools in prose. No “market leader”, no “industry standard”, no “the best AI image generator”. The one superlative on the site is the computed top-pick badge, which says what it measured and where.
- We do not quietly correct. When a number turns out wrong, it is fixed with a new collection date, and the score is re-run rather than nudged.
From dimensions to one number
Dimension scores are normalized against absolute anchors: so a rival’s launch never moves a tool’s score, and combined with a weighted geometric mean, which punishes lopsided profiles harder than a simple average would. A weak dimension therefore costs more than it would in an average, and it costs it from inside the table: there is no adjustment sitting outside the rows. A poor customer record used to be the exception. It was allowed to override the whole composite from outside, and that is gone. It now fills the usability row like any other evidence, weighted by how many people stand behind it, so a reader can point at the cell it moved instead of taking our word for a limit. A rating drawn from a sample too small to carry it is reported as thin evidence and lowers the published confidence rather than the score. Finally, no published score is allowed to reach 100. The composite is secondary output: the four-dimension profile with its confidence is the review.
Version and disclosure
This page describes Vouch Score v11. The methodology is versioned; every published score states the version that produced it. We are reader-supported and may earn an affiliate commission when you buy through our links, at no extra cost to you, no score, no placing and no badge is ever sold, and paid submissions buy review speed, never a listing or a score.