Reader-supported. We may earn a commission from links, at no cost to you. How scoring works →

How we score AI tools

Principles

The four dimensions

DimensionWhat it answers, and where the number comes from
CapabilityHow good is the core output? Sourced from public output-quality arenas (thousands of blind pairwise votes) or from a published third-party accuracy study where no arena exists.
Usability & controlHow easy is it to get what you want? Aggregated from capability-focused review crowds (G2, Product Hunt, Capterra), weighted by sample size.
ValueWhat does a usable result actually cost? Verified official pricing, free-tier availability, billing terms.
Commercial termsWhat the vendor's own documents permit, on which tier: which way the indemnity runs, who owns the output on the entry paid tier, whether the free tier may be used commercially, and the published refund posture. Weighted 0.35 / 0.25 / 0.20 / 0.20 over whichever of the four the documents answer, and blank below two. Every value is a named clause position read off a saved copy of the document, on a stated date, not our verdict on the clause, and not a claim that the vendor honours it.

Bands are published so a number means the same thing everywhere: 90–100 exceptional · 75–89 strong · 60–74 adequate · 40–59 weak · below 40 poor.

When a dimension’s own evidence does not exist

Some tools cannot be measured on some axes, and the reason is almost never the tool. No blind-vote arena grades avatar video, deck generation, product-photo editing or instrumental music. Some products have no review sample that can honestly be attributed to them. Some are billed per second of output, which cannot be placed on a ladder calibrated in dollars per month. We work through four steps in this order, and each step may only use what the step before it produced:

  1. The axis’s own evidence: an arena, a review aggregate, a published price, the contract.
  2. Our own controlled battery, where one has run. None has; evidence grade A is empty site-wide.
  3. A substitute: a different evidence type answering the same buyer question. The cell keeps its name, because the column still reads like-for-like. Capability is filled this way on some cards, from a published accuracy study where no arena covers the category.
  4. A named criterion: a different measurement answering a different question, and then the row takes that criterion’s name on that card, not the axis’s.

Three named criteria are in use, all adopted on 19 August 2026. Delivery envelope asks what the entry paid tier lets you take away: fidelity ceiling, export formats, watermarking, whether delivery can be automated. Access terms asks what the vendor publishes about getting and keeping access: regional availability, usage limits, support channels, service status. Purchase terms asks what a buyer faces at the point of purchase: how the price is published, what activation commits you to, whether a failed generation is billed, whether a balance expires. Each has a fixed rubric of clause-descriptive values, weights renormalised over whichever positions the documents answer, and a floor of two positions before it may publish at all. Every value traces to a saved copy of the document it was read from, and the engine refuses a value outside its rubric, a position with no saved source, or a capture past its stated expiry.

A renamed row is not the axis it sits in, and reading it as one is the mistake this mechanism exists to prevent.A delivery-envelope figure says nothing about output quality: a product handing back a large, unmarked, automatable file of poor work scores above one handing back a small file of good work. So the card prints the criterion’s name and its question where the number is, rather than in a footnote, and a composite containing a renamed row is not the geometric mean of the four dimensions named above. It does not read like-for-like against a card whose rows all answer their own questions, and the pages that carry one say so. The alternative, a hole, tells a reader even less, because a hole cannot be weighed at all.

Evidence grades

Every dimension carries a grade for how we know:

Thin or contested evidence lowers the published confidence; a vendor’s unverified claim can never lift a score on its own.

What we do not do

The omissions matter as much as the method, so they are stated rather than implied:

From dimensions to one number

Dimension scores are normalized against absolute anchors: so a rival’s launch never moves a tool’s score, and combined with a weighted geometric mean, which punishes lopsided profiles harder than a simple average would. A weak dimension therefore costs more than it would in an average, and it costs it from inside the table: there is no adjustment sitting outside the rows. A poor customer record used to be the exception. It was allowed to override the whole composite from outside, and that is gone. It now fills the usability row like any other evidence, weighted by how many people stand behind it, so a reader can point at the cell it moved instead of taking our word for a limit. A rating drawn from a sample too small to carry it is reported as thin evidence and lowers the published confidence rather than the score. Finally, no published score is allowed to reach 100. The composite is secondary output: the four-dimension profile with its confidence is the review.

Version and disclosure

This page describes Vouch Score v11. The methodology is versioned; every published score states the version that produced it. We are reader-supported and may earn an affiliate commission when you buy through our links, at no extra cost to you, no score, no placing and no badge is ever sold, and paid submissions buy review speed, never a listing or a score.