The problem with reading reviews by hand
A product with 2,000 reviews holds the answer to almost every question a seller has: why the rating slipped, what buyers compare you to, which words they use, what they expected and did not get. Reading them takes a full day, and the result is a feeling, not a number. Read the first page and you get the most recent reviews. Read the most helpful ones and you get the oldest.
The sample decides everything
Amazon shows at most 100 reviews per star rating, and "all stars" counts as one filter. So a tool that asks for "the reviews" gets 100 of them, mostly five stars, and concludes that everything is fine.
Collect each star rating separately and the ceiling becomes 500: up to 100 one-star, 100 two-star, and so on. That sample finds the problems a 4.7-star product hides, because it deliberately reads the unhappy buyers.
Then weight it back, or every figure lies
That same sample no longer looks like the product. A product whose ratings are 82% five stars gives a sample that is only 20% five stars. Counted as it is, the average rating drops to 3.0 and every percentage makes the product look far worse than it is.
The fix is what polling institutes do: give each star group its real share of the product's ratings, split among the reviews of that group. In our reports, the same sample gives 3.0 raw and 4.67 weighted, against 4.6 on Amazon.
Count in code, never with a language model
AI reads well and counts badly. Ask it how many reviews mention a problem and it will give you a plausible number. In a report someone is going to act on, that is worse than useless.
So the model's job is to read each review and say what it is about. Every count, every percentage and every average is then computed by code from what it found, and every quote is matched word for word against the review it came from. A quote that does not match is removed rather than shown.
What a good analysis tells you
- Each complaint, with how often it comes up, how serious it is, its likely cause, and what it costs you.
- What buyers love, in their words, so you can say it on the listing.
- The words they use: the vocabulary for your title, your bullets and your search terms, counted in real reviews.
- The objections to answer before the sale, which is what turns a review problem into a returns problem.
- The order of work: what to fix first, placed by impact and by effort.
Two mistakes to avoid
Reading a trend that is not there
In a sample taken star by star, the age of a review depends on its star: the five-star reviews are recent, the one-star reviews go back years. Plot a rating over time on that sample and you will see a crash followed by a recovery. It is the shape of the sample, not of the product. We removed that chart from our own reports for exactly this reason.
Trusting the topics alone
Amazon now shows topics on the product page, free, in its "Customers say" box, and Seller Central shows more in Voice of the Customer. A list of topics is a start. It does not tell you what it costs you, why it happens, or what to change first.
How we do it
We collect up to 500 reviews, 100 per star rating, read each of them with AI, compute every figure in code, weight everything to the product's real rating mix, and write a report you can send to a supplier or an agency: 6 problems with buyer quotes, the fixes in order, listing keywords, and a SWOT. Ready in under 10 minutes, $39, no subscription and no account.
The sample report on this site was run on a best-selling clear iPhone case sold as "Not Yellowing". Its most common complaint, in 39 of the 473 reviews we read, is that it turns yellow.