The Opening Manifesto
Why most product reviews on the internet can't be trusted
And why the exceptions are worth paying attention to.
By Jeff · Published April 23, 2026 · Updated April 23, 2026 · 11 min read · 2,280 words
ManifestoJump to section
There's a polite version of the case against internet product reviews — that most of them are low-effort, that affiliate sites prioritize commissions over honesty, that AI is making everything worse. It's all true, and it's all insufficient. The real problem is that even the reviewers who care — the ones at established publications, with editorial teams and written ethics policies — are working inside a system that produces upward-biased recommendations almost regardless of what they intend.
This essay is about that system. Who it rewards, who it fails, and why breaking it takes more than better intentions.
The honest version of what's happening.
Most product reviews on the internet aren't fake. "Fake" is the easy story — paid shills writing fabricated reviews, fake Amazon profiles, astroturfing — and that story lets everyone else off the hook. The harder story is that the reviews written by real people at real publications, with real testing budgets and real editorial processes, still systematically mislead readers in predictable ways.
They mislead upward. Products that pay well get reviewed favorably more often than products that don't. The best products in any category aren't always the best-rated on review sites; the best-compensating products usually are. Critical reviews of popular products with strong affiliate programs are vanishingly rare. "Best of" lists tend to include exactly the products whose manufacturers run the most generous affiliate programs, often lightly reshuffled.
None of this requires corruption. It requires an ordinary incentive structure applied consistently over time. That's what makes it hard to fix: no single reviewer is doing anything wrong, and the aggregate outcome is still unreliable.
The math problem underneath everything.
Pull up any major product review site and look at the rating distribution of everything they've published in the last year. I've done this repeatedly, across multiple categories, for multiple publications. The distributions all look the same: a tall spike around 4 to 4.5 stars out of 5, a shorter tail at 3.5, almost nothing below 3, and a small cluster at 5.
This distribution is mathematically impossible as a representation of the products they reviewed. If you tested a random sample of consumer products in any category — all the coffee makers on the market, all the fitness trackers, all the sleep masks — you would find roughly a normal distribution. Some products are genuinely excellent, many are mediocre, and a meaningful percentage are straightforwardly bad. That's how any manufactured-goods market works, because that's how manufacturing variance works.
The reviews you're reading aren't a random sample. They're the intersection of (a) products worth publishing a review of — which already filters toward more-visible, better-marketed products — and (b) products the reviewer chose to keep testing and publish about, a second filter that excludes products that performed so poorly the reviewer lost interest. Both filters remove bad products from the set. What's left is a distribution skewed toward "at least decent," which then gets a rating distribution skewed toward "at least 4 stars."
If every product you review is great, you're not a reviewer. You're a catalog.
Layered on top of this: the commercial pressure on a rating of 3.5 or below is enormous. A 3.5-star review earns roughly the same click-through to a buy link as a 2-star review, which is to say almost none. A 4.5-star review converts. The reviewer's publication has no financial incentive to publish the 3.5 unless they have structural commitments to publishing across the full range — and almost none do.
Three incentive failures that shape what you read.
The problem compounds through three specific structural failures, none of which require any reviewer to act in bad faith:
- The selection bias on which products get reviewed at all. Reviewers don't pick products at random. They pick products readers are searching for, which means products with marketing budgets, which means products from companies that can afford to be visible — and those are disproportionately the products with affiliate programs. The unreviewed products are the ones with no marketing and no affiliate program, which includes a meaningful fraction of the actually-good products in any category.
- The publication bias on which reviews get published. A reviewer who tests a product for a week and hates it has spent the week learning that the product is bad. The cost of writing the negative review is roughly the same as the cost of writing a positive one, but the expected revenue is dramatically lower. Over time, publications develop an unspoken practice of quietly not publishing the worst reviews — not out of corruption, but because the opportunity cost is real and no one is fighting for the abandoned piece.
- The structure of "best of" lists. A category roundup of the "best wireless headphones" or "best standing desks" typically ranks 8–15 products. Because the products on the list are paid-partner products, the list isn't really a ranking of the best products — it's a ranking of the best products among those with active affiliate programs. The product that would actually be #1 for most readers might not be on the list at all.
Each of these failures is small. Compounded across thousands of reviews across hundreds of publications across multiple years, they produce a review ecosystem in which the reader is systematically pointed toward products that are good-enough-to-recommend and affiliate-program-visible, rather than toward the products that would actually serve them best.
The Wirecutter case: good intentions, structural constraints.
Let me name names here, because abstract critique is easier to dismiss than specific critique.
Wirecutter, owned by The New York Times, is the most-respected product recommendation site on the English-speaking internet. Their published editorial process is actually strong: they conduct extensive testing, they keep editorial and commercial teams separate, they explicitly refuse paid placements. Their editor-in-chief has been on record saying their product picks are "strictly and entirely through our editorially independent and journalistically rigorous process." Based on everything I've read from and about them, I believe this is a sincere and substantially accurate description of how they try to work.
And yet: Wirecutter is an affiliate-supported publication. Their picks convert to purchases via links that pay them a commission. Their catalog overwhelmingly consists of products from major brands whose affiliate programs Wirecutter is a member of. Their rating rubric rarely produces a truly negative review of a recommended product — instead, they recommend a product or they don't cover it at all.
"We only recommend products we love" is a marketing line that Wirecutter and most similar sites use some version of. It sounds responsible. Read carefully, it's a confession that the catalog is a filtered view: you're only seeing the reviewer's favorite products, not the ones they tested and rejected. The rejections aren't published. The decision-making that led to rejection isn't visible. What you see is, by design, the part of the testing process where the answer was "yes."
This isn't dishonest. It's structurally partial. A reader who only ever reads Wirecutter will end up with an accurate sense of which specific products Wirecutter likes, and an inaccurate sense of which specific products are actually best-in-category. Those aren't the same question.
The CNET problem: when the structure breaks openly.
Wirecutter is what the ecosystem looks like when the people running it are trying hard and still constrained. CNET, owned by Red Ventures through most of the 2020s, is what happens when the constraints win.
In January 2023, Futurism reported that CNET had quietly been publishing articles generated by an "internally designed AI engine" since November 2022, using the byline "CNET Money Staff" — disclosure visible only by hovering over the byline. By the time CNET acknowledged the practice, 77 such articles had been published. An internal audit, by CNET's own editor-in-chief, identified that 41 of those 77 articles required corrections. That's a 53% correction rate. One flagged article had told readers that $10,000 deposited in a 3% savings account compounding annually would earn $10,300 in a year — mathematically incorrect in a way that would materially mislead any reader making a real financial decision.
This was covered at the time by the Washington Post, CNN, The Verge, and most tech media. CNET paused the experiment and updated disclosure practices. The site was eventually sold in a transaction linked to the controversy.
The thing to notice isn't that CNET used AI. Lots of publications experiment with AI. The thing to notice is that a major technology-news site published 77 machine-generated articles in a row without clearly disclosing their origin, that more than half contained factual errors, and that this was only discovered because an outside publication looked. If the same thing were happening at a smaller affiliate site with no scrutiny, no one would know. The assumption that "established publication" equals "trustworthy process" is exactly the assumption that made this possible.
Nothing gets silently corrected on this site. Corrections are logged publicly, because the alternative is a history that pretends it was always right.
Tom's Guide and the mattress.
A third data point, because the pattern matters more than any single example. Tom's Guide, owned by Future plc, is one of the largest consumer technology publications in the English-speaking internet. Their published review methodology is thorough and in many ways commendable — they run benchmarks, they have category-specific test procedures, their reviewers have real credentials.
A 2021 Digiday profile of the site reported that "a mattress buying guide on Tom's Guide sells eight mattresses a day." That's a direct quote from their own publisher at the time, cited as a success of the site's commerce strategy. Eight mattresses a day, every day, generated by a single article.
This is not presented here as an accusation. Tom's Guide's writers work hard, and I have no reason to believe their mattress reviewers aren't honest. But consider what "eight mattresses a day from one article" implies about the structure: the article is a massive revenue-producing asset, kept alive with ongoing commerce optimization. When the article recommends a specific mattress over another, the stakes for the publication — the company, not the individual writer — are enormous. An alternative ranking that moved the top pick from a 20% commission mattress to an 8% commission mattress would cost the company significant revenue.
I'm not alleging that Tom's Guide rigs its rankings. I'm observing that any "best of" article that produces eight sales a day is, structurally, a massive commercial instrument, and treating it as pure editorial content requires either heroic editorial independence or enormous luck with which product turned out to be best.
What gets lost in this system.
The cost of an upward-biased review ecosystem isn't just that readers buy suboptimal products. The deeper cost is that the entire concept of "product review" gets devalued as a source of useful information.
When every review is positive, readers can't distinguish a genuine endorsement from boilerplate. When every "best of" list is similar, readers stop reading "best of" lists. When every affiliate site is bragging about its editorial independence, readers conclude that the claim of independence is itself meaningless — which is unfair to the publications where the claim is actually true, and which erodes the possibility of honest product recommendation as a genre.
The products that suffer most are the products made by small companies that don't run aggressive affiliate programs: companies that can't afford the marketing spend to get into the consideration set of reviewers, whose products might be genuinely superior to the incumbents but who have no mechanism for that to surface. The review ecosystem, in aggregate, is a subsidy from consumers to the best-capitalized brands. Honest reviewing would redistribute some of that subsidy to the better products.
Readers suffer too, in a more diffuse way: they stop trusting information they should be able to trust, because they've been burned by enough plausible-sounding reviews that turned out to be commercially motivated. The long-run equilibrium of a low-trust review environment is that most people give up and buy whatever has the most Amazon reviews, which introduces its own distortions. Nobody wins.
What this site does instead, and why that matters.
The structure of this site is designed to be incompatible with all three of the incentive failures above:
- Products are selected from what readers are searching for — including explicitly negative searches. When readers ask "is [product] a scam?" or "[product] complaints," that's a product I consider reviewing. Products with strong affiliate programs are candidates; products with no affiliate program are also candidates if reader interest is there.
- Negative reviews are published at the same rate as positive ones. The tested-and-rejected products go to the Graveyard — a permanent public log with its own route and schema. Every review that scored below 4.0 is listed there with a one-sentence summary of why. If the Graveyard is empty, the site is failing. If it grows, the reviews on the recommendation side become more credible.
- "Best of" lists aren't published until I've tested enough products in a category to have a defensible ranking. Until then, I publish individual reviews with clear verdicts, and let readers assemble their own comparisons. The category archives at /reviews will not include ranking lists until they can be honestly produced.
None of this fixes the affiliate-incentive problem at the system level. One site can't fix a structural flaw in an entire industry. What it can do is create a single, observable counter-example: a site where the stated commitments are verifiable in the published record. The corrections log, the editorial standards, the methodology, the Graveyard — these pages exist so that if this site ever starts drifting, the drift is visible before the trust has been extracted.
Whether this is enough is a question that will be answered by the reviews, not by this essay.
The first one is coming soon.
Continue reading
How to spot fake Amazon reviews: a 13-point checklist (2026)
Fake reviews have gotten harder to spot and more common than most shoppers realize. Here's a systematic method for identifying them — using language patterns, behavioral signals, and platform indicators that work even when the review looks completely legitimate.
15 min read
The testing window problem: why 30 days is the minimum for most products
Most affiliate product reviews on the internet were written after the reviewer used the product for less than a day — often less than an hour. That's not enough time to know if it actually works. Here's why the testing window is the single biggest predictor of whether a review is useful, and what 'properly tested' means for different product categories.
10 min read
Last updated Apr 23, 2026