Working dataset · 8 analyses
Half a billion Amazon reviews
507,730,787 reviews across 33 product categories, from June 17, 1996 to September 14, 2023, reduced to counts, means, and distributions. No review text, no user IDs, no product IDs — which makes it a corpus you can hand to a class on day one and still learn something real from.
- reviews
- 507.7M
- categories
- 33
- years covered
- 28
- mean rating
- 4.19★
- five-star
- 65%
- verified
- 90%
The distribution
A mean of 4.19 describes almost nothing
Star ratings are not bell-shaped. Across all 507.7M reviews, 75.6% of the mass sits at the two ends of the scale and only 24.4% sits in the middle three. The distribution is J-shaped: 5★ is the mode, 1★ is the runner-up, and the arithmetic mean lands where almost nobody actually rates.
Share of all reviews by star rating
507,730,787 reviews, 33 categories, June 17, 1996 – September 14, 2023.
- 5★65.5%332.5M
- 4★12.6%63.9M
- 3★7.0%35.6M
- 2★4.9%24.7M
- 1★10.1%51.1M
The categories
Three categories are a third of the corpus
Home & Kitchen, Clothing Shoes & Jewelry, Electronics together hold 34.9% of every review ever written. Subscription Boxes, the smallest slice, holds 16,216 — 4,157× fewer than Home & Kitchen. Compare categories on rates, never on raw counts.
All 33 categories
Sort by any column.
- 1Home & Kitchen67.4M4.17★93% v
- 2Clothing Shoes & Jewelry66M4.18★94% v
- 3Electronics43.9M4.10★92% v
- 4Books29.5M4.42★70% v
- 5Tools & Home Improvement27M4.16★94% v
- 6Health & Household25.6M4.20★93% v
- 7Kindle Store25.6M4.43★68% v
- 8Beauty & Personal Care23.9M4.11★91% v
- 9Cell Phones & Accessories20.8M4.01★95% v
- 10Automotive20M4.18★96% v
- 11Sports & Outdoors19.6M4.22★93% v
- 12Movies & TV17.3M4.25★79% v
- 13Pet Supplies16.8M4.09★94% v
- 14Patio Lawn & Garden16.5M4.05★94% v
- 15Toys & Games16.3M4.21★91% v
- 16Grocery & Gourmet Food14.3M4.12★92% v
- 17Office Products12.8M4.21★93% v
- 18Arts Crafts & Sewing9M4.23★95% v
- 19Baby Products6M4.21★90% v
- 20Industrial & Scientific5.2M4.18★95% v
- 21Software4.9M3.94★95% v
- 22CDs & Vinyl4.8M4.50★68% v
- 23Video Games4.6M4.05★86% v
- 24Musical Instruments3M4.26★92% v
- 25Amazon Fashion2.5M3.97★94% v
- 26Appliances2.1M4.22★96% v
- 27All Beauty702K3.96★91% v
- 28Handmade Products664K4.50★95% v
- 29Health & Personal Care494K4.00★90% v
- 30Gift Cards152K4.55★93% v
- 31Digital Music130K4.53★74% v
- 32Magazine Subscriptions71K4.04★82% v
- 33Subscription Boxes16K3.77★88% v
Verification
The least-verified categories are the best-rated
89.9% of reviews carry a verified-purchase flag — but that share collapses in media. Books, Kindle, CDs & Vinyl, Digital Music, and Movies & TV average 71.2% verified against 93.3% everywhere else, and they rate 4.39★ against 4.15★. It is tempting to conclude that unverified reviewers are the generous ones. They are not — see below.
Verified-purchase share against mean rating
One dot per category, sized by review volume. Media categories in amber.
At the review level the relationship reverses. Pooled across all 507.7M reviews, verified purchases average 4.191★ and unverified ones 4.137★ — verified reviewers are the marginally kinder group, and that holds in 29 of 33 categories.
So the chart above is Simpson’s paradox, not a finding about reviewer generosity. Media categories are both less-verified and better-rated, for reasons that have nothing to do with each other: they attract enthusiast raters, and their reviews often predate or bypass an Amazon purchase. Comparing categories recovers the composition; comparing reviews recovers the behaviour, and the two point opposite ways.
A category-level summary cannot show this. It reports the rating distribution and the verified share as two separate margins, and a joint cannot be recovered from margins — you have to count the pairs. The review-anatomy analysis has the full 5 × 2 table by year and category.
Analyses
Questions asked of this corpus
Each analysis states what slice it is computed over. The aggregate pages cover all 507.7M reviews; others work from smaller samples where the review text itself is needed.
- Reviewer-level aggregates · 507.7M reviews
The rater in the rating
Who actually writes reviews, and how much of a star rating is about them rather than the product. One-time reviewers give one star 2.5× as often as regulars.
selectionreviewersinequality10 minRead the analysis - Product-level aggregates · 35M items
How a product’s rating forms
Rating by review index, the first-review effect, and the shape of a contested product. Items whose first review was 1★ run half a star lower forever after.
herdingitem dynamicscross-category9 minRead the analysis - Review-level aggregates · 507.7M reviews
What a review is made of
Helpful votes, length, photos, duplicate text, and the verified-purchase joint that reverses the headline correlation.
helpfulnesstextsimpson’s paradox8 minRead the analysis - Daily series · 1996–2023
Ten thousand days
The full daily series, 1996–2023. The biggest review days in Amazon’s history are not Prime Day or Black Friday — they are the first week of January.
time seriesseasonalityevents6 minRead the analysis - Cross-category matrices · 54M reviewers
What reviewers buy next
Category breadth, the 33×33 co-occurrence matrix, and where a reviewer goes after their last review.
cross-categorynetworksreviewers7 minRead the analysis - Product metadata · 35M items
What’s on the shelf
Prices, brand concentration, and the attribute vocabulary of 35M products — plus two measures the corpus simply cannot answer.
pricesconcentrationmetadata7 minRead the analysis - Full aggregate — 507.7M reviews
When people write reviews
Month, weekday, and hour across all 33 categories — and the December-buys / January-receives pattern hiding in the gift categories.
seasonalitygiftingcross-category8 minRead the analysis - Full aggregate — 507.7M reviews
Twenty-eight years, and 3.7% of them
Volume and mean rating by year, per category. Why the long history contributes almost nothing to a pooled average, and what the 2013 jump and the 2021 peak actually were.
time seriescompositionratings6 minRead the analysis
The data
Thirty-four CSVs, no text, no identifiers
The published aggregates are counts, means, and distributions only — no review text, no user ID, no product ID. That is what makes them safe to hand out and what makes them useless for per-product or NLP work; for that you need the HuggingFace source.
The category summaries — five files
| File | Rows | What it holds |
|---|---|---|
| category_stats_all.csv | 33 | One row per category — volume, mean rating, mean length, verified share, the 1★–5★ split, first and last review date. |
| ts_yearly_all.csv | 798 | Category × year, 1996–2023. The only chronological file. |
| ts_monthly_all.csv | 396 | Category × calendar month. Seasonality, all years pooled. |
| ts_dayofweek_all.csv | 231 | Category × weekday (0 = Monday), all years pooled. |
| ts_hourofday_all.csv | 792 | Category × hour (0–23), all years pooled. |
The behavioural aggregates — 29 more files
The five files above summarise the review table one category at a time. A second set, merged_results_v2/ — 29 CSVs, 13.7 MB — instead groups every review by its author and by its product, which is what makes reviewer-level and item-level questions answerable at all. It is where every finding in the reviewer analysis comes from. Read merged_results_v2/_meta/manifest_v2.json first — it documents every file’s grain, suppression, and caveats, and records two measures it deliberately did not publish.
Grain is load-bearing, not decoration: C is one row per category, G is corpus-wide with no category column, X is a category × category matrix that must not be joined to per-category files, and C+G carries both, separated by a scope column. The two scopes answer different questions and are not comparable.
Plain HTTPS — no credentials
import pandas as pd
BASE = "https://ontopic-public-data.t3.storage.dev/amazon-reviews/merged_results/"
V2 = BASE.replace("merged_results", "merged_results_v2")
cats = pd.read_csv(BASE + "category_stats_all.csv")
oad = pd.read_csv(V2 + "user_one_and_done.csv")S3 protocol
import boto3, pandas as pd
s3 = boto3.client("s3", endpoint_url="https://t3.storage.dev",
region_name="auto")
obj = s3.get_object(Bucket="ontopic-public-data",
Key="amazon-reviews/merged_results/"
"category_stats_all.csv")
cats = pd.read_csv(obj["Body"])The bucket answers anonymous GETs on virtual-host style URLs (bucket.t3.storage.dev/key); the path-style form t3.storage.dev/bucket/key returns 403.
Six ways to get this wrong
Only ts_yearly is a timeline. The monthly, weekday, and hour files pool every year together. Plotting them left to right as a time axis produces a chart that means nothing.
Filter on count before trusting a rate. A category-year holding one review reports rating_5_pct = 100.0. Every rate chart here drops cells under 500 reviews.
Volumes span four orders of magnitude. 67.4M reviews in Home & Kitchen against 16,216 in Subscription Boxes. Normalise before you compare.
Percent columns are 0–100. Not 0–1. Dividing twice, or not at all, is the most common bug against these files.
Every mean rating is an upper bound. The Unknown category is excluded — 11% of ratings but 42% of users, and disproportionately one-and-done. Since sparse reviewers rate lower (3.81★ at one review, 4.26★ at ten or more), dropping them pushes every published mean up.
The variance shares do not decompose. user 30.2%, item 19.3%, category 0.7% are MARGINAL shares of one factor at a time. They are crossed and unbalanced, sum to 50.2% corpus-wide and 105.6% on Gift Cards, and partition nothing. Do not derive a residual from them.
Derived from McAuley-Lab/Amazon-Reviews-2023 and inherits its terms. Aggregation ran on Google Cloud Run, one job per category, streaming each raw_review_* split; the merged CSVs were migrated to Tigris in August 2026 with every object verified by MD5. Charts on these pages read a 33-category JSON built from those CSVs by scripts/fetch-amazon-aggregates.mjs.