How we compile these figures
This page describes what happens to a drive’s listing between being scraped and appearing on a comparison. It includes the parts that go wrong.
We do not review drives
Nothing here is a review in the usual sense. We do not buy, test or handle the drives. This page describes how we handle records — what a manufacturer publishes about a drive — not how we handle the drive.
Building the corpus
Listings are gathered across the searches that define the category, then filtered to the products that actually belong. Search results are full of things that are not drives: enclosures, adapters, cables, docking stations and rack hardware, many of which carry a capacity figure in their own title and would otherwise be read as drives.
Every drive is assigned to a class from its own title and specification table, never from the search that surfaced it. That distinction is not academic — searches for one kind of drive return the other kind in quantity, and a class built from search terms is wrong in both directions.
What this category’s listings get wrong
- A family listing names every form factor it ships in, so a 2.5-inch SATA drive whose title also mentions M.2 can be read as either. Where the specification table and the title disagree, the table wins, because it describes the item actually being sold.
- A drive advertised as an “HDD replacement” is an SSD; the phrase names what it replaces. Listings that state a spindle speed are never SSDs, and that test settles the ambiguous cases outright.
- Capacities are stated in binary and decimal interchangeably. Bands here follow the binary figures the drives actually report, so a 4 TB drive is placed by its 4,096 GB rather than a rounded 4,000.
Baselines
A baseline is the median of a specification within one class and one capacity band, with values beyond 1.5× the interquartile range excluded so a single outlier cannot move it. Baselines are published only where the group holds at least 8 drives; below that a median describes the sample rather than the market, and the page says no baseline can be published rather than printing a number that looks authoritative and is not.
A specification stated by fewer than half a group’s drives is withheld from that group’s baseline for the same reason.
Prices
We publish price tiers, never exact amounts: a price captured on a scrape date is wrong by the time you read it. The one derived figure we do print is price per terabyte, because it is the only number that makes drives of different capacities comparable and no listing states it — and it is shown against a class median rather than as a claim about what anything costs today.
Where this fails
- A published figure can be optimistic, or measured under conditions the listing does not give. We can check a figure against its own class; we cannot check it against a drive.
- Automated cleaning catches the defects it has rules for, and every rule here exists because something got through first. The next class of defect is one we have not seen yet.
- Coverage is uneven. Some fields are stated by nearly every listing and some by a third, and a field stated by few drives supports a much weaker conclusion.