Here’s the uncomfortable part of owning a computer that nobody prints on the box: your storage drive is a consumable. It has a lifespan, it roughly knows how much of that lifespan is left, and in most cases it starts complaining about the end long before it actually gets there. The catch is that the complaining happens in a log full of numbers that most people never open.
So they don’t open it. Then one day the machine won’t boot, the files are gone, and everyone acts surprised. Drive failures only feel sudden because the warning signs are hidden by default. They’re not secret. They’re just boring, and boring is why nobody looks.
This is how to actually check drive health, what the numbers mean, and what you can realistically do once you’ve read them.
First, kill the myth of the sudden death
Drives almost never die in one clean catastrophic moment. They degrade. A spinning drive develops weak spots on the platters and starts quietly remapping them. A solid-state drive burns through its write cycles and retires cells one block at a time. Both processes take months, sometimes years, and both are tracked internally from the day the drive leaves the factory.
The failure mode people experience isn’t the drive dying. It’s the drive dying while they were ignoring it.
The health log your drive has been keeping the whole time
Nearly every modern drive — mechanical or solid-state — continuously tracks its own vitals: temperature, hours powered on, power cycles, error counts, and a set of predictive counters designed to flag degradation before total failure. This data is exposed through a standard diagnostic interface that any operating system can read. The industry calls it SMART; you can think of it as the drive’s own medical chart.
Reading it takes seconds. A free utility that pulls the drive’s internal diagnostic log will show you the whole thing in a plain list of attributes, each with a raw value and a threshold.
The counters that actually matter
- Reallocated sectors — spots the drive found bad and quietly swapped for spares. A handful from the factory is normal. A number that grows between checks means surface damage is spreading.
- Pending sectors — spots the drive is unsure about and hasn’t remapped yet. These get re-tested on the next write. If they stay pending instead of clearing, you have a problem.
- Uncorrectable errors — reads that failed and couldn’t be recovered. Even one or two here deserves attention.
- Spin-up retries and mechanical errors — spinning drives only. The motor struggling to get the platters going is a classic end-of-life signal.
- Wear leveling and remaining life — solid-state drives only. This is your remaining write endurance as a percentage. Below roughly eighty percent, start planning a replacement.
- Temperature — sustained heat above the drive’s rated range accelerates everything above. It’s also the one variable you can actually control.
The counters that are mostly noise
- Power-on hours — a long-running drive is not a failing drive.
- Raw read error rate — on many drives this reports wildly alarming numbers from day one and means nothing. It’s a known reporting quirk, not a death sentence.
- Load cycle count — only worth worrying about if it’s absurdly high relative to the drive’s rated limit, which usually means something is aggressively parking the heads over and over.
The rule that cuts through all of it: trends matter more than absolute values. One bad counter on a five-year-old drive is fine. The same counter climbing every month is a countdown.
The five-minute check anyone can run
- Install a utility that reads the drive’s internal diagnostic log. Any of the free ones will do; you don’t need enterprise software for this.
- Look at the overall health verdict first. If it says the drive has failed or is failing, stop and go straight to backing up.
- Scan the counters listed above. Note anything flagged or close to its threshold.
- Write the numbers down. This is the step everyone skips and it’s the only one that makes the rest useful.
- Repeat every month or two. You’re looking for movement, not perfection.
The deeper checks that actually stress the drive
A surface scan — read-only first
A read-only surface scan walks every sector and reports which ones can’t be read. It’s slow, it takes hours on a large drive, and it’s worth it because it finds weak spots the diagnostic log hasn’t catalogued yet. There are destructive versions that write over every sector to force remapping — those erase the drive, so only run one on a drive whose data you’ve already copied somewhere else.
The practical trick: run the read-only scan first. If it comes back clean, you probably don’t need the destructive one at all.
A filesystem consistency check
The storage hardware and the filesystem on top of it are two different layers, and they fail differently. A drive can be physically perfect while its filesystem is a mess of broken links and orphaned fragments. Every major operating system ships a built-in check utility for this. Run it from a maintenance or recovery mode rather than a live session if you can — you’ll get far fewer false positives and it can actually repair what it finds.
The boring stuff that breaks drives
Before you declare a drive dead, eliminate the idiots:
- Cables. A marginal data cable produces exactly the symptoms of a dying drive: timeouts, corrupted reads, errors in the log. Reseat it, then replace it. They cost almost nothing.
- Power. An aging supply or a shared power rail that can’t hold voltage under load causes random dropouts that look like hardware failure.
- Heat. A drive packed between two others with no airflow runs hotter than it should and ages faster.
- Enclosures and adapters. Cheap USB bridges frequently mangle diagnostics and cause errors that the drive itself never produced. Test the drive connected directly before you believe anything an adapter tells you.
What’s fixable and what isn’t
Usually fixable
- Loose or failing cables and connectors
- Overheating, solved with airflow or spacing
- Filesystem corruption and partition table damage
- Power delivery problems
- Firmware bugs — sometimes a vendor update genuinely fixes recurring dropouts
Not fixable, no matter what a forum tells you
- Growing reallocated or pending sectors. The platter surface is physically degrading. Software cannot undo that.
- Mechanical noise — clicking, grinding, beeping. That’s a head crash in progress. Power it down.
- Solid-state wear-out. The cells are used up. There’s no reset.
There’s a whole cottage industry built around reviving dead drives with freezer tricks and firmware surgery. It sometimes works, briefly, for people who need one last copy of their data. It is not a repair. It’s a hostage negotiation.
When to stop diagnosing and start moving data
Here’s the rule that will save you more data than any tool: the moment you run a second diagnostic on the same drive, you’re already in backup territory. The first one tells you something’s off. The second one is you hoping it isn’t.
So copy everything that matters immediately, verify the copies open, and then keep poking at the drive if you’re curious. A drive that’s already given you one warning is not a drive you want to keep your only copy of anything on.
The real fix is not caring
All of this monitoring is damage control, not protection. The honest truth is that drives are disposable and your data isn’t. Keep at least two copies on separate physical devices, keep one of those somewhere that isn’t in the same building, and check that the copies actually restore instead of just assuming they do.
Do that, and a dying drive stops being a disaster and becomes an errand. You read the numbers, you shrug, you swap the hardware, and nothing is lost. That’s the whole point of the health check — not to save the drive, but to make sure its death is merely inconvenient.
So go look at the log. It’s been sitting there waiting for you this entire time.