Skip to content

Comment on HPE Drive fail at 32,768 hours without firmware update

Comments

Ouch. I wonder how many non-enterprise SSD's come with similar bugs, and zero support by the firmware vendor.

neither the SSD nor the data can be recovered

It looks like such a bug isn't necessarily SSD specific if it completely bricks the drive.

And while 32.768 hours may seem like a long time for a drive, it's under 4 years of continuous operation. Not unheard of if used in a NAS.

It looks like such a bug isn't necessarily SSD specific if it completely bricks the drive.

Maybe. OTOH, plenty of people have been running spinning rust drives with way more than 4 years of power-on operation - if this bricking bug was common there, I'm pretty sure we would've noticed. SSD's are a newer tech and it's more common to replace them anyway as specs improve.

There was actually a similar drive-bricking bug in the SMART implementation of some of Seagate's hard drives about a decade ago. Not quite as deterministic as this one, but essentially uptime-dependant. After denying it for a while, they finally fessed up once a member of the public figured out that affected drives could be unbricked by connecting over the serial debug interface and wiping the SMART log. They ended up offering firmware updates and free unbricking of affected drives at their cost including shipping to their facility. (I don't think the unbricking process was terribly easy either - the publicly-known version required booting the drive with the heads disconnected from the controller to stop it from reading the data, or it'd get stuck in an infinite loop and not respond on serial.)

I have one of these drives. Apparently a model that the known procedure doesn't work on. I'll never buy a shitty Seagate product again.

I would think the typical HPE customer (e.g. us) buys servers and uses them for between 3-5 years, before buying new servers. The old (out of warranty) servers might be discarded, or might be reused as test hardware.

Non tech Fortune 500 IT shops regularly see their refresh budget cut in favor of new projects. Seeing some amount of 5,7,10+ year old hardware still in service isn't unusual.

Oh I'm not implying it's a common bug. I was just saying that there's nothing in the description that makes it look SSD exclusive. To completely brick the drive sounds more like a controller failure. So such a bug could just as easily kill an HDD controller.

SSDs have insanely complex firmware when compared to HDDs so letting this kind of bug slip through is an easier mistake to make in their case.

It's the drive firmware. Drive firmware bricks the disk because the disk is soldered into the drive.

It doesn't need to be continuous. Total operation is the metric.

I read OP's comment as "if this bug happened to a non-enterprise (regular) user and their drive". A regular user has a lower chance of hitting 4 years of non-continuous operation before discarding the drive due to obsolescence. On a 50% duty cycle (12h/day every day) it would take 6.5 years. That's close to how long many people would hang on to their drives. But even a regular user might have a NAS and that accelerates the process.

Outside of the deep-pocket money-doesn't-matter sized enterprises, 4 years can be less than half the expected lifetime for IT kit.

I'm guessing it's not a particularly productive way to store timestamps.

That's a very coincidental reason to go for a 3 year warranty.

Unfortunately, even if you operate these drives continuously from the day you buy them, they will take 3 years, 270 days and 8 hours to fail (as someone else kindly calculated), so a 3-year warranty won't help you in this case...

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.