20 ms·
HPE Drive fail at 32,768 hours without firmware update
- voiper1 7y ago>The issue affects SSDs with an HPE firmware version prior to HPD8 that results in SSD failure at 32,768 hours of operation (i.e., 3 years, 270 days 8 hours). After the SSD failure occurs, neither the SSD nor the data can be recovered. In addition, SSDs which were put into service at the same time will likely fail nearly simultaneously. Looks like some sort of run time stored in a signed 2 byte integer. Oops.
- cesarb 7y agoIt's probably the SMART "hours of operation" field. I see no reason for anything else to be stored in units of hours instead of seconds. Yes, this means that a field meant for diagnosing failures was responsible for a failure. Oops.
- stefan_ 7y agoBut how does that brick the device? I guess the hour counter overflows, goes negative and that screws up a calculation later on, causing the firmware to crash (over and over again..) If only SSD vendors would do the usual cost-cutting measure of loading firmware from the host computer, this could be trivially fixed.
- mnw21cam 7y agoPlease no. I may actually want to boot from one of those devices.
- rini17 7y agoSomewhere in the UEFI kitchensink there certainly is firmware loading support already.
- lonelappde 7y agoWhat do you mean? The fix is to load a firmware update from the host computer.
- stefan_ 7y agoNot once you are past 32768 hours, apparently.
- imtringued 7y agoUsually the SMART counter just wraps around back to 0. In this case it becomes negative because it was read as a signed short.
- zozbot234 7y agoOuch. I wonder how many non-enterprise SSD's come with similar bugs, and zero support by the firmware vendor.
- close04 7y ago> neither the SSD nor the data can be recovered It looks like such a bug isn't necessarily SSD specific if it completely bricks the drive. And while 32.768 hours may seem like a long time for a drive, it's under 4 years of continuous operation. Not unheard of if used in a NAS.
- junglecat 7y agoI'm guessing it's not a particularly productive way to store timestamps.
- zozbot234 7y ago> It looks like such a bug isn't necessarily SSD specific if it completely bricks the drive. Maybe. OTOH, plenty of people have been running spinning rust drives with way more than 4 years of power-on operation - if this bricking bug was common there, I'm pretty sure we would've noticed. SSD's are a newer tech and it's more common to replace them anyway as specs improve.
- Symbiote 7y agoI would think the typical HPE customer (e.g. us) buys servers and uses them for between 3-5 years, before buying new servers. The old (out of warranty) servers might be discarded, or might be reused as test hardware.
- tyingq 7y agoNon tech Fortune 500 IT shops regularly see their refresh budget cut in favor of new projects. Seeing some amount of 5,7,10+ year old hardware still in service isn't unusual.
- 7y ago
- abarringer 7y agoSince most drives are started and used concurrently this bug would blow any RAID set up. There's a dark day coming for some sysadmins.
- cesarb 7y agoThat's only if the sysadmin was trusting a single server with the data, instead of a pair of redundant servers. Which were probably installed and started up at nearly the same time. Oops. This bug has the potential of simultaneously damaging whole sets of servers, if they were bought and installed in bulk. Dark day indeed.
- abarringer 7y agoWe have a cluster of four nodes that were all setup and brought online within hours of each other. The entire cluster could blow up within a couple hours if not patched.
- Scoundreller 7y agoI guess cluster nodes should be scheduled to be taken down for random amounts of time so that they fail in sequence more gracefully. Have an “off on weekends” node. And 24/7 nodes.
- vidarh 7y agoThis is why I never use drives from the same batch, ideally never the same model, and usually not the same manufacturer. It happens way too regularly that drives start failing around the same time.
- GrayShade 7y ago> I never use drives from the same batch Note that it wouldnțt help in this instance, as the bug is caused by the amount of time a drive was running. Different manufacturers would work, yes.
- verytrivial 7y ago> HPE was notified by a Solid State Drive (SSD) manufacturer [...] That's a curious bit of context. It seems to imply they're shifting some of the blame onto their manufacturer? I makes me wonder if this firmware is 100% HPE specific, or if there a 2^16 hours bug about to bite a bunch of other pipelines.
- wtallis 7y agoThis is a 2^15 hours bug, not a 2^16 hours bug. Odds are, somewhere in the SSD firmware source code, there's a missing "unsigned". And in the Makefile, probably a missing "-Wextra".
- richthegeek 7y agoPresumably the bug would still be considered a bug if it occured at 65536 hours? The incorrectly-signed bit only makes it appear sooner, but it's not the bug.
- pfortuny 7y agoSomeone says below that rollover may work OK but not negative times, for whatever reasons.
- air7 7y agoOh! So in that case owners just need to wait 32,768 hours and they'll be fine. /s
- moonbug 7y agoThat's 7.4 years, and well outside warranty.
- eb0la 7y agoIf you purchase a Carepaq, can actually extend the service beyond that? 3,4,and 5-Year 24x7 Carepaqs are available on purchase for the _whole_ chassis and parts... but I wonder if I can keep purchasing 1-year carepaq extensions beyond the 6th year.
- userbinator 7y agoAccording to this page, the SMART hour counter is only 16 bits, and rollover should be harmless: http://www.stbsuite.com/support/virtual-training-center/power-on-hours-rollover http://www.stbsuite.com/support/virtual-training-center/powe... If you look elsewhere on the Internet, you'll find people with very old and working HDDs that have rolled over, so I suspect this bug is limited to a small number of drives. (What that page says about not being able to reset it is... not true.) Likewise, I'm skeptical of "neither the SSD nor the data can be recovered" --- they just want you to buy a new one. Tangentially related, I wonder how many modern cars will stop working once the odometer rolls over.
- garaetjjte 7y ago>Likewise, I'm skeptical of "neither the SSD nor the data can be recovered" If the firmware crashes during boot with negative hour counter, it probably could be only fixed by manually flashing new firmware over JTAG.
- userbinator 7y ago...and likely some of the data recovery companies already know about and are prepared for this.
- djsumdog 7y ago$500 ~ $1k for a JTAG flash. I'm sure they have plenty of other drives that come in and they price at $1k but that take days longer than expected, so it probably all balances itself out eventually.
- Scoundreller 7y agoOr they swap boards and you eventually end up with the same problem...
- simcop2387 7y ago
- jzwinck 7y agoThose who forget history are doomed to repeat it. Just seven years ago Crucial sold tens of thousands of their "M4" SSDs with a firmware bug that made them fail after 5184 hours: https://www.anandtech.com/show/5424/crucial-provides-a-firmware-update-for-m4-to-fix-the-bsod-issue https://www.anandtech.com/show/5424/crucial-provides-a-firmw... Do they still not test these things with artificially incremented counters?
- djsmiley2k 7y agoSome early Intel SSD's did the same thing, prior to M4's... haha
- jzwinck 7y agoDo you have a link for the old Intel bug? Here's one for a new Intel bug after just 1700 power-on hours on some enterprise-class SSDs that are still being sold today: https://www.intel.com/content/www/us/en/support/articles/000038720/memory-and-storage/data-center-ssds.html https://www.intel.com/content/www/us/en/support/articles/000... That's just 71 days of uptime and they hang. There are tens of thousands of these drives deployed as well.
- Bootwizard 7y agoWhy are they allowed to still sell these broken products? That's a scam as far as I'm concerned.
- oarsinsync 7y agoThey released a fix for existing devices, replaced affected devices that had been bricked, and included the fix as part of the manufacturing process for new devices being built. It was certainly very inconvenient having to reboot systems while we waited for a fix to exist, fortunately we didn't lose too many disks before the fault was identified, which reduced the man hours involved in the DCs. I'm not sure if any part of this is a scam. A bug, certainly.
- 7y ago
- pabs3 7y agoWould be nice if the standard firmware update mechanism on Linux (fwupd/LVFS) could be used for HPE products. https://fwupd.org/lvfs/vendors/ https://fwupd.org/lvfs/vendors/ https://fwupd.org/lvfs/devices/ https://fwupd.org/lvfs/devices/
- jabl 7y agoThis so much. Even if you hate uefi with the fire of a thousand suns, there are some good things there. Like GPT, and the UEFI capsule thing that fwupd uses.
- HorstG 7y agoHP actually has a working firmware update mechanism for all their gear. Its a bootable Linux liveDVD that starts into a browser talking to a local Tomcat instance which applies necessary patches. For many cases its also possible to invoke patching from your normal Linux installation. However, a reboot is mostly still necessary, e.g. for disk firmware which the controller applies after its own new firmware has been loaded (sometimes takes more than one reboot). The system is quite a lot older than fwupd and less flakey usually. Google for hpsum or HP SPP
- jabl 7y agoYes, hpsum / SPP is what we use now. Not happy about it.
- pjc50 7y agoAmazing. A repeat of the "Windows 95 crashes after 48 days uptime" and other timer rollover bugs.
- macintux 7y agoI’ve always appreciated the humor of the fact that Win95 was so unstable that no one noticed this bug until years later.
- kube-system 7y agoIt also used to be very common for people to turn off their computers when they were done using them.
- Piskvorrr 7y agoYup. Those were power-hungry times, without deep sleep modes, or even useful hibernation. Heck, most machines then didn't even bother to reduce clock speed when idle.
- buckminster 7y agoThe CPUs used hardly any power so there wasn't much point throttling them. The rest of the computer and the CRT used loads though.
- olyjohn 7y ago"It is now safe to turn off this computer" I remember when I saw my first ATX computer, and it turned itself off. That was cool.
- tyingq 7y agoLinux wasn't so great in 1995 either. We regularly rebooted for various kernel, ip stack, etc, bugs that would crop up after a fairly short amount of uptime. Our Sun workstations were very stable though.
- 7y ago
- bobowzki 7y agoAt the hospital where I work, almost all HP desktops crashed within a few months...
- luma 7y agoThis is for HPE (not HP, which is now a separate company). I haven't heard anything about HP (who makes desktops not storage arrays) experiencing this problem.
- cotillion 7y agoConsidering 900 drives out of 1800 crashed in HP computers at a Swedish hospital in the last few months I suspect there is a connection. Maybe HP and HPE were more tightly connected 3 years, 270 days 8 hours ago.
- strictnein 7y agoThey sure that's hardware related, and not some poorly written ransomware?
- luma 7y agoAll of the impacted devices in this issue are SAS-connected enterprise-class SSDs. While it's not _impossible_ that someone installed a SAS controller and used SSDs that can cost upwards of 10x the price of an equivalent-capacity desktop model using SATA... it's probably pretty unlikely.
- S_A_P 7y agoI just want to know how many of these failed at 32768 hours before they had their oh sh*t moment.
- retrovm 7y agoBeats me but I happen to have a fleet of HP SATA (not SAS) drives and they just crossed this boundary, 32813 power on hours typical. I guess if their SATA firmware had this bug I'd be having a bad week.
- paggle 7y agoYikes! This is why when I built my home NAS I used five different drives and manufacturers.
- EvanAnderson 7y agoI did some recon on eBay looking for used units w/ the affected SKUs for sale and they appear to be Samsung units.
- _bxg1 7y agoWhatever the counter is, the fact that it's 32,768 instead of 65,536 suggests they used a signed int for something that presumably starts at zero and increases monotonically... Avoiding just that mistake would've given them twice as much time - nearly 7.5 years - which seems like it'd be longer than these drives would typically last anyway.
- vortico 7y agoMaybe they're running on a 15-bit architecture where a signed int would be 16384? /s
- imtringued 7y agoIt would have avoided the problem in the first place because the SMART counter is allowed to roll over back to 0.
- gruez 7y ago>By disregarding this notification and not performing the recommended resolution, the customer accepts the risk of incurring future related errors. How is this work legally? For one, how would HPE prove that the customer read the bulletin? I don't imagine they're sending these out via certified mail.
- annoyingnoob 7y agoWhew, dodged that bullet, looks like I'm not using any of the affected drives. Lucky me, for now.
- iveqy 7y agoProbably related to https://news.ycombinator.com/item?id=21471997 https://news.ycombinator.com/item?id=21471997