13 ms·
Backblaze Hard Drive Stats Q2 2020
- sandworm101 6y agoI win. The 8TB drives I used to populate my NAS scored 0.10% better than the 0.81% average. That means my NAS is better than everyone else's. This big takeaway from these numbers is how dramatically low they are in comparison to drives 5/10/25 years ago. If you treat them reasonably, modern drives are rock stable.
- RealStickman_ 6y agoYou could have done 0.7% better :)
- pnutjam 6y agoThank you. I've noticed you can distill most news, religion, and investments down to a simple check, "how does it make me look better then everyone else." Why yes, I'm an American, why do you ask? ;)
- ghaff 6y agoIn the last 5-10 years, I've definitely retired more drives because they became too small for what I was using them for than because they failed.
- jhardy54 6y agoHave you considered using them for RAID? https://en.wikipedia.org/wiki/RAID https://en.wikipedia.org/wiki/RAID
- Xylakant 6y agoYou’d still need the slots and some raid modes are limited by the size of the smallest disk. Drives being as cheap as they are, it’s often not economically viable to retain the old drives.
- zaarn 6y agoIf you use software-raid you can use striping. Basically, divide up each disk into 1TB partitions, then RAID the 1TB partitions together. Synology uses this for their Hybrid RAID. Obviously care needs to be taken to not confuse partitions but this lets you use disks of differing sizes. Unraid also offers RAID with differing sizes for disks, it has quite different failure modes though but it's been fairly reliable for me.
- ghaff 6y agoI know what RAID is :-) Not really, disks are cheap enough that I'd rather keep things simple. I have redundancy in various ways in addition to using Backblaze.
- jandrese 6y agoAffordable RAID controllers end up being more failure prone than the drives attached to them.
- viraptor 6y agoYou don't need hardware raid controllers. (Unless your really need them for some reason) Software RAID is perfectly fine for most use cases.
- akx 6y agoAn mdadm is fine too.
- Yetanfou 6y agoAffordable disk shelves - something like the Netapp DS4243 - on the other hand tend to be reliable and give plenty of space for SAS and SATA drives. I bought a fully populated shelf (4 x powersupply (of which I'm only using 2, the rest are disconnected and meant as spares), 2 controllers and 24 15K SAS drives (mostly HGST, a few Seagate drives) for around €400. I'm currently running it with a number of those 15K drives but will eventually replace some of them with slower but bigger SAS or SATA drives, not so much to get more space but to save power. For now the thing runs fine, the error logs on the drives are mostly clean (2 drives with a few errors, mostly likely due to power glitches as they occurred a long time ago and have not recurred). If you want lots of space for drives with redundant power and cooling these shelves are a good value since there are plenty of them on the market which keeps prices down. I'm using a combination of striping, RAID and mirroring depending on the content, using lvm and mdadm instead of ZFS for the increased flexibility. The server itself contains a HP R410i RAID card which drives 8 internal 2.5" SAS drives, the thing could be expanded for not that much money to a 16 x 2.5" array but I see no reason to do so given the higher prices for 2.5" drives and the abundance of capacity offered by the DS4243.
- Marsymars 6y agoPorts and physical space for drives aren't free. (Though there's some trade-off available between time/money if you want to DIY.) For my own use, I've found it isn't worthwhile expanding beyond eight drives rather than getting rid of smaller ones - adding additional ports costs money and it takes more power and space.
- programmertote 6y agoTangential question: What NAS device and set up do you use? I'm thinking of buying one and would like to learn from a fellow (technically better acquainted) HN user. Thanks in advance!
- sandworm101 6y agoSynology DS1019+ It might not be the cheapest, but 4+1 drives is more efficient than 3+1. Currently I have three 8TB drives: 1xWD Red and 2xIronwolfs (wolves?) with one running as parity (SHR). No complaints. It does everything I need it to do. The interface is simple and everything "just works" as anticipated. If/when I run out of space I plan on adding another 2x8TB drives, probably more ironwolves. I wouldn't go bigger than 8TB as they already take DAYS to add to the array.
- sumtechguy 6y agoI have that same box. Very nice box. Migrated from a 1512. It took probably about 3-4 days to migrate everything. The new box handled it easily. The old one was struggling a bit with rsync. As the default is ssh with it. Once I had to downgraded the encryption then it went a lot faster (from 30MB to 90MB which was close enough to the limit of the cable to not worry).
- lostlogin 6y ago918+ here. The last two drives I put in were 16TBs, and as you say it takes time. It wasn’t helped by me messing the process first round (you need to add the smallest drives first with SHR), so it took over a week.
- programmertote 6y agoThank you for the recommendation! Definitely will track its price now. :)
- xur17 6y agoNot the OP, but I'll answer. I have the Synology DS-918 for the past 2 years, and it's been great. They have fairly decent included management software + an application library. I have it setup to automatically backup to Backblaze every night, which makes me fairly confident anything I store on it will be recoverable.
- ramraj07 6y agoOne of my main questions is how different these stats are for hard drives that are not running all the time. Personal experience over a decade in an underfunded lab was that if you took a drive offline and left it for couple years, you had 10-40% chance of it failing.
- leetcrew 6y agoI had a couple drives that no longer worked after leaving the computer dormant for a few months. most of my drive failures have occurred immediately after hardware upgrades though. seems like I lose at least one drive every time I replace a mobo. I'm always surprised that breaks them, but they always survived trips to/from college in the trunk of my car.
- londons_explore 6y agoI'd give those drives a reformat... Some motherboards have firmware that does odd things to drives, like hiding the first or last sectors to store its own data or for various OEM recovery techniques. That can make your data look corrupt. Yet reformatting (or just putting the drive in with the old motherboard) will magically fix them.
- java-man 6y agoInteresting how HGST has consistently low failure rates [0]. What makes these drives so reliable? Japanese fixation on quality? Specifics of Hitachi design? [0] https://www.backblaze.com/blog/wp-content/uploads/2020/08/Chart-Q2-2020-MFR-AFR-1024x679.jpg https://www.backblaze.com/blog/wp-content/uploads/2020/08/Ch...
- Hamuko 6y agoIsn't HGST just Western Digital now?
- misanthropian00 6y agoOnly if Western Digital was dumb enough to kill the golden egg laying goose. If I were they I would try to keep the Hitachi manufacturing just as it was when they bought it. They do seem to still brand them differently than their own drives and charge a hefty premium and market them to data centers rather than individuals. For me they are way too expensive now. So for me the only Japanese option is Toshiba. I have had much better luck with Toshiba than with Seagate and would easily pay 25% or more of a premium to avoid Seagate. It is so sad because I remember when Seagate was a great and very reliable brand so many years ago.
- cehrlich 6y agoInteresting to note that this wasn't always the case - The Hitachi Deskstar line used to have the nickname "Deathstar" due to their low reliability (this was maybe 15-20 years ago)
- jaclaz 6y agoThis deserves to be cited: HGST drives win the prize for predictability. Boring, yes, but a good kind of boring.
- codezero 6y agoAlso worth calling out that the worst performer, Western Digital, acquired HGST, now let's see if WD improves or HGST gets worse :) Edit: See comment below, looks like this happened a while ago, so not likely a real factor :)
- dv_dt 6y agoWD acquired HGST in 2012, but was strictly independent until 2015. https://en.wikipedia.org/wiki/HGST https://en.wikipedia.org/wiki/HGST
- codezero 6y agoThanks a bunch for the correction, I guess the answer is that potentially the acquisition has helped WD drives a little bit :)
- dv_dt 6y agoFWIW I wonder if they're still run as something of a separate group, but given that the article says the HGST brand was removed in 2018, who knows what speed the long term merge might occur upon. Explains why HGST brand drives are harder to come by I think.
- codezero 6y agoYeah, this wouldn’t surprise me, and it’s not super uncommon. Similar things happened with Cisco and Meraki, though I’ve heard over time that’s started to erode.
- cactus2093 6y agoThe problem is, as far as I can tell you can't really buy them anymore. At least from consumer sites, maybe they still sell large orders to enterprise customers. But it seems like WD is finally discontinuing the HGST name.
- ebg13 6y agoIt's interesting that, despite Seagate being relatively bad compared to HGST, Backblaze keeps installing more and more Seagate drives and not that many HGST drives. I guess the analysis that's missing here is the cost per drive hour?
- atYevP 6y agoYev from Backblaze here -> Yes, we tend to weigh heavily in favor of price and Seagate thus far as been affordable with performance that meets our requirements, so that's why you'll see those drives in higher numbers. We're also trying to diversify our fleet a bit, so it's not purely one vendor, but we do try to keep an eye on cost since we run a pretty lean ship!
- AnIdiotOnTheNet 6y agoYeah, pretty much the same reason people keep buying Seagate drives generally: They cost less and they aren't so much worse that it is worth paying a higher price for a drive that's less likely to fail.
- kube-system 6y agoIt sounds like cost is the primary factor, but diversity can also help mitigate the risk of unknown bugs. Lets say you found that HPE SSDs were super reliable over your test period so you decided to put everything on a bunch of new drives. Then this hits, and 100% of your drives fail at the same time: https://www.pcmag.com/news/time-to-patch-hpe-ssds-will-fail-after-32768-hours https://www.pcmag.com/news/time-to-patch-hpe-ssds-will-fail-... While these stats are a great analysis of random failures -- that's only one type of risk.
- hinkley 6y agoI have seen homogeneity 'kill' twice in my career. The first time was unprovisioned hardware, then second time we lost half of the hard drives assigned for dev machines in the space of about 9 weeks. I had so many processes in place by that point that it was a huge inconvenience (mostly due to whole disk encryption) but zero data loss. The worst thing you can do is to put all of your eggs into one basket. And bulk ordering might get you a set of drives all from the same batch. Bad batches will tend to all fail for a similar reason. Drives fail on a bell curve, right? So the first drive may fail way before any others, but in a RAID array, rebuilds are stressful. Eventually you will hit a statistical cluster. Multiple drives failing close together. If that happens during a rebuild, you will lose a RAID 5 array. If you are very lucky, your RAID 10 array loses two drives in the same mirror. If you have two failures during a long rebuild, even RAID 6 won't save you. I just bought a Synology box for home. This is my third and probably final RAID enclosure for personal use. I was having trouble finding Backblaze-tested drive models to populate it, so I filled it with drives from a Drobo and kept looking. Initially I had populated the Drobo with 4 drives I bought at once. When one failed, I bought 2 HGST drives and replaced a pair. When the new drives arrived, I started trying to cycle them through, and one of the drobo drives failed. I'll give you two guesses which one. There is, as far as I can tell, no prosumer multi-disk filesystem that uses consistent hashing to stripe+mirror files across an arbitrary number of disks, instead of the heavy linear algebra RAID5 relies on. It requires touching the whole disk on every rebuild, and I believe that's why Object Storage is slowly taking over from the top end. It's a simpler form of redundancy. I hope that it's worked its way down to my price range by the time the motherboard on the Synology burns out.
- pwinnski 6y agoI never get tired of these reports, and if I ever again buy a drive with a high failure rate for Backblaze, somebody slap me!
- system2 6y agoWhy wouldn't backblaze use only HGST then? What's the purpose of buying seagates which was the most unreliable hard drives from 15-20 years ago and their stats show higher failure rates. (Still I don't buy them because of the bad taste left.)
- pwinnski 6y agoPresumably cost and availability. I remember reading at one point they were driving from store to store buying every drive they could find.
- hinkley 6y agoThat was during the flooding in Thailand that took out a huge chunk of hard drive manufacturing capacity. Electronics stores started putting quotas on purchases, so BB offered a bounty to friends and family who would bring in hard drives. They had a disassembly line gutting external hard drive enclosures for the hard drives, too. One of the drives in my array was acquired via a similar process. You should read the whole story. I should probably read it again myself. https://www.backblaze.com/blog/backblaze_drive_farming/ https://www.backblaze.com/blog/backblaze_drive_farming/
- xvolter 6y agoI think this comes down to price. They are constantly buying drives and to avoid supply issues, it is acceptable to have a hard drive that fails to a certain degree, if the price is balanced. It would be interesting if these reports included prices, but that might be a problem for Backblaze to reveal so much about their business operational costs.
- uj8efdkjfdshf 6y agoCost and availability, especially at the higher capacities. Note that HGST was bought out by WD in 2012, and WD started sunsetting the HGST brand in 2018.
- atYevP 6y agoYev from Backblaze here -> Some of the responses down below are correct - it really does come down to price. We run a fairly lean ship and so the price per gigabyte is weighed pretty heavily when we make our purchasing decisions. We have a post where we go into the cost of hard drives over time, it's about 3 years old now, but still a good read -> https://www.backblaze.com/blog/hard-drive-cost-per-gigabyte/ https://www.backblaze.com/blog/hard-drive-cost-per-gigabyte/.
- ping_pong 6y agoWhy did they stop buying WDC drives? Is there a known issue with them? Also, why can't I find HGST drives for decent prices? On Amazon they are either refurbished drives or crazy expensive prices for new ones?
- Unklejoe 6y ago> they are either refurbished drives Just a heads up: there's no such thing as a refurbished hard drive. It's not economical. Instead, those drives are actually just used drives with the SMART counters cleared. I've been personally burned by this.
- neilv 6y agoSo they're just rolling back the odometer, and destroying records of accidents (errors)?
- kayson 6y agoYes and no. It's more like they're "recertified". It's pretty easy and non-invasive to replace the logic board, which can fix many problems that might warrant an RMA. I've also gotten refurbs back from Seagate and WDC RMA's that have new top covers on the drive (evidenced by a new, non-retail sticker, no old sticker underneath it). Presumably they're doing some inspection of the platters and heads before sending them out. These drives came with a 30day warranty and cleared SMART data. But I would argue that its fair to reset the SMART data after this kind of refurbishment/re-certification. That being said, there are definitely some sketchy drive resellers on marketplaces like Amazon who just clear the smart data on old drives and sell them as refurbished, or even new.
- jandrese 6y agoRecertified/Refurbished drives are never worth it IMHO.
- 6y ago
- andy4blaze 6y agoAndy at Backblaze here. To put a pin in it, the three main factors for which drives we use are cost, availability and reliability. We have control over reliability as our systems are designed to deal with drive failure. That leaves the market to decide on cost and availability. Assuming a competitive market we can buy the drives that optimize those factors.
- guenthert 6y ago> We have control over reliability as our systems are designed to deal with drive failure. That surely assumes an upper limit of the likelihood of a drive failure. There was a perception that the quality of 3.5" floppy disks declined drastically in the early 21st century. Must we not fear something similar for spinning-rust hard drives once most everyone uses SSDs?
- skj 6y agoBriefly, any drive (floppy disk or tape drive) has some likelihood of failure. You can minimize loss of data (the reliability being discussed) by replicating data in more than one storage item. It just becomes a matter of how many you buy (and how good you are at keeping them all properly organized).
- guenthert 6y agoToday reliability is sufficient that one can meet a given data availability goal by replicating the data 2, 3, 4, 5, <whatever> times as there is only once in a blue moon a bad batch of drives when they tried out a new bearing lubricant or so. But what if the economic incentives decline, the marked breaks apart (as it arguably does), much like it happened for floppy disks once they were (perceived as) obsolete and used only in fringe application (HP logic analyzers come to mind, but also Boing airplanes). Is there not the danger that the quality drops drastically to the point that one would need an unreasonable number of copies?
- kbenson 6y ago
- gruez 6y agoAre there trends for AFR by drive age? ie. for a specific drive or manufacturer, what is the failure rate for drives that have been in use n years? It'd be interesting to see how the failure rate go up/down as they get older.
- chiph 6y agoElectronics typically follow a "bathtub" reliability curve. You'll have a large number of early failures (unflatteringly called "infant mortality") then the curve levels off for a long period of time, and then starts rising again as the items start wearing out. https://en.wikipedia.org/wiki/Bathtub_curve https://en.wikipedia.org/wiki/Bathtub_curve It'd be interesting to see if BackBlaze sees this in their drive populations.
- astrophysician 6y agoI really enjoy reading the Backblaze analysis on this every year, it's such a valuable and interesting data set. I do have one suggestion: it would be great to go one step further and add confidence intervals for the AFR estimates. E.g. if you see 0 drive failures, you don't really expect an AFR of 0 (that is not the maximum likelihood estimate), and the range of AFR's you expect for each drive decreases as a function of the number of drive days (e.g. if one drive has 1 day of use, we know basically nothing about AFR so confidence interval would be ~0-100% (not really, but still quite large), or it would be smaller if you wanted to add a prior on AFR). It would also be interesting to see the time-dependence (i.e. does AFR really look U-shaped over the lifetime of a drive?). That would require a dataset with every drive used, along with (1) number of active drive days, and (2) a flag to indicate if the drive has failed and of course (3) which kind of drive it is. Does Backblaze offer this level of granularity? EDIT: They offer the raw data dumps! https://www.backblaze.com/b2/hard-drive-test-data.html#downloading-the-raw-hard-drive-test-data https://www.backblaze.com/b2/hard-drive-test-data.html#downl... Backblaze, god bless you.
- deleted 6y ago[deleted]
- bronco21016 6y agoI've recently joined the WD shucking crowd so, unfortunately, there are no stats for the drives I'm running. The cost/GB has just gotten too ridiculously low on shucked drives and my array large enough that a failure or two isn't the end of the world, oh and 3-2-1 backup rule. I've always really enjoyed reading Backblaze's reports though and have made past buying decisions based on their information. I wonder if it might be possible for a community sourced version of these reports? A small app that checks SMART data and sends it to a central repository for displaying stats? Many from the self-hosted/homelab crowd are running these shucked drives so there has to be a large pool of stats out there if it can be gathered?
- andy4blaze 6y agoAndy at Backblaze here. We've looked at this and the one "issue" is the consistency of the environments from the community data. In our DCs, the drives are kept a decent temperature, hardly ever moved, and our DC tech sing to them every night (OK, just kidding about that last one). Community drives will comes from all types of environments, from pristine to dust bunny hell. Still, it might be interesting to compare the two data sets if a community cohort could be collected.
- touisteur 6y agoBe careful not to sing too loud at them https://m.youtube.com/watch?v=tDacjrSCeq4 https://m.youtube.com/watch?v=tDacjrSCeq4 Incidentally I've tracked down a HDD (spinning rust) performance (and later breakage) problem to a very loud (and I mean awfully painful, intense vibration on your eardrums) alarm test every wednesday... Do you monitor the noise in your DCs? Maybe for high peaks?
- chromedev 6y agoWhat is Wasabi doing different where they're able to outprice B2, and get better throughput?
- tln 6y agoI'm guessing you work at Wasabi and already know :) Just looking at the pricing pages, pure storage is cheaper at B2, but the API calls and egress are not free.
- ggm 6y agoThe volumes of units are high enough I think this is north of 'too small sample size for statistical validity' So the question in my mind is, how significant is the variance in the failure rates? Is this underlying manufacturing error tolerances, or is this shipping/deployment effects, or is this .. Aliens?