32 ms·
Backblaze Hard Drive Stats for 2018
- alinde 8y agoWould be interesting to also have metrics on failure per TB storage.
- theandrewbailey 8y agoI'm not sure how that would be useful, since terabytes don't fail. When a drive fails, it's effectively a brick with no terabytes.
- alinde 8y agoI was thinking as one failure of a 100TB disk has a very different impact of 10 failures of 1TB disks. It'd give some idea on how much data is lost due failures, no?
- atYevP 8y agoYev from Backblaze here -> Not sure if you'd get that metric from that data. We use Reed-Solomon erasure coding (https://www.backblaze.com/blog/reed-solomon/ https://www.backblaze.com/blog/reed-solomon/) to make sure that data is "rebuilt" should we lose drives (which happens all the time).
- deleted 8y ago[deleted]
- oliveshell 8y agoI suppose, but there’s no such thing as a single HDD that stores 100TB. The biggest you can get currently are (I believe) 14TB helium-filled drives.
- jsgo 8y agoMy guess is they meant 10TB as it would be a more "equal" comparison: 1 10TB drive 10 1TB drives
- oliveshell 8y agoThat makes way more sense. Didn’t think before posting!
- brianwski 8y agoDisclaimer: I work at Backblaze. > When a drive fails, it's effectively a brick with no terabytes. Interesting factoid: that isn't always true. What you describe is actually the CLEANEST type of failure, the drive suddenly becomes a brick. We replace the drive and rebuild it from parity. A way more interesting failure is when disk blocks start going bad at an unacceptable rate. Backblaze splits your data across 20 different hard drives in 20 different machines in our datacenter. The sub-parts we call "shards", a shard sits on one disk. Each shard has a SHA-1 checksum, so we know if each shard has been corrupted. If an individual shard is missing or corrupted, we know it needs to be rebuilt from parity. So when a drive is HALF-FAILED, we even have a procedure to pull the drive out, and then opportunistically copy whatever files we can recover onto a new drive, then put the new drive back into production. Any files we recover where they are in the correct filesystem location and their SHA1 says they have not been corrupted speeds up the rebuild. The reason the speed of rebuild is important is the whole concept of 11 or 12 "nines" of durability. We can't have more than 3 drives fail in any one group of 20 drives, and the faster the rebuild time, the less likely for 4 simultaneous failures. It plugs into the formulas in this blog post we did about durability: https://www.backblaze.com/blog/cloud-storage-durability/ https://www.backblaze.com/blog/cloud-storage-durability/
- zepearl 8y ago(thanks a lot - all extremely interesting) >> We can't have more than 3 drives fail in any one group of 20 drives... Wow, for me, subjectively, an low threshold - and I underderstand that each drive being hosted on a different machine protects you as well from a machine/controller failure (happened to me twice with the controller - both times it was very hard to diagnose and the experience in general has been terrible). Do you have as well "backups"? Or is that in the hands of the customers/users?
- brianwski 8y ago> Do you have as well "backups"? Or is that in the hands of the customers/users? If you store data in Backblaze, there is no "backup" of that data. If Backblaze ever lost 4 drives simultaneously and could not recover the data, the customer would lose data. This is much like Amazon S3. In general, we recommend a 3-2-1 backup strategy where there are 3 copies of the data, at least 2 copies on your site, and 1 copy in the cloud. You can read about that philosophy in our blog post here: https://www.backblaze.com/blog/the-3-2-1-backup-strategy/ https://www.backblaze.com/blog/the-3-2-1-backup-strategy/
- klodolph 8y agoI’m not sure that it’s especially useful to measure that way, which is why they wouldn’t report it. The chance that a given GB of data is on a failed disk is equal to the disk failure rate, regardless of disk size (>1GB). For large deployments, the concern is between failure rates and the amount of time it takes to rebuild data from a failed disk. For small deployments, my main concerns are whether disk failure takes a machine or volume out, causing availability problems. I’m trying to figure where failures per GB would be how you would choose, what scenario we’re you thinking of?
- krob 8y agoHGST look like the best, but they don't have the quantities of the Seagate, makes me wonder if these numbers are skewed :/
- simcop2387 8y agoThey're not going to be skewed. HGST disks are usually more expensive, so that limits the quantity that they buy. The seagate disks are usually cheaper but have a slightly higher (except when it's a brand new line) failure rate. When filling out a single server I go with the HGST disks because the premium price and quality means fewer failures, but it's more cost effective to go for lots of seagate disks when you have more redundancy and can eat more failures.
- level 8y agoI built a server a few years ago, and I determined it was more cost effective for a drive to fail than to use HGST. It would be more inconvenient, but having a drive fail on a home server with only 6 drives didn't seem very likely anyway.
- cm2187 8y agoParticularly if it fails within the warranty period.
- jacobolus 8y agoIf you have a 2%/year failure rate per drive, then that leaves you with a nearly 12%/year chance that at least 1 of 6 drives will fail each year. Or a 31% chance that at least 1 of 6 drives will fail within 3 years. Or a 52% chance that at least 1 of 6 drives will fail within 6 years.
- sangnoir 8y agoThe question then is, would it be cheaper to replace that one drive or get the more expensive disks with lower chances of failing?
- samstave 8y agoBackBlaze spent ~12 million dollars on 12TB Seagate drives (at full retail)
- metalliqaz 8y agoGood thing they charge me $0.94 a month for my B2 storage. Gotta make that budget from somewhere!
- atYevP 8y agoEvery little bit helps ;)
- brianwski 8y agoDisclaimer: I work at Backblaze. > Good thing they charge me $0.94 a month for my B2 storage. We thank you for your business! :-) The absolute beauty of Backblaze B2 (or Amazon S3 or Azure) is that we can build a storage system at scale, and sell off all the little pieces of that. You win because you get a fair price on a sliver, and we win because we add up 100,000 customers like you and make about $1 million per year. The very best business is where the customer and provider are happy with the relationship.
- leowoo91 8y agoAWS nerd here, I recently googled you guys and found out you have 1/4 bandwith cost already. See you soon for my upcoming project.
- metalliqaz 8y agoI like my B2. I also like that Backblaze is one of those companies that does one thing and does it very well.
- samstave 8y agoEDIT: It may have appeared that I was denigrating them for the spend... NOPE I was just curious when I saw the # of drives as the largest chunk in this writeup - to see how much that chunk cost them.
- GNU_IS_UNIX 8y agoDo larger capacity drives have higher failure rates?
- simcop2387 8y agoIt's usually not so much that larger drives are inherently less reliable, but that the larger drives are also the newest lines and so yields and manufacturing issues are more likely to crop up than in the older lines that have had issues sorted out.
- SiempreViernes 8y agoI don't think so, it's just that they didn't include their older and smaller drives, from previous report I remember a segate (I think) 3 TB disk that was abysmal.
- platters 8y agoOnly if you yell at them.
- linsomniac 8y agoAnything interesting in this one? I've stopped reading them because they all seem to be "Seagates fail kind of a lot but we use them because reasons. HGST doesn't fail a lot, but we also have statistically insignificant numbers of them, so <shrug>."
- SiempreViernes 8y agoNo, the opportunity to do a better analysis hasn't gone away.
- metalliqaz 8y ago"because reasons" is an overly negative way to say "because they offer the best value" They can get large numbers of them cheap, and their system is good at detecting and replacing bad drives, so why not use them?
- linsomniac 8y agoIn my head it wasn't negative, it was that there are a lot of reasons that I don't think needed to be gone into. Offering value is one, having systems that are designed to minimize the impact of failures is another. Others that came to mind when I wrote that include: Being able to get them at the quantities they need, having systems that reduce the COST of failures (drive replacement and RMA), cost of RMAing 10 drives is < 10x the cost of RMAing 1 drive, having drives available in the SIZES Backblaze wants. And there are others I could speculate about but don't have as concrete information on (vendor relationships, manufacturer relationships, marketshare, firmware quality/suitability, temperature). Maybe it's just my social/business groups, but "because reasons" doesn't have to have a negative meaning. I use it as more a statement of fact: There are reasons for this.
- metalliqaz 8y agohmmm, I always thought "because reasons" was sarcastic, as in the person would list reasons but they are all bullshit. Of course, sarcasm can be impossible to discern correctly on Internet forums and social media. I admit I could be completely wrong about this.
- peterwwillis 8y agoI'm interested in failure rate per iops. If the drive fails infrequently, great, but if it's also the worst performing drive, screw that. Would rather buy drives that perform as well as possible with the least failure rate.
- walrus01 8y agoIn my recent experience there is not a lot of speed difference anymore between multiple manufacturers' 6TB to 12TB sized hard drives, when comparing between two competing products in the same rpm class (5400 or 7200) and areal density. Assuming similarly sized RAM cache on drive and not something like a drive with 64GB of SSD cache (hybrid drive).
- loeg 8y agoTo put it another way, everyone's pushing up against the same physical limits of spinning rust.
- peterwwillis 8y agoI mean more like benchmarked performance. Two drives with the same specs may end up performing differently. I realize benchmarks are not entirely realistic and tuning can affect the outcome, but if there's a clearly outsized performance difference between two seemingly equivalent products, I want the one that performs better and fails the least. So, aggregate random read and write operations per second over failure rate, as a general spec. (this might also expose flaws in the stats, if one drive model is getting predominately more of a certain operation which results in more failures)
- akulbe 8y agoI realize my comment is a tangent. I'm hoping folks might be understanding and hopefully have some advice. If you have terabytes to back up, are there still any backup services left that'll let you ship them a drive for a faster initial backup?
- ajford 8y agoAWS has the Snowball program. They ship you what is essentially a network storage device with 10G ethernet, and you upload data over your local 10G network using their protocols, then ship back and they ingest into S3/Glacier. Haven't used it personally, but did a fair amount of digging ~3 years ago. It was a potential backup solution, but we ended up shipping a 4U server full of 4TB drives to a partner for off-site backup, since they needed routine access anyways. Check out the AWS page: https://aws.amazon.com/snowball/ https://aws.amazon.com/snowball/
- brianwski 8y agoDisclaimer: I work at Backblaze so I'm biased. :-) > If you have terabytes to back up, are there still any backup services left that'll let you ship them a drive for a faster initial backup? Backblaze offers a "Backblaze Rapid Ingest Fireball" to allow you to ship us 60 TBytes of data on an appliance. https://www.backblaze.com/blog/introducing-backblazes-rapid-ingest-service-fireball/ https://www.backblaze.com/blog/introducing-backblazes-rapid-... If you only have 2 - 10 TBytes, I suggest you get a faster network connection, or carry your laptop to a location (like your work place, or a library, or your neighbor's house) with a fast connection and just upload it. You might be surprised how easy and fast it is to upload a couple of TBytes nowadays. Using the Backblaze Personal Backup with 30 threads, I can upload about 1 TByte every 12 hours or so. So if you can leave your laptop at your workplace for 4 days you can upload 8 TBytes, then bring your laptop back home for the incrementals.
- post_break 8y agoIt's been over a year, can you please ask someone at backblaze to add a single line of code into the snapshots so you can leave a comment? Right now if you make a snapshot, the only info is the date and time. Not the data, what's in the snapshot, literally anything about it. Please!
- zachruss92 8y agoI really appreciate BackBlaze opening up this data. While I don't purchase HDDs often, I always refer to these reports when deciding. They also open sourced their server chassis which is awesome! It's nice to see a reduction of the failure rates as a whole. It looks like the next few years will be some interesting times for the growth in storage capacities for HDDs.
- atYevP 8y agoYev from Backblaze here -> Glad you're enjoying the stats!
- HankB99 8y agoPlease thank those responsible on my behalf (and all of the others who study the numbers before purchasing drives.)
- atYevP 8y agoI'll let Andy know - he might be around here somewhere :D
- ericd 8y agoThese reports were super helpful when we were choosing how to kit out our servers (we went all HGST as a result). It's pretty hard to find good reliability data otherwise. So, a really big thanks from us.
- JustSomeNobody 8y agoThanks! You guys rock.
- sigi45 8y agoI wonder if any of those companies talk to Backblaze about it. Like sending drives back for inspection :) I'm also curious why those companies wouldn't directly talk to backblaze. I read somewhere a blog post on how they bought specific drives online at a sale.
- atYevP 8y agoYev from Backblaze here -> > I read somewhere a blog post on how they bought specific drives online at a sale. Yea, we used to buy drives wherever we could, but that was years ago. We're larger now so we go through more established channels.
- fpgaminer 8y agoTangentially related. Whenever I get a new drive, I always do a "burn-in". A program writes data to the whole drive and then reads it back (reproducible random data). Is there any real justification for doing this kind of test on a new drive? Doing it takes quite awhile, so I've been wondering lately if it's even worth it. I've never found anything with it.
- conbandit 8y agoIf you've never found anything with it, why do you keep doing it (see: the definition of insanity)?
- cataflam 8y agoNot the parent, but probably because 1. It's notorious that hard drives have a higher failure rate at the beginning of their lives than in the middle (see bathtub curve [0]). So it's not absurd to test them hard early on before writing any useful data and to do an early RMA. 2. The failure rate on drives is low enough that his methodology may be right but he still never has any failure in his life. Doesn't it make insane. [0] https://en.wikipedia.org/wiki/Bathtub_curve https://en.wikipedia.org/wiki/Bathtub_curve
- kbutler 8y agoIt would depend on the effort to do the methodology vs. the expected return (savings of finding a failed drive times the probability).
- philliphaydon 8y agoI once bought a new 1TB Drive when they were fairly new. MOVED about 500gb if data to the new drive. Checked it. Seemed fine. Turned computer off and went to bed. Next day the HDD didn’t turn on. Completely dead. :( I’ve never had a failure since but I backup everything now.
- e40 8y agoI have. Last time I bought a bunch of drives for a RAID array, I used the WD utility to test each of them. 1 of the 6 failed the test.
- humantiy 8y agoCurious if anyone knows how they calculate drive days. For example the first drive on their report (hgst 4tb) has a count of 50 but total days of 23069. If I take the 50 by the days it should be 18250, so not sure where the extra 4k in days is coming from. Retired drives or something?
- tzs 8y agoI think it is over the time they have had the drive, not over the reporting interval. It's a measure of the age of the drives. Assuming their drives are operating 24/7, that means that those particular 50 drives have been in service an average of 461 days. I'd expect on next year's report, those particular drives will show up as 49 drives with around 42000 drive days, assuming they aren't replaced by then.
- humantiy 8y agoIf that is the case then wouldn't the Annualized Failure Rate be based of the year total not the drive days if it is total days in service? For example the drive count(50)/drive days(23,236) gives the AFR of 1.58% which equals out their numbers. The drive days is more than the total amount possible for that year.
- mikece 8y agoWhile I've heard lots of great things about the Backblaze reports, I've noticed in comments on NewEgg, Amazon, etc that the SKUs mentioned in the report frequently aren't available anymore. I've never had problems with WD Red drives though I don't purchase them all from the same vendor at the same time to make sure I get drives from different lots in case of a lot defect.
- wmf 8y agoBy the time you have accurate reliability data on any equipment it's obsolete. Maybe this will change with the slowdown of Moore's Law/Kryder's Law.
- mark-r 8y agoA perennial favorite, literally. Thank you.
- atYevP 8y agoYou're welcome!
- charliebrownau 8y agoAnyone that has been in IT for 20-30 years really surprised with the number of seagates that fail
- Rebelgecko 8y agoI wonder if they have states failure rate based on a drive's manufacturing or installation date? It looks like there's a bit of a bathtub curve, and it would be interesting to see if that's attributed to individual drives having a tendency to fail quickly (if they're going to), or if drives are less likely to crap out once their model has been manufactured for a few years
- linux2647 8y agoIs there anything similar for SSDs?
- icelancer 8y agoMy company uses Backblaze since I think it's a good product, but blogs like this really cemented my choice. I appreciate their attention to detail and publishing data openly.
- deleted 8y ago[deleted]
- atYevP 8y agoYev from Backblaze here -> That's awesome to hear! I'm glad you're with us! That's one of the nice side-benefits of this blog and one of the reasons we adopted an "open" policy with the Storage Pods. The first time we published that post it was because folks didn't believe we could store data so inexpensively - so it's nice to hear that we're building some trust along the way!
- chemmail 8y agoRelibality looks decent this year. All my seagates are having tons of errors, i think i'll stick to the other team from now on even though seagates seems to be getting better.
- b3lvedere 8y agoThank you Backblaze! I love your reports. What is your procedure/policy on which disks to use in the pods? Do you try and maybe control the risk by using different harddisk brands in a single storage pod? Or do you just not care, because there have never been 3 pods dead at the same time? :) Do you still use 17 data plus 3 parity shards?
- evil-olive 8y agoTheir Q3 2018 stats had a bit of info on the lifecycle of introducing new disks: https://www.backblaze.com/blog/2018-hard-drive-failure-rates/ https://www.backblaze.com/blog/2018-hard-drive-failure-rates... > In Q3 we added 79 HGST 12TB drives (model: HUH721212ALN604) to the farm. While 79 may seem like an unusual number of drives to add, it represents “stage 2” of our drive testing process. Stage 1 uses 20 drives, the number of hard drives in one Backblaze Vault tome. That is, there are are 20 Storage Pods in a Backblaze Vault, and there is one “test” drive in each Storage Pod. This allows us to compare the performance, etc., of the test tome to the remaining 59 production tomes (which are running already-qualified drives). There are 60 tomes in each Backblaze Vault. In stage 2, we fill an entire Storage Pod with the test drives, adding 59 test drives to the one currently being tested in one of the 20 Storage Pods in a Backblaze Vault.
- b3lvedere 8y agoThank you for the info! Much appreciated.
- deleted 8y ago[deleted]
- arcaster 8y agoTIL - Seagate drives are still basically shit :)
- Arn_Thor 8y agoOdd they don't have the 6TB HGST on the list. Well, maybe not odd, but annoying since I'm curious about that drive's performance especially