5 ms·
Take these stats with a grain of salt. I am becoming more and more convinced that hard drive reliability is linked to the batch more than to the individual dri
by CTDOCodebases 2y ago
Take these stats with a grain of salt.
I am becoming more and more convinced that hard drive reliability is linked to the batch more than to the individual drive models themselves. Often you will read online of people experiencing multiple failures from drives purchased from the same batch.
I cannot prove this because I have no idea about Blackblazes procurement patterns but I bought one of the better drives in this list (ST16000NM001G) and it failed within a year.
When it comes to hard drives or storage more generally a better approach is protect yourself against down time with software raid and backups and pray that if a drive does fail it does so within the warranty period.
- bragr 2y ago>Often you will read online of people experiencing multiple failures from drives purchased from the same batch I'll toss in on that anecdata. This has happened to me a several times. In all these cases we were dealing with drives with more or less sequential serial numbers. In two instances they were just cache drives for our CDN nodes. Not a big deal, but I sure kept the remote hands busy those weeks trying to keep enough nodes online. In a prior job, it was our primary storage array. You'd think that RAID6+hot spare would be pretty robust, but 3 near simultaneous drive failures made a mockery of that. That was a bad day. The hot spare starting doing its thing with the first failure, and if it had finished rebuilding before the subsequent failures, we'd have been ok, but alas.
- tharkun__ 2y agoThis has been the "conventional wisdom" for a very long time. Is this one of those things that get "lost with time" and every generation has to rediscover it? Like, 25+ years ago I would've bought hard drives for just my personal usage in a software raid making sure I don't get consecutive serial numbers, but ones that are very different. I'd go to my local hardware shop and ask them specifically for that. They'd show me the drives / serial numbers before I ever even bought them for real. I even used different manufacturers at some point when they didn't have non consecutive serials. I lost some storage because the drives weren't exactly the same size even though the advertized size matched, but better than having the RAID and extra cost be for nothing. I can't fathom how anyone that is running drives in actual production wouldn't have been doing that.
- itchyouch 2y agoExactly this. I mostly just buy multiple brands from multiple vendors. And size the partitions for mdadm a bit smaller. But even the same model where it's 2 each from bestbuy, Amazon, newegg, microcenter, seems to get me a nice assortment of variety.
- lazide 2y agoIt’s inconvenient compared to just ordering 10x or however many of the same thing and not caring. The issue with variety too is different performance characteristics can make the array unpredictable. Of course, learned experience has value in the long term for a reason.
- Aachen 2y agoI had to re-learn this as well. Nobody told me. Ordered two drives, worked great in tandem until their simultaneous demise. Same symptoms at the same time I rescued what could be rescued at a few KB/s read speed and then checked the serial numbers...
- tempest_ 2y agoI personally like to get 1 of every animal if I can. I just get 1/3 Toshiba, 1/3 WD, 1/3 Seagate.
- FuriouslyAdrift 2y agoNearly every storage failure I've dealt with has been because of a failed RAID card (except for thousands of bad quantum bigfoot hard drives at IUPUI). Moving to software storage systems (ZFS, StorageSpaces, etc.) has saved my butt so many times.
- wil421 2y agoSame thing I did except I only wanted WD Red drives. I bought them from Amazon, Newegg, and Micro center. Thankfully none of them were those nasty SMR drives, not sure how I lucked out.
- dapperdrake 2y agoMy server survived multiple drive failures. ZFS on FreeBSD with mirroring. Simple. Robust. Effective. Zero downtime. Don’t know about disk batches, though. Took used old second hand drives. (Many different batches due to procurement timelines.) Half of them was thrown out because they were clicky. All were tested with S.M.A.R.T. Took about a week. The ones that worked are mostly still around. Only a third of the ones that survived S.M.A.R.T. have failed so far.
- CTDOCodebases 2y agoI didn't discover ZFS until recently. I played around with it on my HP Microserver around 2010/2011 but ultimately turned away from it because I wasn't confident I could recover the raw files from the drives if everything went belly up. Whats funny is that about a year ago I ended up installing FreeBSD onto the same Microserver and ran a 5 x 500GB mirror for my most precious data. The drives were ancient but not a single failure. As someone who never played with hardware raid ZFS blows my mind. The drive that failed was a non issue because the pool it belongs to was a pool with a single vdev (4 disk mirror). Due to the location of the server I had to shut down the system to pull the drive but yeah I think that was 2 weeks later. If this was the old days I would have had to source another drive and copy the data over.
- tempest_ 2y agoZFS is like magic. Every time I think I might need a feature in a file system it seems to have it.
- Kelvin506 2y agoIME heat is a significant factor with spindle drives. People will buy enterprise-class drives, then stick them in enclosures and computer cases that don't flow much air over it, leading to the motor and logic board getting much warmer than they should.
- MrDrMcCoy 2y agoHeat is also a problem for flash. If you care about your data, you have to keep it cool and redundant.
- chuckledog 2y agoThis. My new Samsung T7 SSD overheated and took 4T of kinda priceless family photos with it. Thank you Backblaze for storing those backups for us! I missed the return window on the SSD so now have a little fan running to keep the thing from overheating again
- Kelvin506 2y agoWith the added complication that the controller should be kept cool, but the flash should run warm. The NVMe drives in my servers have these little aluminium cases on them as part of the hotswap assembly. They manage the temperature differential by using a conductive pad for the controller, but not the flash.
- CTDOCodebases 2y agoI have four of those drives mentioned and the one that did fail had the highest maximum temperature according to the SMART data. It was still within the specs though by about 6 degrees Celsius. The drives are spaced apart by empty drive slots and have a 12cm case fan cranked to max blowing over it at all times. It is in a tower though so maybe it was bumped at some time and that caused the issue. Being in the top slot this would have had the greatest effect on the drive. I doubt it though. Usage is low and the drives are spinning 24/7. Still I think I am cursed when it comes to Seagate.
- 10729287 2y agoThis is why it’s best practice to buy your drives from different dealers when setting up RAID.
- cm2187 2y agoWell to me the report is mostly useful to illustrate the volatility of hard drive failure. It isn't a particular manufacturer or line of disks, it's all over the place. By the time Backblaze has a sufficient number of a particular model and sufficient time lapsed to measure failures, the drive is an obsolete model, so the report cannot really inform my decision for buying new drives. These are new drive stats, so not sure it is that useful for buying a used drive either, because of the bathtub shaped failure rate curve. So the conclusion I take from this report is that when a new drive comes out, you have no way to tell if it's going to be a good model, a good batch, so better stop worrying about it and plan for failure instead, because you could get a bad/damaged batch of even the best models.
- immibis 2y agoI looked at this last report and I came to the same conclusion I did in the first report: Seagate drives are less reliable than WD.
- deleted 2y ago[deleted]
- deelowe 2y ago> I am becoming more and more convinced that hard drive reliability is linked to the batch more than to the individual drive models themselves. Worked in a component test role for many years. It's all of the above. We definitely saw significant differences in AFR across various models, even within the same product line, which were not specific to a batch. Sometimes simply having more or less platters can be enough to skew the failure rate. We didn't do in depth forensics models with higher AFRs as we'd just disqualify them and move on, but I always assumed it probably had something to do with electrical, mechanical (vibration/harmonics) or thermal differences.