15 ms·
I wouldn’t feel comfortable with RAID5/6 even on today’s ~10TB drives. The solution isn’t more parity disks, it’s to move away from that RAID set up entirely.
by 88 6y ago
I wouldn’t feel comfortable with RAID5/6 even on today’s ~10TB drives.
The solution isn’t more parity disks, it’s to move away from that RAID set up entirely.
- gbrown_ 6y ago> I wouldn’t feel comfortable with RAID5/6 even on today’s ~10TB drives. I fully agree with this part. De-clustered RAID really should have been mainstream a long time ago. Thankfully dRAID in OpenZFS will bring this to the "masses" in the sense of it being an open source implementation, whereas this has traditionally been a proprietary feature. > The solution isn’t more parity disks, it’s to move away from that RAID set up entirely. But I strongly disagree with RAID being the problem. Rather the increasing of capacity with bandwidth not really improving is the issue. Of course this is a inherent limitation of disk drives.
- Robotbeat 6y agoThey are adding multiple actuators, which actually should help that fundamental limitation.
- altcognito 6y agoCould easily reduce reliability.
- Robotbeat 6y agoIt’s possible, just as increasing the density of NAND with 4 Levels per cell reduces write count. But can be countered by more careful engineering of the device.
- altcognito 6y agoNAND failures tend to be more localized though.
- Robotbeat 6y agoThis is probably not representative because it was using low grade USB flash devices from 10-15 years ago, but I think I’ve had more NAND flash failures than hard drive failures. I don’t think doubling the number of heads would make a huge difference as you wouldn’t be quite halving the reliability.
- altcognito 6y agoOh, I totally agree in that regard. Hard drive technology has been mature for a good number of decades, so a lot of the kinks have been worked out.
- jacquesm 6y agoNo, it will reduce reliability. More parts = worse reliability.
- Robotbeat 6y agoMaybe? The issue is mechanical devices can be engineered until the reliability level needed is achieved. Sure, take the same device and double the numbers, you’ll have half the reliability, but you’ll also have more parts being made and more parts to amortize the reliability improvement engineering over (and more parts to get reliability statistics from... which also means you can test the rest of the hard drive faster, which can provide other reliability enhancements). NAND devices have similar constraints, actually, but they’re reaching more fundamental limits due to quantum effects and they’re generally addressed with software/firmware tricks like wear leveling.
- jacquesm 6y ago> The issue is mechanical devices can be engineered until the reliability level needed is achieved. And if wishes were horses... No, you can not simply state some arbitrary level of reliability and design to that level. At any given level of reliability more parts = less reliability and more time or resources spent on design are not by going to remediate that unless you are willing to accept much higher costs. Engineering is a trade-off and physics determine the sweet spot for the balance of that trade off. Once that sweet spot has been determined all other things being equal more parts will cause your reliability to go down unless you will also accept that your costs will go up and/or other parameters will be affected in a negative way. Software people in general have a very hard time to understand this because to them 'parts' are free, but even in software, assuming zero costs for parts (adding a library, a function, a line of code) has an effect on reliability. That's why computers used to be much more reliable than they are today, and that's before we get into details such as cognitive load while trying to understand complex systems.
- rzzzt 6y agoWhere do those fit, in the opposite corner?
- notacoward 6y agoNot much, actually. Dual actuators only improve parallelism, not media transfer rate or latency. In practice, dual actuators don't even double external transfer rates due to occasional latency (which is unavoidable even for the best-case scenarios) and internal contention elsewhere in the drive. Even if you had four actuators working perfectly in concert, and weren't bottlenecked on the external interface speed (which you would be), a 4x improvement in transfer rate vs. a 8x increase in capacity would still mean a 2x increase in fill/empty time. As with the shift from 5.25" to 3.5" to 2.5" and even 1.8" drives, the only way out of this bind is really to take advantage of the improved areal density to make drives that are the same capacity but smaller and pack more of those into the same volume or power/heat envelope. Drive manufacturers could help this along a little bit e.g. by sharing a motor and some environmental bits between what are otherwise completely separate drives (including separate external interfaces) within a single package, but mostly we'd all better get used to higher drive counts. Dual actuators - and this is far from the first time they've been tried BTW - are mostly a red herring. Background: I worked on exactly these problems for the latter half of a thirty-year career, most relevantly at my last job working on an exabyte-scale storage system at a FAANG.
- effie 6y agoInteresting. What do you think about Seagate's claim that dual actuator decreases their costs because it takes less time to test the drive? Is that the real reason for dual actuator?
- PaulHoule 6y agoAlso what market is there for big hard drives outside of FAANG and a few other places? Retailers in my area (e.g. Best Buy) don't stock hard drives larger than 4 TB; I was going to tell somebody who lived in the Valley that he's lucky to be able to go to Fry's and then Fry's closed down. Those same retailers stock both budget and quality SSD's up to 2TB in size. Most upgraders and the system builders are happy. For the rest of us there is Amazon where the Seagate Exos "enterprise" drive costs half of what similar "consumer" drives cost, has a great reputation and does not seem hard to live with at home. I would not take it for granted at all that a backup, RAID rebuild, restore, metadata scan or any full scan would work on such a disk if I hadn't tested it -- it is just that kind of technology. Synology is not crazy at all when they make you buy branded large drives to go in the enclosure.
- rietta 6y agoRAID 1 still has benefit when coupled with proper backup. Can still replace a drive and keep working faster than doing a restore from backup for large drives.
- rzzzt 6y agoWith current failure rates, isn't one running the risk of hitting a URE while creating the first backup?
- ghaff 6y agoAt the right cost points have a three-way mirror?
- rietta 6y agoA RAID does not increase the risk of a drive failure. It via redundancy reduces the risk, but not to zero. Probability of Single Drive Failure > Probability of Double Drive Failure. RAID is not a backup, so still need that backup. 1. Without RAID, drive fails and you're stuck waiting for restore from backup. 2. With RAID, risk double drive fails < risk of #1 and are stuck waiting for restore from backup. 3. With RAID, risk of single drive fails > 0 and < #2 and continue working while waiting on drive clone while still having the backup in #1 and #2 to fall back on.
- rietta 6y agoCompared to what? The general argument for "no raid" is to have a single drive without redundancy other than backup. That's riskier than a mirror.
- rzzzt 6y ago120 TB drives compared to 4-8-12-ish TB drives. I'm wondering if there is a chance of ever completing operations targeting the entire disk (be it an initial RAID build, resilver, restore or full backup) without hitting a speed bump.
- 6y ago
- wing-_-nuts 6y ago>But I strongly disagree with RAID being the problem. Rather the increasing of capacity with bandwidth not really improving is the issue. Of course this is a inherent limitation of disk drives. I mean, yes, you could have ssds in raid and not have to worry about the drive failing before it's rebuilt, but we're talking about spinning disks here. What's the current recommendations around raid capacity before you have to seriously start worrying about drive failure before it can be rebuilt? This is a genuine question I don't know.
- sandworm101 6y ago>> What's the current recommendations around raid capacity before you have to seriously start worrying about drive failure before it can be rebuilt? That depends on the size of your drives and the number in the array. If you are running a small NAS with only one parity drive, you don't want drives that take more than a day to onboard. That limits you to 4/6TB per drive. If you have two parity drives in the array you can probably risk the 8/10/12 TB size. If you are running commercial-scale arrays of dozen of drives, arrays that can handle multiple failures at once, then the sky is probably the limit. But the story gets more complicated. Are all of your drives the same age? Are all the drives of the same model/manufacturer? Such things increase the risk ofone failure evolving into a multi-drive failure during rebuild. If/when I build a new array I want drives of different ages. So If I had 10 drives, I would start the array on five or six, holding some for later so that the entire array isn't the same age.
- effie 6y agoDrive failure during a rebuild isn't a real concern unless 1) the rebuild takes time comparable to characteristic lifetime of the drive, which on average is >5years. Rebuild times are more in the realm of days, maybe coming to weeks with the proposed 100TB models. Still far from substantial probability of failure. 2) the drives are very old or have bad SMART data. Then the probability of drive failure and (array loss) shoots up. What people often talk about regarding big drives and reliability concerns is probability of rebuild failure for RAID5, RAID6, i.e. some bit is read or written wrong and the array becomes inconsistent. This is much more probable but isn't really a big problem in practice because this can be detected and rebuild can be repeated.
- hinkley 6y agoDe-clustered RAID also has disk replacement times that are proportional to usage instead of capacity, right? Whereas RAID5 and RAID6 require replicating empty blocks? HDD are much faster when they're only half full.
- nsgi 6y agoDepends what you use them for. With a high-capacity hard drive you could store a lot of downloaded movies and you can always download them again if it fails
- 88 6y agoI think for most people having to download tens of terabytes of any content would be extremely challenging.