5 ms·
This is one reason why RAID 0+1 is a best practice, and RAID 5 & 6 are no longer recommended. It takes too long to rebuild the array, leading to a multi-failed
by stephengillie 8y ago
This is one reason why RAID 0+1 is a best practice, and RAID 5 & 6 are no longer recommended. It takes too long to rebuild the array, leading to a multi-failed disk situation.
- chrisper 8y agoRaid 01 has its own risk. If the wrong two disks fail your entire array is toast
- nine_k 8y agoYou can upgrade a 2-disk RAID1 to a 3-disk RAID5, then chain them to RAID0 as normal. It gives you a better chance to keep data intact, hopefully without lowering the write speed seriously. https://en.wikipedia.org/wiki/Nested_RAID_levels#RAID_50_(RAID_5+0) https://en.wikipedia.org/wiki/Nested_RAID_levels#RAID_50_(RA...
- zaarn 8y agoRAID 50 doesn't really solve it, it exposes you to some more risk since you can still die with 2 disks but now you have more disks in each sub array. The correct answer is either a 3-mirror RAID1 or RAID6. Bcachefs also promises some solution to this by allowing both erasure encoding and replication to co-exist, according to it's documentation.
- throwaway2048 8y agoas opposed to raid 5 where if any two disks fail your array is toast, raid 6 increases this to 3. However both raid 5 and 6 have 2 huge problems: Data inflight at write time (power/hardware failures are more likely to corrupt the array, especially silently, which is the worst outcome). Parity calculations require you to spin up the whole raid5/6 array during a rebuild, massively increasing the chance of a multi drive failure and a lost array. If one close-to-EOL drive dies, putting its sister drives through what is essentially an all day full tilt stress test is a terrible, terrible idea, and this idea keeps getting worse (takes longer) as drive sizes grow. raid 0+1 sidesteps these issues mostly at a modest increase in drive count, its a no brainier for most setups.
- StillBored 8y agoData inflight at write time (power/hardware failures are more likely to corrupt the array, especially silently, which is the worst outcome). How is that? RAID doesn't affect data persistence behavior in any meaningful way. FUA/SyncCache/etc are supported by RAID controllers same as the underlying disks in writeback enviroments, parity updates included. Put another way, if you FUA or flush the writeback cache, those operations won't complete in a properly implemented RAID environment until the data is persisted somewhere, even if that means passing FUA down to the underlying storage. Granted there are a number of ways to mess this up, RMW cycles in a controller that doesn't have some kind of persistent memory and flush on power restore. Anyway, none of this is any worse than what happens in any other WB cached storage technology. Finally, all this fearmongering about loss on rebuild is also something that should be more fully explored in the context of the fact that decent RAID systems run background scrub operations on a regular basis. Those operations by themselves are going to "stress test" the array on a regular basis when its consistent and not degraded. I've actually got a fair amount of experience in this area, and I'm here to tell you that if you think this is a risk consider what happens to non-raided unscrubbed drives that have a lot of data silently bitrotting on the platters. That latter effect is nearly always the problem in RAID environments when someone starts a rebuild on drives/sectors that have been unread for extended periods of time. But, in the case of RAID, a properly implemented system won't fail a drive for a single read failure during a rebuild, instead reconstructing from the other drives and leaving the drive online long enough to complete the rebuild and then taking it offline. Basically raid 1 setups don't actually fix any of these problems, except through the use of massive additional parity disks overhead. Overhead that can also be applied to other RAID algorithsm to much better effect. AKA a mirrored RAID 6 provides far more protection than a mirrored raid 0. Similar levels can be had with 6+6 in environments where that is possible, with trivial capacity overhead.
- throwaway2048 8y agoRaid 5/6 require parity calculations before data can be written to disk. This is a significant amount of data, especially at high writing speeds. That is what causes the inflight data problem. Battery and flash backup on controllers dosen't fix the problem of hardware failure (which is significant, especially on big hot controllers.
- stephengillie 8y agoMultiple "0" drives can be added for further redundancy.
- astrodust 8y ago"Zero" drives are the ones that when you lose them you have zero data. "One" drives are the ones with a copy.
- stephengillie 8y agoI haven't worked with arrays for years. Sorry for the mistakes.
- deleted 8y ago[deleted]
- yellowapple 8y agoThe normal answer here is to make sure that each side of the RAID10 (RAID01 is something different and much less common) mirror uses drives from a different vendor, thus giving each side a different bathtub curve / failure rate and mitigating the impact of a bad batch. This is a nice advantage over parity-based setups like RAID6 (since replicating this with RAID6 would require finding a unique vendor for each array member, and there are only so many vendors). For archival purposes, though, you're probably better off with a normal RAID1 + some kind of JBOD setup (like with LVM); striping makes data recovery more difficult should you indeed lose all RAID1 sides of a given member.
- nine_k 8y agoRAID6 should be fine rebuilding online (in RAID5 mode) even under a moderate write load. Of course one should source RAID disks form 3 different vendors, to ensure that they are from different batches, and are not going to fail at approximately the same time.
- stephengillie 8y agoDo other manufacturers produce this size of drive? It's difficult to source from 3 vendors if there's only one making the product.
- thfuran 8y agoGet one from amazon, one from newegg, one from the manufacturer directly or some such.
- lostapathy 8y agoI try to buy hot spare or the last drive in a raid6 later than the rest of the array to try to spread them out too.
- tracker1 8y agoGood advice, though I once had about half a dozen drives (12 drive RAID Z2 with 2 as hot spares) fail within a few weeks of each other in separate batches from sourcing. (Seagate 3TB drives, I think there's been articles on how bad that series was).
- tfigment 8y agoI don't know how i survived those Seagates. Lasted maybe a year and started dropping like flies. Synology seems to recommend identical drives as i recall but work fine with different sizes and makers afaict.
- Hei1Fuya 8y agoZoL 0.8 will have sequential resilver which should be able to restore a disk in a few hours.
- ahoka 8y agoInteresting. What do you think the advantage of raid01 instead of raid10? The latter looks safer at first sight.
- stephengillie 8y agoI get RAID 01 and 10 mixed up all the time. These names are too similar. Please understand that I meant the better of the 2.
- jmpman 8y agoCRUSH algorithms are used to overcome rebuild limits in modern arrays. https://www.ssrc.ucsc.edu/Papers/weil-sc06.pdf https://www.ssrc.ucsc.edu/Papers/weil-sc06.pdf
- kdkeyser 8y agoCRUSH is an example (and not the first) of a "distributed rebuild" approach: you have an array of N drives (with N large, e.g. 100), and if 1 drive fails, you read in parallel from all (N-1) remaining drives, while distributing the reconstructed data across the remaining available capacity of all (N-1) remaining drives. In effect, you get the total bandwidth of (N-1) HDD's working in parallel. And the bandwidth of 100 HDD's doing sequential IO in parallel is really massive ( ~ 10 GB/s). Examples of companies claiming to use this approach are Qumulo (rebuild in couple of hours), Infinidat (couple of 10's of minutes), ClusterStor GridRAID (now part of Seagate I think), or "Declustered RAID" in GPFS (IBM)
- pinewurst 8y agoGridRAID is owned by Cray now, who were the primary OEM from Seagate. Thanks for pointing out that declustered/distributed rebuild RAID has many historical precedents (also 3PAR BTW) pre-CRUSH/Ceph.