15 ms·
Place I work got hit HARD by this. Year and half before I started at this company they setup 4 SSD's in a RAID 5 config. One day while working on things I noti
by ScoJoh 8y ago
Place I work got hit HARD by this. Year and half before I started at this company they setup 4 SSD's in a RAID 5 config.
One day while working on things I noticed a message on our server indicating a PDR1001 Error (Predictive Failure). So we ordered a new one. The new SSD arrived. We popped in the new one and the RAID started to rebuild.... Lo and behold during that operation Drive 1 threw the same error and the whole thing came crashing down...
We ended up losing the whole array. I had NO idea we had SSD's in the system. I had no idea that no one was monitoring their life.... The moment I saw we had this issue I saw the writing on the wall. 4 SSD's in a raid 5 all installed at the same time... means all SSDs end up with the same approximate critical end of life.
All I could do was shake my head at the whole thing... Pretty sure those in the charge who setup the array still don't understand why this situation was 100% avoidable...
- arminiusreturns 8y agoBesides the ssds, who in their right mind uses raid 5 anymore? It's been dead to me for years... since the first time I had to rebuild an r5 of 4tb disks and did the math on the time window for cascades. Also, not to nitpick, but you shouldn't wait to order a replacement that gets shipped to rebuild, you should have spares ready to go at all times for all raid systems. That might have been the difference between a cascade and a normal rebuild.
- simcop2387 8y agoYea any time I'm using a raid setup (as opposed to a cluster or other more intelligent system like zfs) I insist that there be at least one hot spare and a cold one ready. And never let anyone convince you a raid is a backup
- Narkov 8y agoHot and cold spares are rather wasteful. Just RAID1 and then use software (i.e. ceph) to make that redundant across multiple enclosures and ultimately, geographic zones/DC's.
- hobs 8y agoDepends on your RPO and RTO tbqh, if you have the budget to do that in software you are either fairly small or able to implement changes without the business yanking your chain and many are willing to just pay more for their storage layer than re-architect existing apps.
- mmt 8y agoThat strikes me as a bit contradictory, since I would consider RAID1 (plus the redundancy across enclosures) to be the equivalent to even "hotter"[1] spares. Unlike the hot (warm) and cold spares for the RAID5/6, all the drives in the RAID1 are continuously in use, which means they're subject to wear-related failures [2]. If you're getting performance benefits from the RAID1, then the extra drives may not be wasted, but that's a separate topic. [1] Which I argue is a misnomer. To me, "warm spare" makes more sense, since it's powered up but not actively synced in any way. Some systems, configurably, even spin down such spares. [2] As I mentioned in https://news.ycombinator.com/item?id=17855632 https://news.ycombinator.com/item?id=17855632 some failures are power-on/spun-up related, which makes the availability of a truly cold spare even more beneficial.
- apta 8y ago> or other more intelligent system like zfs I thought zfs was not incompatible with RAID. How else would you run it out of curiosity?
- bcaa7f3a8bbc 8y agoWhen using ZFS on top of external RAID hardware or software, instead of using the builtin RAID-Z feature, many unique crash-safe and data-corruption prevention features provided by ZFS are lost. The best practice is avoiding using other RAID mechanisms when possible.
- alexsb92 8y agoAs someone looking to backup a few terabytes of photos and videos, where I do have them on a raid, what constitutes an actual backup?
- Max_aaa 8y agoAn actual backup, would be a copy of said photos and vidoes on: - An external Drive. (which is only connected to copy files) - An external cloud service. - Another computer that you have. You should probably do a couple of these. I personally backup home computers (using borg) to a home server, that server has a 2.5 2TB external HDD connected to it (2 other 2.5 external drives are kept outside of the house). A backup of important files from the nas (including the computer backups) gets copied over to the external drive nightly. Weekly the drives gets rotated. The really important stuff is also backed up offsite on a daily basis.
- AnIdiotOnTheNet 8y agoMy general rule is that there are at least 3 copies in different physical locations, at least 2 of which are not within "tornado distance" of each other. The way to think about backup and DR is "what would it take to destroy all of this?" and keep making the answer more and more extreme until it is so horrible that if it actually happens you won't care about what you lost. PS: Also always remember that a backup you don't test isn't actually a backup.
- simcop2387 8y agoBasically like others said, there's some kind of physical separation. That way if the machine were to explode or you loose too many drives in the array you still have a copy of them. RAID protects you from a hardware failure (to some extent), a backup protects you from hardware failure, software failure, and human failure (or it should anyway). The idea is that if the cannonical version of things gets destroyed somehow the backup is available to rebuild, so it can't be alive in the same machine as the cannonical one (except maybe when doing the backup itself). The other "rule" that many people follow is 2 is 1 and 1 is none. The idea is that anything that isn't backed up isn't protected and can't be relied upon to exist. A cloud service like backblaze, google drive, amazon cloud drive, etc. is a good secondary backup for a lot of people even if it'll take you a month or two to get your data there to begin with.
- jacquesm 8y agoWith the current drive capacities even RAID6 is at its limits (or in fact, already past them).
- metaphor 8y ago> since the first time I had to rebuild an r5 of 4tb disks and did the math on the time window for cascades. Care to elaborate on this? Disk specs? Resilivering time?
- vernie 8y agoWhat superseded RAID 5?
- ProblemFactory 8y agoGenerally variations on RAID 10. Disks are large and cheap enough to afford losing 50% of your capacity, and it's much faster while in use and when rebuilding.
- mmt 8y agoRAID6, which provides two drives worth of parity compared to RAID5's single drive worth of parity, is the most common successor, especially in hardware (ASIC) implementations. Other examples are ZFS's RAID-Z, which can support even more parity for even more resiliency to drive failure. Contrary to a sibling comment, RAID1+0 does not supersede RAID5, as it has existed at least as long and has always had different trade-offs. How many drive failures one may need to be able to survive, given historical drive failure rates [1], current drive capacity vs. transfer rates (i.e. minimum rebuild time), and individual parameters (e.g. number of drives per array, acceptable magnitude of performance degradation during rebuild), is left as an exercise to the reader. As the OC suggested, this number can still be 1 (i.e. RAID5) for SSDs. [1] optionally including or excluding "black swan" events such as the flooding in Thailand that wiped out thos disk factories, rendering an entire "generation" of HDDs, manufactured in haste elsewhere, far less reliable
- zaarn 8y agoI'm planning to use RAID5 in an upcoming N/DAS build. Though it's not a server-level setup, so I'm using SnapRAID+MFS which can tolerate more than 1 drive failure in this case, atleast without loosing the entire array. RAID5 definitely has a purpose, notable small arrays where RAID6 would be wasting space or if you have nodes to failover too (ie, you have 3 nodes, run each with raid 5, if one goes belly up during a rebuild you still have 2 left and you can reprovision the third). RAID6 would likely be at the limit of most modern implementations so you'd have to jump to (n,n-3) RAIDs or higher. Of course there is always RAID1(0) if you like 50% storage efficiency.
- mmt 8y ago> RAID5 definitely has a purpose, notable small arrays where RAID6 would be wasting space This doesn't quite make sense to me. Why would the size of the array make any difference as to whether or not the extra drives needed by RAID6 can be characterized as waste? > or if you have nodes to failover too In essence, that's RAID5+1, which makes me wonder: If you have node-level redundancy, why bother with the intra-node redundancy? > RAID6 would likely be at the limit of most modern implementations so you'd have to jump to (n,n-3) RAIDs or higher. It's not clear to me what limit you're referring to, but, considering how long RAID6 has been implemented, especially in hardware, and how far computing power has increased since then, such a claim is dubious, if not extraordinary.
- zaarn 8y ago> Why would the size of the array make any difference as to whether or not the extra drives needed by RAID6 can be characterized as waste? RAID risk is about your rebuild failure rate, in this case URE (Unrecoverable Read Error). If one such is likely to occur during your rebuild, that can lead to a lot of problems. In RAID5 until you hit about 2 TB disks and up to about 4 or 5 disks total the risk of a URE during the full read of the array is fairly low. Once you go above that you risk loosing data to URE errors. In RAID6 you essentially multiply the URE rate together, which means you have a much much lower error floor and you can repair bigger arrays. In a cost-benefit analysis this means that if your array is small enough, a RAID5 gives you more effective disk space with little additional risk. >If you have node-level redundancy, why bother with the intra-node redundancy? A failover is still a failover and can reduce performance and it reduces your remaining failover margin. You want to keep your failover rate low, though if you doN't particularly care you can use RAID 0 too. >It's not clear to me what limit you're referring to, but, considering how long RAID6 has been implemented, especially in hardware, and how far computing power has increased since then, such a claim is dubious, if not extraordinary. This is not a CPU power limit, rather RAID6 will in the next few years hit the spot where the risk of loosing 2 drives during a rebuild, due to the large drive sizes, becomes large (URE is still low IIRC), so at that point you want more redundancy to keep the rebuild risk low.
- kuwze 8y agoCould you dumb it down? Are you saying that because all the SSDs were from the same manufacturer and installed at the same time their chance of collectively wearing out simultaneously was high?
- ApolloFortyNine 8y agoMost definitely yes. But in addition to this, standard Raid5 does not periodically read the data, so it's actually rather common for issues to only arise on a resilver. This is why proper maintenance in ZFS is to run ZFS Scrub (basically check every file) once a week.
- mmt 8y agoThat's also proper RAID5/6 maintenance. My main/recent familiarity is with LSI hardware RAID implementation, where they call it a "patrol read". I believe mdraid has checkarray. I'm not sure if you meant to imply that ZFS is different from standard RAID in this regard, but it doesn't seem as though it is.
- Siecje 8y agoWe use ZFS, how can I check if we are doing a ZFS Scrub every week?
- simcop2387 8y ago
- true_tuna 8y agoThere’s so much wtf in this. Raid5? No. Also, you have to be sure not to fill the drives up. Creates a pathological wear situation.
- wtallis 8y ago> Also, you have to be sure not to fill the drives up. Creates a pathological wear situation. It's a sliding scale. SSDs intended for use in servers usually have more spare area than client/consumer SSDs, and most of their specifications assume they're full. A lot of enterprise SSDs also include features to allow the user to adjust the usable capacity, and the write endurance rating and warranty will scale to match.
- tinus_hn 8y agoThis happens with spinning disks as well, you have to read all the disks completely from time to time. Many parts of the disk are only used during a sync and that is the only time you don’t want to find out it’s unreadable.