5 ms·
No. Here is a good explanation: http://www.zdnet.com/blog/storage/why-raid-5-stops-working-in-2009/162 http://www.zdnet.com/blog/storage/why-raid-5-stops-workin
by simplexion 12y ago
No. Here is a good explanation: http://www.zdnet.com/blog/storage/why-raid-5-stops-working-in-2009/162 http://www.zdnet.com/blog/storage/why-raid-5-stops-working-i...
- CHY872 12y agoI don't believe the contradicts what I've said at all. His claim is that when one disk fails completely, we expect a sector or two on each of the two remaining drives to result in a 'bad sector' error. > So the read fails. And when that happens, you are one unhappy camper. The message "we can't read this RAID volume" travels up the chain of command until an error message is presented on the screen. 12 TB of your carefully protected - you thought! - data is gone. Oh, you didn't back it up to tape? Bummer! So, at this point, we've got two hard drives containing millions of sectors for which one or two are bad, and the software claims that the whole thing is broken, and we have to find a new array? As far as I know, RAID 5 is block level - the block level failure should only destroy one block. All of the others (apart from the other dead ones) are fine. This sort of thing happens in all scenarios with a single disk - eventually an operating system will hit a bad sector which it will have to deal with. In other words, why does the RAID controller crap itself when it can't read (with no recourse at all, according to these articles) when it could just do what every other hard drive does and return 'sector unreadable'. Then the operating system can just remap it etc. I know in some situations one would want to be notified of any miniscule error, but it should be possible to ignore the warnings.
- micro-ram 12y agoI think the assumed issues here are the long rebuild times and the expectation of an additional failure happening during the rebuild which could then trigger another disk error. Then the entire array of data would be taken offline and possibly corrupted. I have build many RAID5/6 arrays and have rarely lost any data, but I am leery and now and tend to just stick with smaller RAID 1 or 10's due to the large size of current disks. We really need native ZFS (BTRFS?) on everything now. Data should be automatically distributed to multiple disks and the file system should be able to guarantee via checksum the data read is what I wrote.
- venus 12y agoAre any consumer NAS that implement ZFS even available? Sounds like ZFS's proactive sector sweeps across all managed drives would handily solve the problem the article raises with conventional RAID.
- lmz 12y agoiXsystems sells the 4-bay FreeNAS mini on Amazon: http://www.ixsystems.com/storage/freenas/ http://www.ixsystems.com/storage/freenas/ That probably counts as "consumer".
- mitchty 12y agoSo I built a zfs raid nas box with 6x3Tb drives. I have it resilver every 2 weeks. So far, NOTHING has failed checksums for almost 2 years. So while the maths behind a 3-4 tb drive returning incorrect data I'm sure is technically correct, I've not seen issues. If you want "proof" i can dump out zpool status/info/log/etc to show i'm not lying. Note the pool is ~50% in use so its not a great example. Also its raidz2 (raid6) so not a direct comparison. I also bought each drive from different lots to hopefully ensure if a drive failed i'd have 2ish days to get a replacement.
- dannyperson 12y agoWhat happens when a URE is encountered and all the disks are online? It seems that they could be detected early and fixed before a rebuild is necessary by doing a weekly sweep of the entire array, reading every data block.
- harshreality 12y agoWith raid 5 if you get a parity block mismatch, you don't know which drive was wrong. You could compare parity bit by bit, to find which bits (it could be anywhere from 1 bit up to the stripe size) was wrong, but you won't be able to figure out how to fix those bit(s) on the stripe without additional information, either a hash or FEC data. Checksumming filesystems let you find the faulty drive by reconstructing data from each n-choose-(n-1) drive set and finding the set with the correct hash. Filesystems using FEC instead of raid (5/z1, 6/z2 ...) can also correct data errors, but I'm not aware of any consumer-level filesystems that implement it. I'm not sure why. Doesn't Amazon use it for S3? Data block and FEC data layout on a disk array has to be a solved problem.