5 ms·
I've had ZFS pools survive (at different times over the span of years): - A RAIDz1 (RAID5) pool with a second disk start failing while rebuilding from an earli
by seized 3y ago
I've had ZFS pools survive (at different times over the span of years):
- A RAIDz1 (RAID5) pool with a second disk start failing while rebuilding from an earlier disk failure (data was fine)
- A water to air CPU cooler leaking, CPU overheated and the water shorted and killed the HBA running a pool (data was fine)
- An SFF-8088 cable half plugged in for months, pool would sometimes hiccup, throw off errors, take a while to list files, but worked fine after plugging it in properly (data was fine after)
Then the usual disk failures which are a non-event with ZFS.
- gigatexal 3y agoThis is why I always opt for ZFS.
- ilyt 3y agoI recovered from 3 disk RAID 6 failure (which itself was organization failure driven...) in linux's mdadm... ddrescue to the rescue, I guess I got "lucky" the bad blocks didn't happen in same place on all drives (one died, other started returning bad blocks), but chance for that to happen are infinitesly small in the first place So shrug
- radlad 3y agoYes, I also had a "1 disk failed, second disk failed during rebuild" event (like the parent, not your story) with mdadm & RAID 6 with no issues. People seem to love ZFS but I had no issues running mdadm. I'm now running a ZFS pool and so far it's been more work (and things to learn), requires a lot more RAM, and the benefits are... escaping me.
- avianlyric 3y agoHow do you know you got lucky with corrupted blocks? mdadm doesn’t checksum data, and just trusts the HDD to either return correct data, or an error. But HDDs return incorrect data all the time, their specs even tell you how much incorrect data they’ll return, and for anything over about 8TB you’re basically guaranteed some silent corruption if you read every byte.
- Dylan16807 3y agoThose specs aren't true, or at least the distribution is very far from uniform and they're failing to explain it even a fraction of the way. You don't get corruption that often.
- kalleboo 3y ago> and for anything over about 8TB you’re basically guaranteed some silent corruption if you read every byte This is simply not true. I've been running bi-weekly ZFS scrubs on my file server which has grown from 80 TB to 200 TB of data over three years now with zero failed checksums. That is petabytes of reads with zero corruption. The oldest drives are nearing 5 years old (they were in a Btrfs Synology NAS before, again, zero failed checksums). The wildly pessimistic numbers on the spec sheets are probably just there so they can deny warranty replacements by saying that some errors are in-spec.
- abrookewood 3y agoMy son has a ZFS pool where he just realised one of the drives has half a billion errors on it and the data is still fine. New drive has been ordered!
- unixhero 3y agoPerhaps you ought to run to a Best buy and buy external drives and do backups ASAP. RAID is not backup, always remember that!
- abrookewood 3y agoIt's actually all replaceable data and he's planning to recreate the pool anyway. Just insane that ZS handles it though.