4 ms·
On the last round of zpool scrubs of the 13PB-used set of zfs filesystems here (raidz3, mostly SAS drives, all hosts with ECC RAM), 50MB of data was repaired.
by ilikejam 7y ago
On the last round of zpool scrubs of the 13PB-used set of zfs filesystems here (raidz3, mostly SAS drives, all hosts with ECC RAM), 50MB of data was repaired.
50MB isn't much compared with 13PB, but silent corruption is a real thing.
- anon9001 7y agoThat makes sense, but mdadm also scrubs weekly on a cronjob. Is there something about ZFS that makes corruption more recoverable than mdadm's RAID?
- InvaderFizz 7y agoOne thing is that ZFS tells you exactly what file is corrupted. I recently took over a ZFS setup that was built with plain ZFS on top of a 24-drive wide RAID6. Drives failed, rebuilt the RAID6, nothing appeared to be wrong. ZFS scrub revealed two corrupted files. I was able to pull backups of those two files, replace them on the FS, and rerun the scrub. Obviously I moved off the RAID6 soon after, but this is a powerful example of why the checksuming of ZFS is so useful.
- ilikejam 7y agoRAID5 corruption is non-recoverable - you can't tell which is right, the data or the parity. RAID6 is recoverable - you can get consensus between the data and the two parity bits. ZFS stores hashes of each block, so as long as you have two copies, you can just pick the data+hash that's correct and restore the corrupt block, even with only two copies of the data.
- saltcured 7y agoThe core ZFS feature being references here is storage checksums. The general argument is that as storage volumes have increased in size and/or workload, the conventional checksum methods have not been changed to increase their reliability in a way that prevents real world corruption on real world storage during their real world lifetimes. A RAID system can only reliably recover when an IO error is reported. In this case, it can assume which element of the stored data is no longer trustworthy and use the other parts to reconstruct it. A RAID scrub operation that finds inconsistencies, without a reported IO error, cannot really determine why it is inconsistent. It doesn't know which element is corrupt and which other elements to use to reconstruct it. In effect, ZFS adds another data integrity layer over the physical storage, so you can have a lower probability of silent failures where the disk controller or IO channels do their own "parity" or other checks and report a successful read when they actually returned corrupt values. Something like mdadm or LVM would need to add a new layer of checksums over the physical devices to provide this level of protection. Each block/chunk/stripe written to disk would need another integrity checksum written as well. During operation, these checksums would be verified before deciding that a read operation was actually successful. If the verification fails, a virtual IO error is synthesized to flag the read as bad, so that the other RAID recovery techniques can then be applied while knowing which block/chunk/stripe is invalid.