4 ms·
MDADM RAID1 does not seem to have the ability to recover from silent data corruption. In the standard two-disk RAID1, it's actually impossible to recover, since
by zootboy 5y ago
MDADM RAID1 does not seem to have the ability to recover from silent data corruption. In the standard two-disk RAID1, it's actually impossible to recover, since it doesn't know what copy is the uncorrupted one. With three disk RAID1 it would theoretically be possible to use a majority vote, but I don't believe MDADM has such code.
Btrfs raid1, on the other hand, uses the checksums of the data stored in its metadata trees to validate which copy of the data is correct and repair the corrupted copy. 3-disk makes this even more robust, as you get an exponential growth of data-metadata pairs to pull from. If any one pair matches, that data can be assumed to be correct.
- butlerm 5y agoNot sure how well this works out in practice, but if one drive returns a read error for a sector and the other returns the corresponding sector without an error, it should be a simple matter to trust the latter. Of course that assumes that storage corruption occurred after uncorrupted data was written to disk. Modern drives have rather large per physical sector checksums to detect the occurrence of these errors and presumably retry a couple times before returning an I/O error if configured correctly. In addition a good I/O bus should have error detection checksums for data in transit, and good hardware should have ECC RAM as well.
- zootboy 5y agoIndeed. That's why I specified /silent/ corruption. If the stack returns an explicit error, that is correctable (assuming the other disks don't error and return the correct data). Sadly, a large number of SSDs (and nearly all cheap USB thumb drives / SD cards / eMMC devices) will happily return silently corrupted data.