4 ms·
This not provide the same integrity as ZFS. Dm-integrity only protects against corruption on the disk itself, while ZFS attempts to provide protection against a
by kroeckx 6y ago
This not provide the same integrity as ZFS. Dm-integrity only protects against corruption on the disk itself, while ZFS attempts to provide protection against all sources of corruption, including software/firmware bugs, RAM and I/O path.
ZFS is using a Merkle tree and so knows which checksum to expect before reading the data. With raidz/mirror, when it detect corruption, it will attempt to read from the other drive(s) and combines those disks that give the correct checksum. If it can't find a combination that works, just what depends on that block is not available. If it can recover the block it will attempt to repair the disk(s) with corruption by writing the correct data to it again.
Standard mdraid can not repair an array in case of corruption since it doesn't know which of the data it has is correct or not, you need to figure out which of the disks has corrupt data on it, and manually remove the drive from the array and hope none of the other disks has corruption. In theory in case of a read error mdraid could attempt to fix it, but it doesn't and just removes the drive from the array. Dm-integrity provides more guarantees that the data you get is correct, but since mdraid and dm-integrity are different layers, they don't actually work together. There is no attempt to repair in case of corruption, dm-integrity will just return an read error and mdraid will remove the drive from the array.
- cmurf 6y agoWhether uncorrectable read error (bad sector) reported by drive firmware, or dm-integrity detecting corruption, the affected LBA's are propagated up to md. And then md can determine the location of a copy (mirror or reconstruct from parity). It then overwrites the bad location. This mechanism is often thwarted with consumer drives, when their SCT ERC timeout can't be set or is longer than the kernel's SCSI command timer default of 30 seconds. Once a command hasn't returned a result of some kind within 30s, the SCSI driver does a link reset. On SATA this has a pernicious effect of clearing the entire command queue, not just the one that was hung up in "deep recovery". LBA's aren't returned so it's indeterminate where or what the problem was caused by, no fix up happens. This results in bad sector accumulation. This misconfiguration is common, and routinely costs people their data. https://raid.wiki.kernel.org/index.php/Timeout_Mismatch https://raid.wiki.kernel.org/index.php/Timeout_Mismatch This may be easier and more reliable to do with a udev rule; but the concept is the same. Also, while this is linux-raid@ wiki, it doesn't only apply to mdadm raid, but LVM, and Btrfs as well. I don't know if it applies to ZoL because I don't know about all the layers ZFS implements itself separate from Linux. But if it depends at all on the SCSI driver for error handling, it would be at risk of this misconfiguration as well.
- kroeckx 6y agoIt seems that a HGST/WD Ultastar, which is their data center drive line, has ERC disabled by default. I've now set it to the suggested 7 seconds.