4 ms·
Yeah the RAID1 part is weak sauce. However, my specific setup mitigates these flaws: * The intention is to detect at-rest bit flips before they progressively p
by pslam 11y ago
Yeah the RAID1 part is weak sauce. However, my specific setup mitigates these flaws:
* The intention is to detect at-rest bit flips before they progressively pollute everywhere. I don't mind if it's not instantaneously detected on-access. I perform a nightly full scan - so there's up to 24 hours where an at-rest bit-flip may lie undetected, but it won't progress past that.
* I use ECC at all cache hierarchy levels possible, and ECC DDR, with active background scrub. So bit flips here won't occur (with vanishingly small probability).
* I use software-RAID1 only. The path between DDR, thru CPU, and to storage controller is unlikely to have a bit flip. From there onwards, the data is essentially written twice, so there is vanishingly low probability of both having a bit-flip, except for systematic failure for that bit pattern in two attempts to two different disks and two different controllers.
So there's still some places in the stack where errors can be introduced, but the most common areas of fault are either duplicated or covered by detection mechanisms.
My goal is to detect, but not correct, transport and at-rest bit flips. I'll discard everything when an error is detected, and use backups.
For my next setup I'll probably switch to a filesystem with a better end-to-end error detection story, such as ZFS, and ditch RAID altogether.
- cmurf 11y agoHow do you detect bit flips with software RAID? A scrub can detect differences, but without checksums it's ambiguous which copy is wrong. And in the case of mdadm, which has no concept of a file system, it won't report the affected file. I'm not even sure you get an LBA, but if you do, then you have to go look that up, accounting for offsets, with the chosen file system to get a file. Conversely, Btrfs scrubs will report a corrupt file path if there's a bit flip in a single copy. If there's another copy available then there's just a kernel note that there was a data csum error and the problem was fixed, no file name path.
- pslam 11y agoThe entire point of my system is that I don't need to care which copy is good. If it scrubs and detects any error, I'll mount disks independently as two filesystems, and diff the contents. I'll then restore corrupted data from backup. Yes, a checksumming filesystem would be better, but at the time of creation (and even now?) none of the filesystem choices were mature, or proven to be sufficiently reliable. Ext4 and software RAID1 are both mature and proven reliable.