3 ms·
No, it is not. ZFS scrub only checks that the checksums is valid. A scrub will not check in detail for what this post calls "structural error" (it only will fin
by diegocg 5y ago
No, it is not. ZFS scrub only checks that the checksums is valid. A scrub will not check in detail for what this post calls "structural error" (it only will find errors when doing normal walk of structures, which is not very detailed - doing a detailed file system checking is so resource intensive that some fsck implementations have a "low memory" mode in order to avoid OOM).
The checksum/mirroring mechanisms cannot fix any structural error when the filesystem is doing something and finds and inconsistency.
ZFS chose to not have a fsck out of pure arrogance, not because scrub is a proper substitute. ZFS developers believed that corruption bugs produced by code can be fixed by providing Bug Free Code (tm). That, and the fact that errors due to media corruption will be fixed with checksums and mirroring, made them believe that they could make fsck a thing of the past. Other modern file systems mimicking ZFS are developing a fsck, despite having scrub-like functionality.
...as I said, and your reply proves again:
> It's surprising the large amount of people I have found who are incapable of conceiving the notion of what this post calls "structural error" in modern file systems such as ZFS
People _really_ wants to believe ZFS has some kind of magic.
- XorNot 5y agoYour post though isn't providing a single example of a "structural error" which could occur in ZFS and wouldn't be fixed by the checksums and redundancy data. What's your point here? That ZFS can't correct logic bugs in it's own implementation that would lead to structural errors? (What system could?)
- Jenda_ 5y agoI don't have experience with ZFS, but in btrfs, you can get "corrupt leaf" on a FS passing scrub. This probably happens either because of bug or because the structures got corrupted (e.g. bad RAM) when they were constructed, before the checksum (of the already wrong data) was calculated. This is kind of expected behavior, as the checksums only assert that we have read from the device what was written to it previously, but they have absolutely no relation to constrains like "the blocks form a valid B-tree". Another example that I have here right now, is a directory that says the following on "ls" # l pg_stat_tmp/ ls: cannot access 'pg_stat_tmp/global.stat': No such file or directory [...] -????????? ? ? ? ? ? db_0.stat The btrfsck says stuff like "parent transid verify failed" and "ERROR: child eb corrupted". scrub finishes without errors. > That ZFS can't correct logic bugs in it's own implementation that would lead to structural errors? Again, I can't speak for ZFS, but the problem with btrfs is that for example in ext4 you have fsck that will fix such errors (sometimes losing the affected files). But in btrfs, the fsck is mostly "beta and do not use and it can't fix that, just move the data elsewhere, create the FS from scratch and move data back and hope it won't happen again".