3 ms·
It's "silent corruption" so a scrub would not detect it. This is the worst possible scenario which is why this is getting a lot of publicity in contrast to bugs
by PaulCarrack 3y ago
It's "silent corruption" so a scrub would not detect it. This is the worst possible scenario which is why this is getting a lot of publicity in contrast to bugs in the past.
What's worse is that the bug has been in production for 18 months (since 2.1.4) so if you've ever put a file on a ZFS filesystem in the last year or so, you may be affected.
The only true way of knowing whether you are are victim of this is to look at every single file and compare it against the known good that has never been on a ZFS filesystem.
- gavinhoward 3y agoOh...this is bad. I have a filesystem that is not ZFS, but it is a backup of a ZFS filesystem. Am I screwed?
- csdvrx 3y ago> Oh...this is bad. Yes, it's extremely bad, and the title of the original submission from 2 days ago may not have caught your attention. It should have been "If your version of ZFS is less than 18 months old, you may have silent data corruption". I editorialized the git title as little as I could while still describing the essential problem. > I have a filesystem that is not ZFS, but it is a backup of a ZFS filesystem. Am I screwed? If it's a backcup of a ZFS filesystem created with a version of OpenZFS more recent than 18 months, maybe. That's because even if it's a 100% perfect backup, you can't know if the backup contains files that where silently corrupted when they were on the original ZFS filesystem (unless you have copies of the files before they were on the ZFS) The bug is deemed "unlikely" from 2.1.4 to the version 2.2 which introduced block cloning and increated the probability of the bug showing up, but you won't know if 0% or 0.1% (or any other proportion) of your files are affected until after you compare checksums. I'd suggest to wait until more is known. If this bug flew under the radar for so long, it should be rare.
- gavinhoward 3y agoUnfortunately, I'm on Gentoo and tend to update regularly. And within the last two weeks, I destroyed many of my oldest snapshots. So yeah, I'm screwed. I guess I should just hope for the best.
- csdvrx 3y ago> And within the last two weeks, I destroyed many of my oldest snapshots. The goal of this repost is to save some trouble to people who may also delete old snapshots if they don't understand how they can be impacted by this bug as the original title was very unclear
- gavinhoward 3y agoYes, that was a good idea. Hey, are there any tools to search for all zero blocks? Maybe that might help detect corruption in this case.
- csdvrx 3y agoLast time I checked, the discussion about how to best detect affected files was still going on. IIRC having a null-byte prefix may be caused by the special scripts to artificially and forcefully trigger the bug. In naturally corrupted files, the position of null bytes may be more random than that. New tests reveal the IO load may matters a lot: if you kept your IO low, you may have as few as 0.01% corrupted files. All this brings me back to my original idea: the best solution may be using metadata from old backups (ctime, checksum) and looking for discontinuities (same ctime, different checksum) like an IDS would, as checking for null bytes would require too much calibration. This method will require acquiring this metadata from old backups to figure which file got corrupted when, and what's the best backup to restore it from