10 ms·
A data corruption bug in OpenZFS?
- commandersaki 3y agoExcellent writeup robn!
- dannyw 3y agoFascinating write up. As someone with a ZFS system, how can I check if I’m affected?
- moviuro 3y agoIt's a very rare race condition, odds are very low that you were impacted. If you were, you would have noticed (heavy builds with files being moved around where suddenly files are zero). [0] https://bugs.gentoo.org/917224 https://bugs.gentoo.org/917224 [1] https://github.com/openzfs/zfs/issues/15526 https://github.com/openzfs/zfs/issues/15526 (referenced in the article)
- dist-epoch 3y agohttps://github.com/openzfs/zfs/issues/15526#issuecomment-1810800004 https://github.com/openzfs/zfs/issues/15526#issuecomment-181... > zpool get all tank | grep bclone > kc3000 bcloneused 442M > kc3000 bclonesaved 1.42G > kc3000 bcloneratio 4.30x > My understanding is this: If the result is 0 for both bcloneused and bclonesaved then it's safe to say that you don't have silent corruption.
- keep_reading 3y agobclones were only one way to trigger the corruption. This is not a good way to check. It's also not worth checking for because this bug has existed for many years. Your data probably wasn't affected. None of the massive ZFS storage companies out there ran into it by now either. Your data is fine. Sleep easy.
- MenhirMike 3y agoPeriodic reminder to check if your backups are working, and if you can also restore them. It doesn't matter which file system or operating system you use, make sure to backup your stuff. In a way that's immune to ransomware as well, so not just a RAID-1/5/Z or another form of hot/warm storage (RAID is not a backup, it's an uptime/availability mechanism) but cold storage. (I snapshot and tar that snapshot every night, then back it up both on tape and in the cloud.)
- bgro 3y agoIt’s always amazing to me how frequently backups silently fail. Every backup software or general common tool to back things up that I’ve seen has many points of silent failure where it just gives up copying at some point in the process or skips over files for some reason without indicating what or why. If you don’t delete files as you go, now you have an unknown partial backup state that basically doubles your needed space. If you delete as you go, sometimes something happens and the process stops or corrupts so your data is now split and you may have lost something. Even trying to log all the failures during the process is amazingly difficult and solutions to work around that specific problem, themselves, somehow introduce more and new types silent failure in some type of irony.
- MenhirMike 3y agoYes! The worst is that even if you set up all kinds of reports etc. on what you expect, if the backup runs for weeks/months successfully, you just stop paying attention and then when something fails, you won't notice it. I do think that file systems that support snapshots - like ZFS, but I think LVM can be used for stuff like ext4, and Apple APFS does too - is the way to go. Not sure how well NTFS's Shadow Copies/Volume Shadow Service work, I heard horror stories, but not sure if those are one-off freak accidents. Probably worth considering ReFS anyway these days on a Windows Server. But with a Snapshot, you're at least insulating yourself mostly from changes to the data you're backing up. At the expensive of managing snapshots, that is, getting rid of old ones after a while because they keep taking up space.
- 3y ago
- hulitu 3y ago> This whole madness started because someone posted an attempt at a test case for a different issue, and then that test case started failing on versions of OpenZFS that didn’t even have the feature in question. One will expect more seriosity from filesystem maintainers and serious regression testing before a release.
- amelius 3y agoShouldn't we expect formal verification methods, even? Or is that too much to ask for?
- urbandw311er 3y ago[flagged]
- viraptor 3y agoIt's a diff without the author/commiter lines. If you go out of your way to look for the author in the GH link, that's not on the poster.
- xoa 3y agoAs they wrote (and if you're that far into it you really should have been reading) in the paragraph before that, that's the source of the bug in modern OpenZFS but that doesn't mean it was a bug at all in 2006 Sun ZFS. A lot of changes have happened since then, and there may have been something else in ZFS back then or in Solaris itself that meant it wouldn't ever be hit. It's a genuinely interesting tidbit that a little change like that 17 years ago could pop up this way, there is nobody who'd somehow hold Sun (long dead) responsible let alone whichever team did it (even if one dev made the change it should have then needed review/clearance, and if not that's still the problem of Sun not them). Kind of a strange thing for you to zero in on out of a whole interesting article.
- joshxyz 3y agoanyone know what diagram tool did he use? thanks
- mgerdts 3y agoWhen I think of a fs corruption bug, I think of something that causes fsck/scrub to have some work to do, sometimes sending resulting in restore from backups. From the early reports of this, I was having a hard time understanding how it was a corruption bug. This excellent write up clears that up: > Incidentally, that’s why this isn’t “corruption” in the traditional sense (and why a scrub doesn’t find it): no data was lost. cp didn’t read data that was there, and it wrote some zeroes which OpenZFS safely stored.
- cesarb 3y agoIMO, part of the issue is that something which used to be just a low-level optimization (don't store large sequences of zeros) became visible to userspace (SEEK_HOLE and friends). Quoting from this article: "This is allowed; its always safe to say there’s data where there’s a hole, because reading a hole area will always find “zeroes”, which is valid data." But I recall reading elsewhere a discussion about some userspace program which did depend on holes being present in the filesystem as actual holes (visible to SEEK_HOLE and so on) and not as runs of zeros. Combined with the holes being restricted to specific alignments and sizes, this means that the underlying "sequence of fixed-size blocks" implementation is leaking too much over the abstract "stream of bytes" representation we're more used to. Perhaps it might be time to rethink our filesystem abstractions?
- ajross 3y agoIndeed, sparse files are simply a mistake to have included in Unix in the first place (I think we blame this on early SunOS? Not sure, though almost certain that 3BSD and v7 didn't have them). Yes, they have been used productively for various tricks, but they create a bunch of complexity that every filesystem needs to carry along with it. It's a bad trade.
- cogman10 3y agoThis a feature I was completely unaware of. Why would you choose to use a sparse file instead of multiple files?
- vlovich123 3y agoThe number of file descriptors you can have open by a single program is limited and eats up kernel resources.
- ajross 3y agoEven for a torrent client, the number of active file descriptors is a function of the number of peer connections (e.g. a few dozen). It doesn't scale with the size of the output file.
- lupusreal 3y agoIs anybody using bcachefs yet?
- frankjr 3y agoI'm keeping an eye on it but it's not there yet e.g. https://github.com/koverstreet/bcachefs/issues/619#issuecomment-1869606009 https://github.com/koverstreet/bcachefs/issues/619#issuecomm...
- ktm5j 3y agoWell, to be fair they tried and failed to reproduce the corruption that was reported. While I agree that I'm not ready to dive into bcachefs, I'm not exactly swayed by this bug report.
- LanzVonL 3y agoIt's important to note that the recent showstopper bugs have all been in OpenZFS, with the Oracle nee Sun ZFS being unaffected by either.
- nimbius 3y agoOracle laid off basically every Solaris developer in 2017. They are by all observation simply not interersted in the product anymore. its probably the most mournful thing ive seen in tech in a very long time. OpenZFS is a mighty filesystem hobbled by an absolutely detestable license (the CDDL.) Its greatest single contribution was in all likelyhood to BSD, although it didnt seem to make the OS more popular as a whole. the latest and greatest from the OpenZFS crowd seems to be bullying Torvalds semi-annually into considering OpenZFS in Linux...which will never happen thanks to CDDL and so the forums devolve into armchair legal discussions of the true implications of CDDL. You'll see a stable BTRFS and a continued effort to polish XFS/LVM/MDRAID before openZFS ever makes a dent. One could argue OpenZFS is a radioactive byproduct of one of the most lethal forces in open source in the past 20 some years: Oracle. They gobbled up openoffice and MySQL, and went clawing after RedHat just shortly after mindlessly sending Sun to the gallows. Theyre an unmitigated carbunkle on some of the largest corporations in the entire world, surviving solely on perpetual licensing and real-world threat of litigation. That they have a physical product at all in 2023 is a pretty amazing testament to the shambling money-corpse empire of Ellison. Ultimately the FOSS community under Torvalds is on the right track. Just because Shuttleworth thinks he cant be sued by Oracle for including ZFS in Ubuntu with some hastily reasoned shim doesnt mean Oracle wont nonchalantly send his entire company to the graveyard just for trying. Oracle is a balrog. stay as far away as you can.
- rincebrain 3y agoWho on earth is trying to bully Linus into anything? Where have you seen that?
- LanzVonL 3y agoHe did go away for a vacation-style treatment a few years ago after offending Intel. Like a re-education camp.
- frankjr 3y agoI wonder if any large storage provider has been affected by this. I know Hetzner Storage Box and rsync.net both use ZFS under the hood.
- mappu 3y agoWasabi Cloud Storage have a Sponsored-By tag on the git commit fixing the issue, so I assume they're highly involved somehow.