6 ms·
This bug shouldn't really scare people. It's requires such an incredibly specific workload to hit Here's a post by RobN (the dev who wrote the fix) on the ZFS
by AndrewDavis 3y ago
This bug shouldn't really scare people. It's requires such an incredibly specific workload to hit
Here's a post by RobN (the dev who wrote the fix) on the ZFS On Linux mailing list
> There's a really important subtlety that a lot of people are missing in this. The bug is _not_ in reads. If you read data, its there. The bug is that sometimes, asking the filesystem "is there data here?" it says "no" when it should say "yes". This distinction is important, because the vast majority of programs do not ask this - they just read.
> Further, the answer only comes back "no" when it should be "yes" if there has been a write on that part of the file, where there was no data before (so overwriting data will not trip it), at the same moment from another thread, and at a time where the file is being synced out already, which means it had a change in the previous transaction and in this one.
> And then, the gap you have to hit is in the tens of machine instructions.
> This makes it very hard to suggest an actual probability, because this is a sequence and timing of events that basically doesn't happen in real workloads, save for certain kinds of parallel build systems, which combine generated object files into a larger compiled program in very short amounts of time.
> And even _then_, all this supposes that you do all this stuff, and don't then use the destination file, because if you did, you would have noticed that its incomplete.
> So while I would never say that no one has ever hit the problem unknowingly, I feel pretty confident that they haven't. And if you're not sure, ask yourself if you've ever had highly parallel workloads that involve writing and seeking the same files at the same moment.
https://zfsonlinux.topicbox.com/groups/zfs-discuss/Tcf27ae8f8cdd24ca-Md584bcc9d1f07edb6a6e042f https://zfsonlinux.topicbox.com/groups/zfs-discuss/Tcf27ae8f...
Here's another writeup by another ZFS dev Rincebrain https://gist.github.com/rincebrain/e23b4a39aba3fadc04db18574d30dc73 https://gist.github.com/rincebrain/e23b4a39aba3fadc04db18574...
I think the only reason this has gotten so much attention is because it came up as a block cloning bug (which it's not) and that being a new feature created a massive scare that it's widespread. This isn't the first or the last bug ZFS has had - it's software.
- cangeroo 3y ago> ask yourself if you've ever had highly parallel workloads that involve writing and seeking the same files at the same moment. Uhhhh, databases?
- rincebrain 3y agoDatabases don't involve using SEEK_HOLE to find gaps in a sparse file, usually, so it wouldn't come up here.
- thyrsus 3y agoSo you needn't do so: SEEK_HOLE does not occur in the GitHub.com repos for postgresql, mariadb, nor sqlite. Are there system libraries they incorporate which use SEEK_HOLE?
- buildbot 3y agoI feel like there has been kind of a weird concerted effort to push that zfs is bad due to this bug and how trust has been lost etcetera etcetera - super annoying when most other filesystems just corrupt your data and nobody will ever know it happened. I’ve experienced bad data corruption on xfs, btrfs, ext2, and ext4. So far zfs is been nothing but perfect.
- antongribok 3y agoThe reason, and the difference, is that all these other filesystems have check and repair (and sometimes multiple) tools. Please correct me, but ZFS has none.
- gavinhoward 3y agoYou're completely wrong. ZFS's scrub is both a check and a repair tool. It's already saved some of my data.
- hdjdkdbdbe 3y agoYou are partly right. Zfs scrub is a repair tool when it has parity / mirrored copy of data to recreate it.
- danparsonson 3y agoIt's safe to make the assumption that the tool isn't magic.
- Borealid 3y agoA scrub can also repair data when using `ncopies` greater than one even outside any mirrors or parity.
- deleted 3y ago[deleted]
- nightfly 3y ago
- sandreas 3y agoThis is a great explanation, thank you.
- qwertox 3y ago>> So while I would never say that no one has ever hit the problem unknowingly, I feel pretty confident that they haven't. And if you're not sure, ask yourself if you've ever had highly parallel workloads that involve writing and seeking the same files at the same moment. It makes it sound unlikely, but if I have a couple of VMs in datasets (all formatted as ext4 internally and some running DBs inside them) each is one big `raw` file which is getting a lot of reads and writes, I assume. How are they at risk? Also, what about ZVOLs mounted as ext4 drives in these VMs?