3 ms·
This article is just plain wrong. The author doesn't seem to understand how filesystems work - specifically the difference between writing data and renaming dir
by throwaway09223 3y ago
This article is just plain wrong. The author doesn't seem to understand how filesystems work - specifically the difference between writing data and renaming directory entries. LWN should either issue a correction or take it down.
First: The author is confusing atomicity techniques with durability techniques. rename() isn't at all relevant to durability. Never has been.
Second: Nothing about using rename() is relevant to what happens when pages are lost on crash. Using rename() or link() has no bearing on dirty writes. This whole section should just be deleted.
Third: When pages are lost, they're not back-filled with random unallocated space (and if they were, see the second point above). Filesystems are designed to ensure these regions are zeroed out before the area can be used - either proactively or in fsck.
Finally, this quote: "Which brings us to the present day fsync/rename/O_PONIES controversy, in which many file systems developers argue that applications should explicitly call fsync() before renaming a file if they want the file's data to be on disk before the rename takes effect"
There is no such controversy because rename() is utterly unrelated to the durability of data on a disk. This is a pure hallucination.
What a bizarre article.
- throwaway09223 3y agoI read a bit more and figured out why the article was written. In 2009, ext4 (and ONLY ext4) adopted this behavior: https://git.kernel.org/pub/scm/linux/kernel/git/tytso/ext4.git/commit/?id=dbc85aa9f11d8c13c15527d43a3def8d7beffdc8 https://git.kernel.org/pub/scm/linux/kernel/git/tytso/ext4.g... I don't think any other unix filesystems do this. It's not posix, and it's not behavior that any apps should be depending on if they care about their data. I would suggest updating the LWN article to clarify that this is a piece specifically about ext4 -- not POSIX, or even Linux in general.
- nh2 3y agoWhat you say is correct, and ext4's current default behaviour can be quite confusing: * ext4 turns rename() into an invisible fsync(), if the target file existed. * ext4 turns close() into an invisible fsync(), if the opened file existed. This leads to surprising results: If you unzip something the first time, it's nice and fast. If you unzip the same file again, force-overwriting existing files, suddenly it's 10x slower. You now spend hours finding out why, and start cargo-culting vodoo that doing `rm` before makes it inexplicably fast again. Invisible fsync()s that the application programmer cannot turn off are quite bad. My speculation is that Ted Ts'o agrees with that in principle, but that he got tired of having to explain people that if they don't fsync(), there are no guarantees; so he added this hack that makes 95% of the cases go away with surprising special-case behaviour.
- nh2 3y agoFor people who want to read more about this in _up-to-date_ documentation: * `auto_da_alloc` section of https://www.kernel.org/doc/html/latest/admin-guide/ext4.html https://www.kernel.org/doc/html/latest/admin-guide/ext4.html * https://en.wikipedia.org/wiki/Ext4#Delayed_allocation_and_potential_data_loss https://en.wikipedia.org/wiki/Ext4#Delayed_allocation_and_po...
- hedora 3y agoThat’s not quite true. There used to be a performance bug in ext3 that caused it to order the writes before the journal flush that created the files. Then, gnome(?) and only gnome(?) relied on it, leading to massive filesystem corruption, so the kernel team gave up and made the old braindead behavior the default. (The current default behavior is braindead because it makes rename extremely slow for correct programs that don’t care about durability across crashes, and therefore don’t call fsync).
- throwaway09223 3y agoYeah, all filesystems are allowed to synchronize to disk as often as they want, in whatever way they want in /addition/ to what's specified. Filesystems can even make their own guarantees, above and beyond generalized specifications like fsync(). I could write a filesystem that synchronized files after every 42 megabytes, or once every 69 seconds. Apps might even come to depend on this behavior. All filesystems will have predictable, implementation dependent idiosyncrasies. But these aren't specified behaviors. They're not part of an interface and they shouldn't be depended on. The article implying that they ought to be is, I think, shortsighted. Other filesystems don't do this and your code will break if you assume it. I agree it's braindead.
- lxgr 3y agoFrom the article: > However, the ordering effect of rename() turns out to be a file system specific implementation side effect. It only works when changes to the file data in the file system are ordered with respect to changes in the file system metadata. In ext3/4 [...] Seems pretty clear to me that this is talking about an ext3/4 implementation detail that people have started to rely on...? I also pretty clearly remember that from back in 2009. > This article is just plain wrong. Articles like this helped me understand the distinction between theoretical POSIX semantics and what Linux was actually doing at the time. It seems like you are reading more into this article than at least I got out of it – maybe it's because I still remember that 2009 controversy, but I wouldn't have drawn any of the (incorrect, and I agree on that!) conclusions you're listing above.
- fao_ 3y ago> The author doesn't seem to understand how filesystems work At the time of writing (2009) the author had 10 years of experience under Sun Microsystems, IBM, Intel, and Red Hat -- all working on filesystems. Including ZFS, ext2/3, and ChunkFS. (It's literally on her Wikipedia page) So I'm more likely to regard her comments in high regard versus a driveby post written by a throwaway account.
- throwaway09223 3y agoIt's pretty common for people with some experience to not understand filesystem nuance - there's quite a bit of it right here in the hn comment section. You'll note below I also corrected someone significantly more credentialed than the author of this LWN post. My suggestion to you: Worry less about measuring resumes and more about who is actually correct as a matter of demonstrable fact. If you have actionable questions, ask them. Use facts, not appeals to (frankly not very substantial) authorities.
- nh2 3y agoThe throwaway acccount is right. In POSIX, file content data and metadata (directory entries) are separate. Atomicity and durability are separate. It is likely that the LWN post author understands this very well, but omitted this important info from the article. -- I can also understand throwaway's criticism, e.g. on sections like this: > Given this situation, application developers came to rely on what is, on the face of it, a completely reasonable assumption: rename() of one file over another will either result in the contents of the old file, or the contents of the new file as of the time of the rename(). No. There is nothing reasonable about this. This is programming. When you're programming against a spec (POSIX), you don't "make assumptions". You rely only on what it says in the spec. It would be "reasonable" to wish that somebody creates spec that ensures the mentioned rename semantics. But assuming that a spec says something it doesn't is wishful thinking. People are quick to lazy it out, assume, copy-paste, etc, instead of critically thinking "wait, does the code I write here really guarantee the desired effect, e.g. to write my file to disk". Reject assumption-based programming. Go check. Read the docs. That makes good programs. Or, as the throwaway says, rely on "demonstrable fact".
- jcalvinowens 3y ago> The author is confusing atomicity techniques with durability techniques. No, the author isn't confusing anything. The author is describing a controversy in the linux community, and presenting the arguments being made in that controversy for you, the reader, to evaluate. It's called journalism. I happen to agree with you that it was a sort of silly controversy... but the controversy was very real. > There is no such controversy. This is a pure hallucination. What a bizarre article. There factually was controversy on the mailing list. It's history, it happened. That's what the article is about.