6 ms·
Yes, but as I said above that doesn't make sense. The atomicity we're discussing here is with respect to time, and sequencing (not say size, like an atom). We
by throwaway09223 3y ago
Yes, but as I said above that doesn't make sense. The atomicity we're discussing here is with respect to time, and sequencing (not say size, like an atom).
We say that rename() is atomic because you will either see file A or file B. There are no other possible states.
With a crash, writes might not be committed. You might get an earlier filesystem state.
- aidenn0 3y ago> With a crash, writes might not be committed. You might get an earlier filesystem state. And the whole point of TFA is that this is perfectly fine as long as the metadata writes aren't committed before the data writes. That is: Precondition: File at path A exits with data X 1. Create file at path B 2. Write data Y to file B. 3. Close file B 4. Rename B to A 5. Crash --- As long as the contents in the file at path A are either X or Y, then we have achieve atomicity in the sense used in TFA. This is what we mean by rename acting "atomically across crashes" Note that we only care about the file at path A; all of these are fine: - File at path B exists and has data Y - File at path B exists and has garbage data - File at path B doesn't exist - File at path B exists and is of size 0
- rcxdude 3y agoSpecifically (to hopefully clarify your point), the operation we wish to be atomic is 'replacing the mapping path A->data X with path A->data Y'. This is hoped to be done by placing data Y in path B (which does not need to be atomic )and then using the atomicity of the rename to move around the path. This then puts an ordering constraint that any observer (whether after a crash and looking at the written filesystem, or during the normal operation of the system) sees data Y going into path B before the rename B to A, and this is what POSIX does not guarantee.
- throwaway09223 3y agoYeah, I get it. The issue is that filesystems don't work like you've assumed. I outlined why this idea cannot work here: https://news.ycombinator.com/item?id=38407472 https://news.ycombinator.com/item?id=38407472 Again, manipulating links has nothing to do with file data.
- rcxdude 3y agoIt blatently can work, it's just not guaranteed by the standard (at least without an fsync). I don't think anyone is talking about situations where B is being actively written to while the rename happens.
- throwaway09223 3y agoThis is like saying sync() works. You can sync to durable storage at any point, yes. But you cannot do it in a semantically useful fashion. "I don't think anyone is talking about situations where B is being actively written to while the rename happens." Well, you specifically are ignoring this. I'm not. Filesystems are free to generate extra syncs whenever they like. You could even have a filesystem sync every single write operation to durable storage - why not?
- throwaway09223 3y agoThis is fundamentally not how unix filesystems work. No file data is written in the above scenario with rename(). Rename changes links - not files. Let me rewrite this for you using more correct language: Given a link to a file at path A exists. The file contains data X. There may be other links at paths C, D and E. 1. Acquire a filesystem-global lock on link ops (link/unlink et al, but not read/write). 2. Create a link to X at path B 3. Remove the link at path A 4. Release the filesystem-global lock. Note: The data in X is not relevant. The data in X may be undergoing active modification while the above 4 steps are performed. The data in X may be memory mapped and may change many times during the "atomic" rename() syscall. The atomic properties of rename() are only vis a vis link/unlink semantics within a single filesystem. Other files can't be created or deleted while the rename() is in progress. File data however is under no such restriction -- and in fact it cannot be because the file may be memory-mapped and modified without any syscall interactions. There is no system call sequence point to gate reads and writes. What you are describing is fundamentally incompatible with how the system actually works. What you describe cannot be made to work, ever. It is architecturally invalid at a fundamental level.
- aidenn0 3y agoIt works in ZFS, XFS, and even ext4 with data=ordered. It worked well enough in ext3 as well, that I didn't see issues with it despite crashing the kennel a lot. > File data however is under no such restriction -- and in fact it cannot be because the file may be memory-mapped and modified without any syscall interactions. There is no system call sequence point to gate reads and writes. Note a complete lack of mmap() in the list of operations above; you are trying to trust what I am asking for into something that is impossible, when what I'm asking for is definitely possible. [Edit] I'm also well aware of how rename works; the ask is that the link is not committed before the data. This is possible to do at the expense of some performance, but possibly less performance then using existing commit primitives (e.g. fsync or fdatasync)
- throwaway09223 3y agoThere are all kinds of cases where filesystems might issue an extraneous sync, but it is not atomic, merely ordered. You've switched from talking about one to talking about the other and they are not the same. This is the core of our disagreement. Because it is not an issue of atomic behavior but merely ordering there's no reason to modify the kernel or touch syscalls. It would be far more reliable to provide this in a library -- rename(3) -- which would then automatically provide the benefit on every posix compatible filesystem. It would also allow for a use case of wanting to rename a file without forcing a full data sync -- which could be performance critical behavior in some scenarios. To recap: No filesystems provide atomic syncing of data during rename(). The words don't even make sense. Ext doesn't do it, zfs doesn't do it, xfs doesn't do it. Achieving this would require mechanisms that don't exist. Some filesystems do guarantee that an fsync will occur alongside a rename. There is no specific guarantee as to what file state will be synced. These filesystems do this to mask the impact of folks mis-using the API, paying an efficiency cost as a result. In ext4, for example, the sync does NOT occur within the same transaction as the rename. The sync is NOT atomic (it would be wild if this were the case, as noted above -- architecturally inconsistent and extremely non-performant) It would be ideal for this to happen in libc rather than in the kernel.