3 ms·
This is fundamentally not how unix filesystems work. No file data is written in the above scenario with rename(). Rename changes links - not files. Let me rewri
by throwaway09223 3y ago
This is fundamentally not how unix filesystems work. No file data is written in the above scenario with rename(). Rename changes links - not files. Let me rewrite this for you using more correct language:
Given a link to a file at path A exists. The file contains data X. There may be other links at paths C, D and E.
1. Acquire a filesystem-global lock on link ops (link/unlink et al, but not read/write).
2. Create a link to X at path B
3. Remove the link at path A
4. Release the filesystem-global lock.
Note: The data in X is not relevant. The data in X may be undergoing active modification while the above 4 steps are performed. The data in X may be memory mapped and may change many times during the "atomic" rename() syscall.
The atomic properties of rename() are only vis a vis link/unlink semantics within a single filesystem. Other files can't be created or deleted while the rename() is in progress. File data however is under no such restriction -- and in fact it cannot be because the file may be memory-mapped and modified without any syscall interactions. There is no system call sequence point to gate reads and writes.
What you are describing is fundamentally incompatible with how the system actually works. What you describe cannot be made to work, ever. It is architecturally invalid at a fundamental level.
- aidenn0 3y agoIt works in ZFS, XFS, and even ext4 with data=ordered. It worked well enough in ext3 as well, that I didn't see issues with it despite crashing the kennel a lot. > File data however is under no such restriction -- and in fact it cannot be because the file may be memory-mapped and modified without any syscall interactions. There is no system call sequence point to gate reads and writes. Note a complete lack of mmap() in the list of operations above; you are trying to trust what I am asking for into something that is impossible, when what I'm asking for is definitely possible. [Edit] I'm also well aware of how rename works; the ask is that the link is not committed before the data. This is possible to do at the expense of some performance, but possibly less performance then using existing commit primitives (e.g. fsync or fdatasync)
- throwaway09223 3y agoThere are all kinds of cases where filesystems might issue an extraneous sync, but it is not atomic, merely ordered. You've switched from talking about one to talking about the other and they are not the same. This is the core of our disagreement. Because it is not an issue of atomic behavior but merely ordering there's no reason to modify the kernel or touch syscalls. It would be far more reliable to provide this in a library -- rename(3) -- which would then automatically provide the benefit on every posix compatible filesystem. It would also allow for a use case of wanting to rename a file without forcing a full data sync -- which could be performance critical behavior in some scenarios. To recap: No filesystems provide atomic syncing of data during rename(). The words don't even make sense. Ext doesn't do it, zfs doesn't do it, xfs doesn't do it. Achieving this would require mechanisms that don't exist. Some filesystems do guarantee that an fsync will occur alongside a rename. There is no specific guarantee as to what file state will be synced. These filesystems do this to mask the impact of folks mis-using the API, paying an efficiency cost as a result. In ext4, for example, the sync does NOT occur within the same transaction as the rename. The sync is NOT atomic (it would be wild if this were the case, as noted above -- architecturally inconsistent and extremely non-performant) It would be ideal for this to happen in libc rather than in the kernel.
- aidenn0 3y ago> No filesystems provide atomic syncing of data. Yes everyone in the discussion agrees on this fact > ...but it is not atomic, merely ordered. You've switched from talking about one to talking about the other and they are not the same If the operations we are talking about are ordered, then this provides the atomicity that TFA is asking for. Properly ordering operations is a very common way of making a series of steps atomic.
- throwaway09223 3y ago> "If the operations we are talking about are ordered, then this provides the atomicity that TFA is asking for." No, it doesn't. I think we're reaching the heart of the disagreement here - around what "atomic" means. Atomic means indivisible. It means that other things cannot happen at the same time - that the individual (possibly ordered) steps of an operation cannot be observed while in progress. This is why it's absolutely nonsense to talk about atomicity across a crash. The fundamental design of unix files and memory preclude this. It can't be done.