6 ms·
Things Unix can do atomically (2010)
- MintPaw 8mo agoNot much apparently, although I didn't know about changing symlinks, that could be very useful.
- 0xbadcafebee 8mo agoYou can use `ln` atomicity for a simple, portable(ish) locking system: https://gist.github.com/pwillis-els/b01b22f1b967a228c31db3cf2789ee13 https://gist.github.com/pwillis-els/b01b22f1b967a228c31db3cf...
- akoboldfrying 8mo agoReally nice explanation of a useful pattern. I was surprised to discover that even the famously broken NFS honours atomicity of hardlink creation.
- michaelcampbell 8mo agoI use `mkdir` for a locking mechanism quite frequently. It also provides me an area to put stuff during the script runs.
- exac 8mo agoSorry, there is zero chance I will ever deploy new code by changing a symlink to point to the new directory.
- sholladay 8mo agoWhy? What do you prefer to do instead?
- gib444 8mo agoAnything less than an entire new k8s cluster and switching over is just amateur hour obviously
- iberator 8mo agowhy? it works and its super clever. Simple command instead some shit written in JS with docker trash
- alpb 8mo agoNobody's saying you should deploy code with this, but symlinks are a very common filesystem locking method.
- slopusila 8mo agothat's how some phone OSes update the system (by having 2 read only fs) that's how Chrome updates itself, but without the symlink part
- x4132 8mo agonot surprised about the chrome part, but pretty shocked at the phone OS part. I know APFS migration was done in this way, but wouldn't storage considerations for this be massive?
- slopusila 8mo agowhat would be more massive would be phones not booting up because of a botched update. this way you can just switch back to the old partition
- marmarama 8mo agoNot really, because only the OS core is swapped in this way. Apps and data live in their own partitions/subvolumes, which are mutable and shared between OS versions. The OS core is deployed as a single unit and is a few GB in size, pretty small when internal storage is into the hundreds of GB.
- dizhn 8mo agoNo snapshotting at all? Thinking about it.. The filesystem does not support it I suppose.
- LiamPowell 8mo agoAndroid does use snapshots: https://source.android.com/docs/core/ota/virtual_ab https://source.android.com/docs/core/ota/virtual_ab
- dizhn 8mo agoOh cool. I was a bit confused about not using snapshots and relying on symlinks but it couldn't be so simple. I guess it's just a simple userspace cow mount. https://source.android.com/docs/core/ota/virtual_ab#compressed-snapshots https://source.android.com/docs/core/ota/virtual_ab#compress...
- bandrami 8mo agoWorks pretty well for Nix
- mananaysiempre 8mo agoAnd for Stow[1] before it, and for its inspiration Depot[2] before even that. It’s an old idea. [1] https://www.gnu.org/software/stow/ https://www.gnu.org/software/stow/ [2] http://ftp.gregor.com/download/dgregor/depot.pdf http://ftp.gregor.com/download/dgregor/depot.pdf
- bandrami 8mo agoI really liked stow. My toy distro back in the day was based on it.
- atmosx 8mo agoWorked pretty well in production systems, serving huge amount of RPS (like ~5-10k/s) running on a LAMP stack monolith in five different geographical regions. Just git branch (one branch per region because of compliance requirements) -> branch creates "tar.gz" with predefined name -> automated system downloads the new "tar.gz", checks release date, revision, etc. -> new symlink & php (serverles!!!) graceful restart and ka-b00m. Rollbacks worked by pointing back to the old dir & restart. Worked like a charm :-)
- gonzus 8mo agoThen you are locking yourself out of a pretty much ironclad (and extremely cost-effective) way of managing such things.
- 1718627440 8mo agoIsn't that the standard way to do that? Why wouldn't you?
- silisili 8mo agoI don't do devops/sysadmin anymore, so this would have been before the age of k8s for everything. But I once interviewed for a company hiring specifically because their deployment process lasted hours, and rollbacks even longer. In the interview when they were describing this problem, I asked why the didn't just put all of the new release in a new dir, and use symlinks to roll forward and backwards as needed. They kind of froze and looked at each other and all had the same 'aha' moment. I ended up not being interested in taking the job, but they still made sure to thank me for the idea which I thought was nice. Not that I'm a genius or anything, it's something I'd done previously for years, and I'm sure I learned it from someone else who'd been doing it for years. It's a very valid deployment mechanism IMO, of course depending on your architecture.
- kjs3 8mo agoClearly, there's a wildly baroque stack of abstractions meters deep that is way better than a simple, reliable, well tested, ubiquitous, idiomatic one line solution. Modern software at it's finest.
- sega_sai 8mo agorename() is certainly the easiest to use for any sort of file-system based synchronization.
- compressedgas 8mo agoAs long as you don't run into or want freedom from possible path races, for that you need the missing: frenameat2(srcdirfd, srcfd, srcname, dstdirfd, dstfd, dstname)
- ta8903 8mo agoNot technically related to atomicity, but I was looking for a way to do arbitrary filesystem operations based on some condition (like adding a file to a directory, and having some operation be performed on it). The usual recommendation for this is to use inotify/watchman, but something about it seems clunky to me. I want to write a virtual filesystem, where you pass it a trigger condition and a function, and it applies the function to all files based on the trigger condition. Does something like this exist?
- Brian_K_White 8mo agoincron
- ta8903 8mo agoThanks, I didn't find this when I was looking for a solution for my problem. This is pretty much the exact solution for my usecase, though for some reason inotify feels more complicated than some kind of filesystem mount solution for me.
- direwolf20 8mo agoare you asking for if statements? if(condition) {do the thing;}
- ta8903 8mo agoI know this is trivial to do programmatically, but I was looking for a way this will be handled by the filesystem. For instance, if I have some processes generating log files, and I have a script that converts them to html, I wanted the script to be called every time a log file is updated, without having a daemon running in the background to monitor the directory, just some filesystem mount. This would have made some deployments easier.
- laz 8mo agoSounds half baked. What context does this function run in? Is it an interpreted language or an executable that you provide? Inotify is the way to shovel these events out of the kernel, then userspace process rules apply. It's maybe not elegant from your pov, but it's simple.
- zzo38computer 8mo agoEven though it can do some things atomically, it only does with one file at a time, and race conditions are still possible because it only does one operation at a time (even if you are only need one file). Some of these are helpful anyways, such as O_EXCL, but it is still only one thing at a time which can cause problems in some cases. What else it does not do is a transaction with multiple objects. That is why, I would design a operating system, that you can do a transaction with multiple objects.
- ptx 8mo agoWindows had APIs for this sort of thing added in Vista, but they're now deprecating it "due to its complexity and various nuances which developers need to consider": https://learn.microsoft.com/en-us/windows/win32/fileio/about-transactional-ntfs https://learn.microsoft.com/en-us/windows/win32/fileio/about...
- akoboldfrying 8mo agoI don't follow, sorry. Are you saying that if we run: mv a b mv c d We could observe a state where a and d exist? I would find such "out of order execution" shocking. If that's not what you're saying, could you give an example of something you want to be able to do but can't?
- jstimpfle 8mo agoI don't think that's happening in practice, but 1) it may not be specified and 2) What you say could well be the persisted state after a machine crash or power loss. In particular if those files live in different directories. You can remedy 2) by doing fsync() on the parent directory in between. I just asked ChatGPT which directory you need to fsync. It says it's both, the source and the target directory. Which "makes sense" and simplifies implementations, but it means the rename operation is atomic only at runtime, not if there's a crash in between. It think you might end up with 0 or 2 entries after a crash if you're unlucky. If that's true, then for safety maybe one should never rename across directories, but instead do a coordinated link(source, target), fsync(target_dir), unlink(source), fsync(source_dir)
- amstan 8mo agoMissing (probably because of the date of the article): `mv --exchange` aka renameat2+RENAME_EXCHANGE. It atomically swaps 2 file paths.
- oguz-ismail2 8mo agoTitle says Unix, renameat2 is Linux-only.
- jasode 8mo ago>Title says Unix, You're misinterpreting the title. The author didn't intend "Unix" to literally mean only the official AT&T/TheOpenGroup UNIX® System to the exclusion of Linux. The first sentence of "UNIX-like" makes that clear : >This is a catalog of things UNIX-like/POSIX-compliant operating systems can do atomically, Further down, he then mentions some Linux specifics : >fcntl(fd, F_GETLK, &lock), fcntl(fd, F_SETLK, &lock), and fcntl(fd, F_SETLKW, &lock) . [...] There is a “mandatory locking” mode but Linux’s implementation is unreliable as it’s subject to a race condition.
- monibious 8mo agoBut I also don't think the auther meant Things you can do in Linux but not Unix
- deleted 8mo ago[deleted]
- deleted 8mo ago[deleted]
- jasode 8mo ago>But I also don't think the auther meant Things you can do in Linux but not Unix I wasn't claiming that. I just thought the ggp had a useful comment about renameat2() which led to gp's "correction" which wasn't 100% accurate. IBM z/OS UNIX also has renameat2(). It doesn't have the Linux specific flag RENAME_EXCHANGE. https://www.ibm.com/docs/en/zos/3.1.0?topic=functions-renameat2-change-name-location-file https://www.ibm.com/docs/en/zos/3.1.0?topic=functions-rename...
- maximgeorge 8mo ago[dead]
- klempner 8mo agoThis document being from 2010 is, of course, missing the C11/C++11 atomics that replaced the need for compiler intrinsics or non portable inline asm when "operating on virtual memory". With that said, at least for C and C++, the behavior of (std::)atomic when dealing with interprocess interactions is slightly outside the scope of the standard, but in practice (and at least recommended by the C++ standard) (atomic_)is_lock_free() atomics are generally usable between processes.
- senderista 8mo agoThat's right, atomic operations work just fine for memory shared between processes. I have worked on a commercial product that used this everywhere.
- Igrom 8mo ago>fcntl(fd, F_GETLK, &lock), fcntl(fd, F_SETLK, &lock), and fcntl(fd, F_SETLKW, &lock) There's also `flock`, the CLI utility in util-linux, that allows using flocks in shell scripts.
- cachius 8mo agoWhat are flocks in this context? Surely not a number of sheep...
- pjmlp 8mo agoIn UNIX/POSIX file locks are advisory, not enforced, it only works if all processes play ball.
- zbentley 8mo agoSure, but the discussion is around whether they’re atomic, not whether they’re advisory.
- zbentley 8mo agoAren’t flock and POSIX locks backed by totally different systems?
- zbentley 8mo agoToo late to edit, but it appears that they are per this comment on a different article and the documentation it references: https://news.ycombinator.com/item?id=46607265 https://news.ycombinator.com/item?id=46607265
- andrewstuart 8mo agoAnywhere there is atomic capability you can build a queuing application.
- ncruces 8mo agoI use several of these to implement alternative SQLite locking protocols. POSIX file locking semantics really are broken beyond repair: https://news.ycombinator.com/item?id=46542247 https://news.ycombinator.com/item?id=46542247
- pjmlp 8mo agoUnless they can be guaranteed by the POSIX specification, they are implementation specific and should not be relied upon for portable code.
- kccqzy 8mo agoWhich of these are not guaranteed by the POSIX specification? It’s been a while since I studied it, but if I recall correctly the ones mentioned in the article are guaranteed.
- jeffbee 8mo agoI wonder why the author left out atomic writes with O_APPEND.
- zbentley 8mo agoUnsure. Aren’t there filesystems which make O_APPEND less durable than it’s specified to be, which might be interpreted to adversely affect atomicity? Could that be it?
- ozgrakkurt 8mo agoThis requires O_SYNC and O_DIRECT afaik. Even then it is only some file systems that guarantee it and even then file size updating isn’t atomic afaik. Not so sure about file size update being atomic in this case but fairly sure about the rest. Matklad had some writing or video about this. Also there is a tool called ALICE and authors of that tool have a white paper about this subject. Also there was a blog post about how badger database fixed some issues around this problem.
- jeffbee 8mo agoI don't think any part of your post is right. Aside from NFS, there should not be filesystems where this doesn't work. If there are, those are just bugs. The flags you mentioned are not required or relevant. Setting the fd offset to the end of the file atomically is the entire purpose of O_APPEND.
- ozgrakkurt 8mo agoIt depends on what you mean by atomic. If it is only writing to page cache and you are writing a small amount then yes? If there is a failure like a crash or power outage etc. then it doesn’t work like that. You might as well be pushing into an in-memory data structure and writing to disk at program exit in terms of reliability
- jeffbee 8mo agoYou are projecting imaginary features onto O_APPEND and then hypothesizing that your imaginary features might not work. POSIX says that for a file opened with O_APPEND "the file offset shall be set to the end of the file prior to each write." That's it. That's all it does.
- KevinChasse 8mo agoNice catalog. One subtle thing I’ve found in building deterministic, stateless systems is that atomic filesystem and memory operations are the only way to safely compute or persist secrets without locks. Combining rename/link/O_EXCL patterns with ephemeral in-memory buffers ensures that sensitive data is never partially written to disk, which reduces race conditions and side-channel exposure in multi-process workflows.
- nialv7 8mo agoThe mmap/msync one is incorrect I believe? (Correct me if I am wrong). msync() sync content in memory back to _disk_. But multiple processes mapping the same file always see the same content (barring memory consistency, caching, etc.) already. Unless the file is mapped with MAP_PRIVATE.
- DSMan195276 8mo agoYeah I agree that one isn't very clear, perhaps the idea is to use `msync()` as a barrier to achieve consistent ordering of the writes without having to handle that yourself with more complex primitives. But then, they do mention some of those primitives at the bottom of the article, so it's hard to say what exactly the idea is.
- icedchai 8mo agommap/msync is behavior is also very platform specific. On some systems (like AIX, at least older versions), even without msync, memory mapped data is synced back to disk periodically. I worked on a code base that was portable between Linux, AIX, and some other Unix flavors. mmap/msync was a source of bugs. Just imagine your system running for days, never syncing any data to disk... then someone pulls the plug. Where'd my data go? Even worse, it happened "in production" at a beta site. Fortunately we had a way to recover data from a log.
- deleted 8mo ago[deleted]