15 ms·
Moving beyond fork() + exec()
- ComputerGuru 4mo agoI'm not surprised Chen's patch was rejected; that's an extremely niche usecase not worth supporting. With my shell developer hat on, I agree with the closing "developers would likely welcome a native implementation that isn't (unlike the current implementation) hiding fork() and exec() under the covers".
- smj-edison 4mo agoIt sounds like they're interested in the concept though, just not that specific implementation.
- sanderjd 4mo agoYeah this seems like a promising discussion.
- Chu4eeno 4mo agoIt has been for decades at this point. thiago's blog posts which introduced me to the topic over a decade ago (and is still one of the best explainers) points out that posix_spawn was introduced in POSIX.1-2001: https://web.archive.org/web/20120718152158/http://www.macieira.org/blog/2012/07/forkfd-part-1-launching-processes-on-unix/ https://web.archive.org/web/20120718152158/http://www.maciei...
- sanderjd 4mo agoYeah fair enough.
- hparadiz 4mo agoMaybe tangentially related but I always think it's silly that every linux process has the same libgcc_so.so.1 loaded into memory for each process even though the raw binary for the library is exactly the same so you end up with like 800 copies of libgcc_so.so.1 in memory. I mean maybe this has been optimized for already and I don't know what I'm talking about but maybe someone with more knowledge about the kernel knows? Is this something we simply can't optimize for because of security implications?
- 201984 4mo agoShared libraries (and mmapped files in general) are deduplicated; it's nowhere near as bad as you think. The kernel loads a .so into memory once and then maps that memory into every process that mmaps it. Editing to add: this deduplication is one of the greatest upsides to dynamic linking. Common libs like libgcc and libc only have to exist in memory once and can stay in CPU caches, whereas if they were statically linked into every binary, each binary would have a copy of that library that wouldn't be shared with anything else and you'd waste a lot of memory.
- sjmulder 4mo agoDoesn't the loaded code have to be patched for relocations?
- ptspts 4mo agoIt does, so not 100% is reused. The patched parts are in different sections though, so the entire .text (code) section ends up being reused.
- monocasa 4mo agoNot on modern archs that provide decent support for PIE (position independent executables).
- 201984 4mo agoHow do you think position independent code can call functions from other .so's without being patched with their addresses? They can't, so even PIC code still has to have a relocation table that gets patched. It's in a different page than the code though, so code does still get reused.
- monocasa 4mo agoThat's not really patching though, any more than any use of function pointers is patching.
- Sophira 4mo agoI'm guessing that a big part of the problem with moving away from fork() in general is that each new process needs a copy of the parent process' environment anyway, right?
- lokar 4mo agothe environment is not that big
- dijit 4mo agoI'm a bit naive, but I don't think that's necessarily a requirement. It might be commonly held convention, and thus, an assumption, in Linux (and, broadly, UNIX) but I don't think it's true inside VAX or even Windows, so I don't think it's a requirement. Unless I've missed something (which is totally possible, this is not an area of OS design I've spent much time).
- zerobees 4mo agoThe LWN article is incorrect in saying that it "must copy the entire process state (including memory) for the child process". There are some kernel structures and page tables that need to be initialized, plus you need a new stack, but it's not nearly as dramatic as implied. Most of the parent's memory is "incorporated by reference", so to speak. In fact, if you profile it, in the fork() + execve() model, execve() is far more expensive, because not only does it replace the old process with a new one, but it also involves running the dynamic linker, which opens, parses, and mmaps library files. It still makes sense to get rid of the fork() overhead if you're going to throw away the cloned process state soon thereafter, but if you wanted to make process execution radically faster, rethinking the exec architecture would probably offer more significant gains.
- ktpsns 4mo agoThere is lots of discussion on this old API here on hacker news, for instance https://news.ycombinator.com/item?id=31739794 https://news.ycombinator.com/item?id=31739794
- deleted 4mo ago[deleted]
- lokar 4mo agoThis seems unnecessary to me. In the example, the core of git should be a library yo can link so you don't need to run the binary. That would be better in every way.
- sanderjd 4mo agoThere are lots of reasons to want to spawn fresh processes, which aren't solved by linking a library.
- aerzen 4mo agoSpawning processes should not be on the hot path of any program.
- 1718627440 4mo agoWhy? That's a very useful processing primitive.
- lokar 4mo agoIt’s a hack with many disadvantages. Sometimes a hack is the right answer, but the kernel should it add a primitive for it.
- MBCook 4mo agoShould bash link in every program the user might want? Load them up as dynamic libraries?
- m132 4mo agoNode, Python, PowerShell, and the rest do (almost) just that. launchd and systemd famously strived to remove as much shell from the start up process as possible because it was harming boot times and introducing unpredictability.
- sanderjd 4mo agoI just ran into this recently, where I had an obscure bug caused by needing to close more file descriptors in the forked process. "I want a clone of the current process" is just way less common in my experience than "I want a completely new process". It feels crazy that we don't have a way to directly express the latter thing, and can only approximate it by cloning and then fixing things up in post.
- dnw 4mo agoWhat do you mean by "a completely new process"?
- sanderjd 4mo agoA process that shares nothing with the process that spawned it.
- jerf 4mo agoA thing that makes that complicated is that while you want that conceptually, you don't want that in reality. For instance, if the spawning process is in a container of some sort and it spawned a process that "shares nothing with the process that spawned it", the spawned process would no longer be in that container, because the state of "being in the container" is one of the things it shares with the parent process. This is just an example of I don't even know how many things a modern-day process will share from its parent. By "complicated" I do not even remotely mean "unsolvable". I just mean that if you really dig down into what it means to "share nothing" in a modern operating system, it's a lot richer than it was back when fork+exec was a practical solution. There's a lot of fuzzy things that could go either way when you say "shares nothing".
- dcrazy 4mo agoIt’s such a bad idea that every OS except Linux implements it? On macOS it’s posix_spawn, on Windows it’s NtCreateProcess.
- debatem1 4mo agoThere are a lot of slightly different fork-exec-like things in the concept space and it's hard to imagine one approach satisfying them all. IMO it would be interesting to take an approach analogous-ish to sched_ext_ops where you built the rough flow chart of a combined fork-exec, but with hooks built to enable ebpf to change behavior or skip the bits these sophisticated users don't want/need.
- MBCook 4mo agoFork/exec is great if you actually want the traditional copy of your process for some reason. For launching something totally new, like the example in the article of some tool calling git, I think it does make a ton of sense to make something new. Especially since I suspect that is by far the more common case. I suspect “I want a clone of me“ is relatively rarely used at this point.
- debatem1 4mo agoRelatively rarely, but in some performance sensitive use cases. Mine happens to be fuzzers, where a very cheap fork-like primitive would be a really big win.
- surajrmal 4mo agoAndroid and chrome both benefit greatly from fork exec as part of their zygote model iirc. It substantially reduces the memory cost and latency of spawning new apps and tabs.
- uecker 4mo agoThe elegance of the fork() + exec() model is that every kind of configuration can be done after the fork using all the usual APIs. Every attempt to replace it with a combined call that I have seen so far seemed fundamentally poorer because it needs to add all configuration options as parameters to the call and then do this in away that you can extend it later and does not become a mess.
- amluto 4mo agoI have the entirely opposite opinion. IMO a big mistake of the UNIXy model is that so much state is preserved across the creation of a process. For example, there are APIs to have a specific thing be fd number 4 so you can run a program and have it find that thing at fd 4. This is weird. Windows, for all its many, many faults, did not use fork+exec and instead mostly has options for how one creates a process. It wasn’t done elegantly, but it was the right decision.
- 1718627440 4mo agoIs it weirder, that you can pass an variable precisely into argument 4? You do need to pass information to a subprocess and there needs to be some agreement on what means what. Sure, maybe you could use names instead of fds, but that sounds needlessly complicated.
- jonhohle 4mo agoThat’s like saying you could use positions to specify function argument access (as in assembly) instead of variable names. File descriptors being numbers that are likely array indexes in a file handle seems like a leaky abstraction. Having a namespace that a parent process share with its children seems like a much cleaner design.
- amluto 4mo agoA way to pass a defined list of handles to a subprocess (or a friendly other process) makes sense. Having that mechanism be direct inheritance of those handles with the same numbering as the source is obnoxious.
- mrkeen 4mo ago> fork() is a relatively expensive system call; it must copy the entire process state (including memory) for the child process. Many optimizations have been made over the years, but a fork is still a fundamentally costly operation. To make things worse, a fork() call is often immediately followed by an exec(), which will discard all of that memory that was so carefully copied for the child. It's weird to leave out a mention of copy-on-write - the optimisation that means that you don't copy over all the memory.
- FooBarWidget 4mo agoIt says state. Copy on write still means it's O(number of page table entries) even if you don't copy the contents. It's a well known issue that forking a program with large virtual memory size is slow.
- mort96 4mo agoIt says "(including memory)". It's pretty natural to read this as "(including the contents of allocated pages)".
- m00x 4mo agoOn modern hardware a cow page copy should only take 1-5ms. Redis forks to save the db to disk and it's been a solid design choice. I guess it depends on how sensitive your application is to main thread pauses.
- rom1v 4mo agoRelated to the discussion: "A fork() in the road": https://www.microsoft.com/en-us/research/wp-content/uploads/2019/04/fork-hotos19.pdf https://www.microsoft.com/en-us/research/wp-content/uploads/... > ABSTRACT > The received wisdom suggests that Unix’s unusual combination of fork() and exec() for process creation was an inspired design. In this paper, we argue that fork was a clever hack for machines and programs of the 1970s that has long outlived its usefulness and is now a liability. We catalog the ways in which fork is a terrible abstraction for the modern programmer to use, describe how it compromises OS implementations, and propose alternatives. > As the designers and implementers of operating systems, we should acknowledge that fork’s continued existence as a first-class OS primitive holds back systems research, and deprecate it. As educators, we should teach fork as a historical artifact, and not the first process creation mechanism students encounter.
- pizlonator 4mo agoFork is marvelous for the zygote pattern Hard to come up with an optimization that is equally efficient and elegant
- toast0 4mo agoThe zygote pattern[1] is a great optimization to deal with the cost of forking, but IMHO, being able to inexpensively spawn a carefully tailored process regardless of the size and scope of the current process would be better. I would guess it would be a small difference in measurable performance between zygote and a direct clean spawn, but it's one less trick an application needs to do, and it would be very helpful for libraries that spawn things. Spawning inside a library isn't always a great thing to do, but some things would really benefit from process level isolation. [1] In case one isn't aware, the zygote pattern involves forking a 'zygote' process during application startup, and having that process do any forks that need to happen during application runtime. This reduces the cost of forking in large applications, because the zygote will have few fds open and use little memory. This lets your large application spawn new processes without delaying the application or the startup of the new processes. Some applications will spawn many zygotes to allow parallelism for spawning at runtime.
- burnt-resistor 4mo ago> "If you are repeatedly creating large processes, you are already doing it wrong. The fix is in user space, not the kernel." Every couple of years, someone claims they have "the solution" implying everyone else who came before them didn't know what they were doing.
- yxhuvud 4mo agoIt can also mean that neither the hardware side or the software side is static, but change over time. That means that their demands and what they allow also change over time. This leads to the insight that what was perhaps a good idea on 70s hardware/software is not necessarily a good, or even ok, idea 50 years later on modern hardware executing OSes and programs that have been kept up to date.
- mike_hock 4mo agoThe most astonishing part is that this is dated June 5th, 2026. I.e. a year that starts with 20, not 19.
- JdeBP 4mo agoThese discussions were definitely had back in the 20th century too. The spawn model versus the fork+execve model has been an on-going debate since the time of MS/PC/DR-DOS.
- jcalvinowens 4mo agoIt is a weirdly common misconception that that fork() is cheap... it is O(N) on the size of the process, and it always has been. Yes, it's copy on write... but there is a linear relationship between the size of the process and the number of page table entries required to represent it.
- themafia 4mo ago> the number of page table entries This is not exactly fixed since you can vary the amount of memory each page maps with things like hugepages and the same process can run with different page sizes.
- IshKebab 4mo agoYou can in theory, but it's rare in practice because it isn't always enabled and it requires root to configure.
- Panzerschrek 4mo agoThe whole approach of using fork seems to be unnatural for me. In many cases (even in the majority of them) it's not needed to inherit the whole structure of the parent process, but to start a given executable. Windows does this better with its CreateProcessW interface.
- ajkjk 4mo agoFork always seemed conceptually terrible even when I first learned about it.. If you want to do one thing (start a process) you should not have to use a mysterious incantation that does a different unrelated thing (forks your process) in order to do it. I am curious about what the best way to handle the example in the article of one process spawning many git subprocesses is. Surely it just doesn't make sense to repeatedly start git from scratch in the course of a long-running parent operation. What's the low cost abstraction for the same result, though?
- wmf 4mo agolibgit2 exists. You could imagine communicating with some gitd over a pipe/socket but I don't know why that would be a good idea. Short of that you have to spawn processes.
- trumpdong 4mo agoOn Windows maybe it would be a COM server, using IPC built into the OS. The client sees it like a local function call.
- spacechild1 4mo agoYeah, as someone who originally came from Windows, the fork+exec model never made sense to me. Now I know it's just a historical quirk, but for some reason there are still people who pretend that fork+exec is actually a good thing...
- kps 4mo agoFork is conceptually simple. Without bringing in any other layers, you start a process with the one thing known to exist: yourself. Otherwise you need multiple steps to create a process, fill it with something to run, and arrange for it to execute. Or like Win32 you permanently smush them together with other layers, like filesystems and object loaders and linkers.
- Too 4mo agoFill with what stuff exactly? The only thing I want to inherit from the parent process is its cwd and environment variables, even those are often overridden. The rest can easily be passed explicitly through other channels like pipes or command line arguments. Back to the example from the article. It makes no sense that a git-subprocess forked from a web server need to have any process state inherited from the web server.
- ggm 4mo agoAesthetically I have no intention of moving beyond. I'm content with my kernels scheduler and how it maps "heavyweight" processes to cores. I do use threaded code. It's significantly harder to write and reason about. (45 years in to a CS career, ageing out) You have to be clever to do better than clever people. Clever people bootstrapped me into fork()/exec() and I know my limits.
- redleader55 4mo agoWhen cores start needing more than 9 bits to be represented and RAM is in terabytes, many of the old assumptions need to change. Schedulers need to be implemented in userspace, RAM needs to be allocated in GB, not in 4k, io needs to require less round-trips between kernel and user space and NICs need to do a lot more work before the data reaches the CPU.
- skydhash 4mo agoDoes it need to be the same OS? Most consumer device are in the low 16GB range for memory with some outliers in the 64 and 128 GB. 32 cores are still in the realm of specialized devices. Yes, we’re not the one paying for Linux development, but its subsystems are so complicated for general purpose computing. Like fitting formula 1 car parts onto a camry.
- tadfisher 4mo agoOur software is littered with the consequences of these kinds of assumptions, and they have an impact on consumer use cases. x86 still runs in real mode on boot despite dropping the PC BIOS. Lots of software still assumes a 4kb page size, to the point where migrating Android to 16kb is an ongoing multi-year effort involving far too many people. And this is an OS for phones, which you might assume would lack the memory to benefit from a larger page size. And one of the most popular consumer CPUs for enthusiasts, the Ryzen X3D chips, broke assumptions in both Linux and Windows schedulers that all cores have access to the same amount of L3 cache. I would probably not assume the kinds of hardware limitations that we have now will persist into the useful lifetime of current software. Splitting the OS into "consumer" and "enterprise" variants is one of those moves that would bake in a ton of assumptions and make things messier in the future.
- a-dub 4mo agoi thought this was all fixed with special modes of clone that are optimized and don't actually copy anything (ie, it creates a new deficient process that can pretty much only exec)?
- zbentley 4mo agoKind of. Those exist, but because Linux’s formal ABI is syscalls and not libraries that combine them in known-safe ways, the clone speedups that make fork faster are a confusing and fragile API for low-level programmers to use. That, and even those clone-without-pagetable-copy improvements leave a lot of slowness on the table. Being able to skip even disable-able functionality intended for fork would simplify code. Also, for programs that launch the same subprocess many times, a better API might allow caching away some of the pre-entrypoint initialization of exec.
- asveikau 4mo agoThe things you can do between fork and exec are sometimes underestimated. Off the top of my head, you can call dup2(), you can set a process group id, probably a few other things. If you contrast that with win32, where you optionally pack a bunch of initial values into a struct, win32 is a much more narrow, less pleasant, less freeform interface, where it is harder to introduce more features. But I think there is already posix_spawn to imitate that philosophy on Unix-like OSs.
- loeg 4mo ago> The things you can do between fork and exec are sometimes underestimated. Off the top of my head, you can call dup2(), you can set a process group id, probably a few other things. What do you mean underestimated? You can do anything between fork and exec; there are no limitations.
- dcrazy 4mo agoThat’s not true. man 7 signal-safety
- loeg 4mo agoYou're talking about libc design choices, not constraints imposed by the kernel. To the kernel, a post-fork pre-exec process is just any old process. GP was suggesting post-fork processes were constrained in the syscalls they could invoke; they are not.
- deleted 4mo ago[deleted]
- asveikau 4mo agoI did not say they are constrained in what syscalls they can make, as if some nanny at the syscall entry point will punish you for doing wrong. I said that it interacts poorly with threads due to inherent race conditions. See the other comment.
- codedokode 4mo agoThe problem with replacing exec/fork is that you usually want to configure new process: for example, set up signal handlers, close or open FDs, switch namespaces, setup seccomp, adjust permissions. And all the system calls to do it apply only to the current process and you need something to replace them. The proposal in the article was to create a new API for this. My idea is that we could make a new syscall, for example "spawn", that creates a new empty process, loads some lightweight "loader" into it, and passes arbitrary configuration data. The loader configures the process and exec()'s the main program. This allows to avoid forking the memory and keep existing APIs, but still requires to fork file descriptors and other things.
- nyrikki 4mo agoLuckily someone with a time machine saw your post and added it to POSIX.1-2001 :) (Sorry if you weren't joking) but yes, posix_spawn() has been a thing and in glibc fork is just a alias to clone() Not exactly that OP idea, but fork/exec is legacy really.
- MayCXC 4mo agoother people in the thread say that posix_spawn is more or less implemented as a fork+exec wrapper though? it sounds like the idea is more like if there were a separate deferred_fork that made an intermediate "process factory" that let you set up a process without actually creating a new one until the exec. obviously the if() construct would have to be replaced with an in-process handle that mimics calls to the posix api.
- trumpdong 4mo agoI liked the other proposal where you can create a blank process and then force it to make syscalls, ending with execve. That doesn't require a bunch of special data structures to hold the syscalls you want to do.
- stevefan1999 4mo agoIf fork and exec can exhibit persistent and algebraic behavior (beyond its CoW nature) that would not only be more useful but more interesting to use, for example using it for doing lazy evaluation
- LoganDark 4mo agoHuh, LWN has moved to (sometimes) requiring a click to proceed past the subscription pitch to the actual article. I feel like this may have an inverse effect (insistent begging to the point of inserting additional obstacles = angry/insulted users that are less likely to pay).
- corbet 4mo agoIt's an experiment. Compared to the text-obscuring popovers that are prevalent elsewhere on the net, it seems pretty low-key; as far as I know, this is the first complaint I've seen. I don't know if we will continue experimenting with those or not...better ideas for getting people to subscribe to the site would be more than welcome.
- foo-bar-baz529 4mo agoThis isn’t moving beyond fork and exec at all. It’s adding a complicated API for a marginal gain for a niche use case, and ignoring the actual big bottleneck of fork
- sathyayoshi 4mo ago[flagged]
- mpweiher 4mo agoI've always liked the Mach approach. You've got a few primitives: - address space - memory objects - threads Mix and match. A Task (process) is not a primitive, but a composite object combining address space with one or more threads. How you fill the address space with actual memory objects is up to you. Map from disk or COW your own address space...have fun! https://developer.apple.com/library/archive/documentation/Darwin/Conceptual/KernelProgramming/Mach/Mach.html https://developer.apple.com/library/archive/documentation/Da...
- tus666 4mo agoHow can you write for LWN and not have heard of clone(CLONE_THREAD) and multithreading?
- medoc 4mo agoRelated: executing small commands from the Recoll indexer: https://www.recoll.org/pages/idxthreads/forkingRecoll.html https://www.recoll.org/pages/idxthreads/forkingRecoll.html
- high_byte 4mo agoall this for 2%