4 ms·
Spoiled web developer here. Since I have no way to test this, I have some questions! It seems like this whole thing could've been simplified to calling exec. Wh
by upghost 2y ago
Spoiled web developer here. Since I have no way to test this, I have some questions! It seems like this whole thing could've been simplified to calling exec. Whats with the whole pipe/fork thing if you are just going to kill one process, clone the other (with exec), and the kill the original. Is the pipe crucial to transferring the contents of memory, or something?
- _jackdk_ 2y agoYes, it's about transferring the contents of memory. If you just call exec(), you get a new process with the file descriptors (connections) open, but you have no idea what the game state (e.g., which connection corresponds to which player) before the exec() call. The fork() is a way of keeping that memory around, and the pipe lets the forked copy send its state to the new code that's restarting.
- skissane 2y agoI don’t see why the fork is strictly necessary, one could also just write it out to a temporary file. That would have the advantage it might be easier to debug if the copyover failed for some reason. Other options (some of which the article briefly alludes to) include POSIX shm_open, Linux memfd_create, Linux O_TMPFILE. So long as you get either an FD or a filesystem path, which can then be passed over the exec in either argv or envp If you use a custom memory allocator with its own heap, you could then store that heap in a shared memory segment and then remap it after the exec. I guess the risk is the memory layout may have changed so you can’t remap it at the same address, in which case you either crash or have some code which rewrites all the pointers in the heap (potentially very painful to do reliably…) Or you could just not allow raw pointers in the objects in that custom heap, requiring them all to be offsets from the heap start
- _jackdk_ 2y agoAll of this sounds pretty reasonable, but my goal was to document the simplest way I saw it work back in the day. You can keep the process alive if copyover fails by writing to a temporary file and designing the server so that you get a "rescue console" if it fails to come up. Then you could inspect the dumped file, see what's going wrong, and exec() into another binary to try again. You could even use sqlite as your state file, which would let you interactively explore and edit it during debugging. A custom allocator sounds like an extremely interesting approach. I'm still not sure whether I'd want to rely on struct layouts not changing, but if you use it diligently you at least make it easier to find everything you've allocated.
- o11c 2y agoYou could use a file instead, but `pipe` is the most portable way to not leave clutter on the system if something goes wrong. If you need traffic in both directions, `socketpair` is useful (and I think portable across platforms that support networking as we know it). Less-portable or less-reliable alternatives include `memfd`, a deleted-but-still-open tempfile, etc. To avoid problems with failure after `exec`, it is very important to fork+exec the server first just to see if it actually starts. If you don't need the PID to remain the same you can actually just use that instance without the other fork or another exec (this still keeps the same PGID). Note that "failure after `exec`" can be fairly sharply divided into "error in C runtime (usually, library compatibility)" and "error in something `main` calls", which require very different approaches. --- Aside: I've had catastrophic failures using the "save after fork" approach (mentioned later in the article) when the child process crashed before the save completed, and the parent just kept trying to have the child save (I also managed to completely crash GDB while investigating this). It is very important for the parent to commit suicide if the child can't do its job, to minimize the amount of lost data. The other thing I would recommend to anyone working with this kind of thing: use systemd and commit all the way; it will make your life 1000% easier than trying to reimplement it badly yourself.
- skissane 2y ago> `socketpair` is useful (and I think portable across platforms that support networking as we know it) Windows is a notable platform not supporting socketpair. Some languages offer socketpair on Windows, e.g. Python, but they are implementing it themselves using TCP loopback connections. Somewhat irrelevant to this discussion given Windows lacks exec (starting a new executable requires a new process which gets a new PID), and in practice lacks fork too (the low-level NT API has undocumented support for forking, but most of the higher-level APIs get confused by it and break when you use it, making it unusable by the vast majority of applications)
- skissane 2y ago> Somewhat irrelevant to this discussion given Windows lacks exec (starting a new executable requires a new process which gets a new PID), Well, you can do exec() on Windows if you implement it yourself in user space, see e.g. https://github.com/polycone/pe-loader https://github.com/polycone/pe-loader (a bit of a dated project, 32-bit only, but the I believe the same principles apply to 64-bit) Cygwin uses a hack – exec() creates a new Windows process, but every Cygwin process has two PIDs, a Windows PID and a Cygwin PID, so exec() changes the Windows PID but keeps the Cygwin PID the same – so it looks like an exec to Cygwin processes, but just like an ordinary spawn to non-Cygwin processes Less insane approach: if you put most of your code in a DLL, you can unload the DLL and then load a new version of it
- ecdavis 2y agoNakedMUD (and probably SocketMUD) uses a temporary file during copyover. https://github.com/avidal/nakedmud/blob/fcab059439515be9ad939f3d70b2794959a6e96f/src/socket.c#L1038 https://github.com/avidal/nakedmud/blob/fcab059439515be9ad93...