9 ms·
How a fix in Go 1.9 sped up our Gitaly service by 30x
- jsiepkes 9y ago"Recompiling with Go 1.9 solved the problem, thanks to the switch to posix_spawn" I never understood why so many people use fork() instead of POSIX spawn(). For example OpenJDK (Java) also does this as the default for starting a process. Which leads to interesting results when you use it on a OS which does do memory over committing like Solaris. Since the process briefly doubles in memory use with fork() your process will die with an out of memory error.
- jsiepkes 9y agoTypo: "does do" should read "doesn't do".
- dboreham 9y agoBecause decades of written material about Unix says fork() is really cool (even though it isn't)?
- f2f 9y agofork was alright before other people tacked on multiples of cruft like threads and whatnot onto commercial unixes and they became mainstream. the current problem is that you don't want to have to copy all file descriptors if all you're going to do is call "exec" and reduce them to three: in, out, err. for example, here's the caveats section from the macOS fork man page: There are limits to what you can do in the child process. To be totally safe you should restrict your- self to only executing async-signal safe operations until such time as one of the exec functions is called. All APIs, including global data symbols, in any framework or library should be assumed to be unsafe after a fork() unless explicitly documented to be safe or async-signal safe. If you need to use these frameworks in the child process, you must exec. In this situation it is reasonable to exec your- self. That spells defeat :) Earlier in the game, copy-on-write had to be created for the same reasons.
- cstrahan 9y ago> the current problem is that you don't want to have to copy all file descriptors if all you're going to do is call "exec" and reduce them to three: in, out, err. To be clear, exec does not necessarily close all but the first three fds -- by default all fds will be inherited. However, you can set the close-on-exec flag on each individual fd (in fact, that's what the Go stdlib does behind the scenes). Search for FD_CLOEXEC in fcntl(2) and open(2) and you'll see what I'm referring to. http://man7.org/linux/man-pages/man2/fcntl.2.html http://man7.org/linux/man-pages/man2/fcntl.2.html http://man7.org/linux/man-pages/man2/open.2.html http://man7.org/linux/man-pages/man2/open.2.html
- khc 9y agobecause posix_spawn() in linux often calls fork(). I just looked at the manpage now and it says under some conditions it'd call vfork() instead, but I don't remember that being the case when I last looked at this (6-7 years ago?)
- noselasd 9y agoNowadays, posix_spawn() calls clone() on linux, with the CLONE_VM flag, behaving much like vfork() as far as I can tell. That means the child and parent process shares the memory (until exec() is performed). Especially if the parent process is multi-threaded this avoids a whole lot of pagefaults that would occur if using fork() when another thread touches memory, possibly triggering a lot of copy-on-writes in the time window between calling fork() and the child calling exec() Code: https://sourceware.org/git/?p=glibc.git;a=blob;f=sysdeps/unix/sysv/linux/spawni.c;h=6b699a46dd404aa39d49349d09a191909e3e93ab;hb=65f6c94e682e3c21bde4c3aea74011c00d41ac04#l352 https://sourceware.org/git/?p=glibc.git;a=blob;f=sysdeps/uni...
- ploxiln 9y agofork() is a pretty simple way to be able to modify the environment for a process you will spawn. fork(), the child can modify its own environment using various orthogonal system calls, e.g. to redirect stdout/stderr or drop permissions, and then exec the target executable. Threads throw a wrench in things. But fork() existed for decades before threads. O_CLOEXEC etc helps. Lots of command-line utilities don't use threads. fork() isn't the fastest way - but in many situations it's not a problem, it's just convenient. In that respect it's somewhat like using python when you could have used go.
- jacobvosmaer 9y agoAnother nice example is changing the working directory for the new process. With fork+exec, you can do a chdir after fork but before exec. With posix_spawn you're stuck with the working directory of the parent.
- nitwit005 9y agoI thought of creating a fix myself way back, and the issue was that Go made use of system calls directly. You basically have to re-implement posix_spawn in Go. If you look at their change, it includes updates to chipset specific files, and the fix only seems to work on a CPU that reports as amd64.
- ithkuil 9y agoI must say I didn't go to look at the sources of the patch, but what you say sounds so odd that I'll take the chance and suggest that perhaps the fact that in golang "amd64" is, for historical reasons, the name of the architecture more neutrally known as "x86_64", is the source of confusion (I.e. it doesn't just work on AMD or on CPUs that claim/report having a specific model/maker etc). Low level syscall ABI is architecture dependent.
- linkregister 9y agoamd64 is the original name of the instruction set. Intel did beat AMD to a 64-bit instruction set: that of the Itanium processors, IA-64. Itanium had performance issues and lots of errata. Most importantly, IA-64 was not natively backwards-compatible with x86 instructions. amd64 became the standard. x86_64 is a common name for the amd64 architecture, and is a way to describe both the AMD and Intel implementations. In my opinion, amd64 is a less ambiguous name and is more historically accurate. https://en.m.wikipedia.org/wiki/X86-64 https://en.m.wikipedia.org/wiki/X86-64 Yes, I am aware that my point is undercut by the fact that the article title is x86-64, but I stand by my statement.
- cesarb 9y agoIt's not just the article title. Follow the two footnotes in the "History" section of that article, to the press releases from AMD announcing the new ISA. They consistently call it "AMD x86-64" or "AMD's x86-64" or just "x86-64". The oldest snapshot I could find of the x86-64 web site (https://web.archive.org/web/20000817014037/http://www.x86-64.org:80/ https://web.archive.org/web/20000817014037/http://www.x86-64...) also calls it x86-64. The most recent snapshot of that site, however, calls it AMD64; it seems to have changed sometime in the middle of April 2003. That is, both x86-64 and AMD64 are historically accurate (2003 was early enough in the ISA's lifetime), but x86-64 is the earlier name.
- revelation 9y agoBecause every straightforward way of running an external command on Unix involves fork(). So someone wrote that API not thinking much of it. Then shock horror they realize running a throwaway command is fork()ing the main process. But now everyone is too angsty to change it because someone out there might rely on the environment copy functionality, even when they shouldn't.
- kevincox 9y agoIt's often to remember the other point of view when you see huge performance differences. > A bug in Go <1.9 was causing a 30x slowdown in our Gitaly service.
- 0003 9y agoThe author does say fix.
- deleted 9y ago[deleted]
- Animats 9y agoEach Gitaly server instance was fork/exec'ing Git processes about 20 times per second so we seemed to finally have a very promising lead. What's really wrong here is that they're apparently spawning processes like crazy. Do they spawn a new process for each API call? That's like running CGI programs under Apache, like it's 1995.
- dchest 9y agoNothing wrong with it. In fact, I wish spawning processes was more common. It's beneficial for security.
- mschuster91 9y ago> It's beneficial for security ... and for RAM usage. Java applications all have a tendency to bloat the longer you keep them running.
- paulfurtado 9y agoIt's not bloat or memory leaks per se, the JVM just does not return memory to the OS after it is freed. To limit its memory usage, tune the heap size. To fully allocate the heap on startup for consistent usage, use -XX:+AlwaysPreTouch
- axaxs 9y agoJava has always been 'use memory to increase speed...sometimes'. You can tune it some, sure, but that's what it's known to do.
- chinhodado 9y agoLast time I checked, I still couldn't control the max heap free ratio, because apparently that option/flag just doesn't work with Java 8's default GC.
- nurettin 9y agoBack when I tried running a jruby application on 800mb ram, it bloated, then started throwing "OutOfMemoryError"s and "Insufficient Class Space" or something similar. Apparently jruby was generating too many new types at runtime to accommodate rails framework. Garbage collector was pretty garbage at it's job back in 2011.
- empath75 9y agoRelated post from a few weeks ago: Fork is not my favorite syscall: https://news.ycombinator.com/item?id=16068305 https://news.ycombinator.com/item?id=16068305
- jorangreef 9y agoWe currently have the same problem in Node, where fork is still being called synchronously from the event loop instead of asynchronously from the thread pool. Calling exec() or spawn() in Node is therefore not asynchronous and can block your event loop for hundreds of milliseconds or even seconds as RSS increases. https://github.com/nodejs/node/issues/14917 https://github.com/nodejs/node/issues/14917
- alecthomas 9y agoThat looked like a very frustrating exchange.
- zebra9978 9y agoAre you guys planning to migrate gitlab to golang ? I think the biggest feature that everyone wants is better performance. Is the migration path that tough ?
- romanovcode 9y agoThey have so many features that I don't see it happening ever.
- carussell 9y agoThey wouldn't need to stop the world and do a full rewrite. It would be feasible if they stop writing new components in Ruby and began replacing the existing parts piecemeal.
- connorshea 9y agoThere are no plans to migrate all of GitLab to Go. The main Rails app is going to stay a Rails app for the foreseeable future. There are a few reasons for this. For one it'd be such a huge project, but also Rails is working well for us, it's great for our pace of feature development. We are working on moving the git layer to Gitaly[0] which is written in Go (and is what this blog post is about). It was one of our major bottlenecks and we've seen a lot of benefit from having made the switch. It's not done yet, but a lot of the calls to git that the application makes are now done through Gitaly. [0]: https://gitlab.com/gitlab-org/gitaly https://gitlab.com/gitlab-org/gitaly
- YTGRK 9y agohttps://youtu.be/I2l8xkjTUh4 https://youtu.be/I2l8xkjTUh4
- stonewhite 9y ago> Having solid application monitoring in place allowed us to detect this issue, and start investigating it, far earlier than we otherwise would have been able to. Yet apparently nobody either caught or investigated the latency spike after the previous deployment.
- tuna 9y agoPretty sure that using stdlib and trying to limit shell script in Go would help performance. Case in point, forking "du": https://gitlab.com/gitlab-org/gitaly/blob/master/internal/service/repository/size.go#L24 https://gitlab.com/gitlab-org/gitaly/blob/master/internal/se...
- haikuginger 9y agoThe article linked to a set of posix-spawn benchmarks[1] that seemed to indicate that while fork/exec time scaled linearly with resident memory on Linux, it did not scale at all on macOS. First of all, I was somewhat confused by that due to the availability of copy-on-write; I wouldn't have expected fork/exec time to scale up that way. Second, I was surprised that there wasn't an attempt to explain the behavior difference between the two systems. Can someone familiar with either or both point towards an explanation for why that's the case? It seems very odd. [1]https://github.com/rtomayko/posix-spawn#benchmarks https://github.com/rtomayko/posix-spawn#benchmarks