9 ms·
Optimizing M3: Halving Our Metrics Ingestion Latency by Forking the Go Compiler
- lostmsu 7y agoGreat tech article!
- iampims 7y agoAt scale, details matter. Great read.
- Go0the0gophers 7y agoAt scale, details matter and Go also :)
- shereadsthenews 7y agoI seem to recall there are a couple of github issues around reusing goroutines and their stacks or being able to specify the stack size of a new goroutine instead of making it a runtime constant. Either would be very helpful for those of us using Go at scale.
- nine_k 7y agoWould the segmented stack of the early Go implementations be helpful in such a case?
- richieartoul 7y agoYeah I think this issue would not have cropped up with the early segmented stack implementation, but segmented stacks had their own issues (which is why the Go team migrated away from them)
- jerf 7y agoWhile I won't claim this is unique to Go, I've had some similar good experiences cloning out various bits of Go for my own crazy purposes. The standard library and compiler are relatively clean code for what they are, and it's relatively easy for a pro developer to fork them temporarily like this, or pick up a standard library that almost does what you need and add what you need. I've forked encoding/json to add an attribute that collects "the rest" of the JSON attributes not automatically marshaled, both in and out. I've forked encoding/xml to add a variety of things I needed to write an XML sanitizer (in which you are concerned with things like "how long the attribute tag" is during parsing; it's too late to be presented with a 4 gigabyte attribute in the user code, it's already brought your server to its knees). I saved weeks by being able to start with a solid encoder/decoder backend and be able to follow it and bring it up to snuff, rather than start from scratch. A coworker forked crypto/tls because it was the easiest TLS implementation to break in deliberate ways to test for the various SSL vulnerabilities that have emerged over the years. Of course I recommend this more as a last resort than the first thing you reach for, but it's a fantastic option to have in the arsenal, even if you don't reach for it often. I encourage people to at least consider it.
- derefr 7y agoIt's certainly not unique to Go; this is encouraged (if rarely done) in the Erlang community. In fact, this is the reason that "Erlang/OTP" is a concept distinct from "Erlang." OTP is a distribution of Erlang, a "fork" that remixes together the core Erlang components in a certain way. (Specifically, OTP is Ericsson's distro of Erlang targeted at building telecom switches. That means it contains not just regular language stdlib stuff, but libraries like "Megaco/H.248: a protocol for control of elements in a physically decomposed multimedia gateway".) You're free to make your own distro of Erlang; it's extremely easy (and in fact, everyone is implicity doing it every time they use Erlang's release building process to build an app. An Erlang "release" is a customized, derivative Erlang distribution! You can build a new SDK with the release tooling just as easily as you can build an app runner.) Even if it's rarely taken advantage of, it's clear that "forking your own Erlang distro and customizing it" is the idiomatic approach for many potential user stories. If you study the Erlang "kernel" library, you'll notice that there are many customizations which the kernel "expects" people to make, but not in the sense that there is any stable ABI grip-point to slot in a plugin via a configuration stanza. Rather, Erlang just has a trivial implementation of the logic in the kernel, sitting there in well-factored module. The kernel authors are expressing a clear intent by doing so: if you want some logic more fancy than this trivial version, just fork the kernel and plop in your own version of this module that does something else! Personally, I love this way of thinking. Rather than the usual problems cropping up where something in the runtime that wasn't "quite right" for the application's use-case, was then reimplemented as a non-runtime-integrated library, with other warts and higher overhead as a result; instead, users are encouraged to just solve their problem in the runtime. Sometimes that results in upstreaming a patch, but that's not-at-all the goal. The goal is just to have software that has e.g. exactly one IO path, which everything uses. (Side note: I find it mystifying that Ericsson's OTP release of Erlang is still considered by the community to be "the" Erlang SDK. It's as if Ubuntu was the only Debian-alike and there was no Debian, even though it's very clear exactly what Debian would look like. A "core" Erlang distro—with all the generally-useful OTP stuff like supervisors, but without the telecom-specific stuff—is totally possible; as is rebasing OTP to be a downstream distro of it. But nobody really seems to care about doing so.)
- stubish 7y ago
- curiousDog 7y agoGreat read, so a key takeway would be to make sure to prime the connection pool and also re-use it. Isn't re-using goroutines a bit of an anti-pattern though?
- aflag 7y agoI don't think that's the takeway. Reusing goroutines was only necessary in a very specific situation at a very large scale. The article is very good in providing ideas and tools you can try to use whenever you find yourself in a similar situation.
- andrewfromx 7y agoSummary: developers were calling a method over and over 30 levels deep inside the stack and just barely not going over the 2k golang stack size initial limit. i.e. they were getting great performace because everything happened to be 1.9k or 1.8k, or just not quite 2k or more. Then, a change, and performance went terrible. On the opposite side of 2k with 2.1k or 2.2k an entire extra 2K more had to be allocated for a total of 4k to fit everything. Engineers stop at nothing to find the RCA looking at assembly of the binary and yes, forking the go compiler.
- gen220 7y agoFascinating read. Although the idea of using thread pools evokes the pthread management, this post is rather convincing that such "hand-holding" is necessary in applications with intense SLAs. Alas, the magic the Go team has worked with routines doesn't yield a free lunch for everybody. If we accept that pooling is necessary in some cases, I'm curious – is there a common source that these applications use? In trying to answer my own question, I found that M3 has a mature-looking implementation of such an abstract solution. https://github.com/m3db/m3/tree/master/src/x/sync https://github.com/m3db/m3/tree/master/src/x/sync. Elsewhere, I couldn't find anything similar in the usual suspects. CockroachDB has one-off, specific implementations in the places where they've decided pooling is worth it. Looks like Kubernetes uses the stdlib's `sync.Pool` interface in a similar way, but doesn't use a full-fledged "routine pool". Do people at Uber think this is a robust enough solution to be used outside of m3? Seems like it might be useful in the stdlib as an implementation of `sync.Pool` :)
- richieartoul 7y agoWe use this https://github.com/m3db/m3/blob/master/src/x/sync/pooled_worker_pool.go https://github.com/m3db/m3/blob/master/src/x/sync/pooled_wor... all over our code base so its definitely stable enough to use in your own projects if you have a need, although it does require some tuning. My guess is that the Go team would not consider this critical / core enough to include in the standard library and I'd be inclined to agree with them.
- gen220 7y agoAwesome, thanks for the ref. I'd agree that it probably doesn't belong in the stdlib, as there aren't many programs that would really benefit from it. OTOH, it's good one to keep in pocket, for the few applications that would.
- jadbox 7y agoWhat fix/patch could the Go team do here to mitigate the issue from happening? It does seem like it would be nice if you could specifically to the runtime the amount of pre-alloc stack size needed. Would this help?
- abalone 7y agoAh, memory management. Here’s my basic understanding: - Go initially allocates 2KB stack per routine. When it exceeds it it copies all of it into 2x the space. - This was happening once or twice per request. They didn’t explain exactly why all that stack memory was being used (maybe someone can chime in), but contributing factors were a 30 function deep call stack and a minor code change that tipped it into the next stack growth tier. - Also this doesn’t get freed up until garbage collection runs. - They worked around it by implementing a kind of go routine pool that keeps assigning work to the same (stack-expanded) routines, staying ahead of the garbage collector. My takeaways: 1. Fantastic analysis and job well done. 2. Pooling does not seem to be how things “should” work in Go. It’s more of a hack around undesirable allocator / garbage collector behavior. 3. I’m really interested in reference counted languages like Swift on the server for these reasons. I know ARC means more predictable latency when it comes to garbage collector behavior (which is only indirectly the problem here). Now I’m really curious how Swift allocates the stack and whether it would avoid this “morestack” growth penalty that Go has.
- e12e 7y agoDoes this mean go doesn't do (enough) inlining of functions?
- shereadsthenews 7y agoEveryone agrees this is a problem. Today Go only inlines a function under certain very narrow circumstances.
- GordonS 7y agoIn C# you can apply an `AggressiveInlining` to remove certain limiting restrictions, making it more likely that a method will be inlined - does Go have something similar?
- fwip 7y agoThis has actually improved in the latest release, version 1.12. I'm sure it's still behind other languages, though.
- billsmithaustin 7y agoHopefully Uber won't have to maintain their forked Go compiler for too long.
- richieartoul 7y agoWe don't! We only briefly forked the compiler to prove our RCA, but we resolved the issue with a custom worker pool (linked in the blog post). We compile all our code using the same compiler as everyone else :)
- bdamm 7y agoVery nice! Also another proof point for why it's nice to have open source tools.
- robocat 7y agoExcept it was the second growth just exceeding the 4096 stack size that was causing the issue: "it looked like the goroutine stack was growing from 4 kibibytes to 8 kibibytes"
- kjksf 7y agoIt seems to me they could have used the following hack to fix such issue: var dontOptimizeMeBro byte // go:noinline func makeStackBig() { var buf [16386]byte dontOptimizeMeBro = buf[0] + buf[len(buf)-1] } Call this at the start of the goroutine. What it does, I hope, is extend stack to 16kb, once (as opposed to going from 2kb to 4kb then to 8kb then to 16kb and paying for coyping the memory multiple times). The stack stays big for the remaining lifetime of the goroutine.
- richieartoul 7y agoYep that works too although I didn't benchmark the difference in performance. The nice thing about the worker pool though is that it auto-tunes the stack size based on the workload.
- pcwalton 7y agoThis is a case in which generational GC can help. If you allocate goroutine stacks in the nursery, then you can use a bump allocator, which makes the throughput extremely fast. Throughput of allocation matters just as much as latency does! (By the way, Rust used to cache thread stacks back when it had M:N threading, because we found that situations like this arose a lot.)
- shereadsthenews 7y agoYou can't bump-allocate 6K contiguous to an existing 2K allocation, and Go currently requires the stack to be contiguous.
- pcwalton 7y agoThat's right, but you can bump-allocate a new 6K and copy over. It's a lot faster than falling through to a general-purpose malloc implementation.
- shereadsthenews 7y agoThat’s not obviously true. Page sized and page aligned allocation is pretty fast in Go. Having a generational GC where all pointers are movable could have systemic impact in the performance of the whole program.
- pcwalton 7y agoIf it were fast enough, then caching stacks wouldn't be a win. Bump allocation is like 6 instructions. The idea that generational GC would not be a win does not match the experience of any other language. Generational GC is virtually always a win for languages like Go. This is just another reason why Go should adopt it.
- shereadsthenews 7y agoIt's possible. Go has been replaying the entire history of GC research at about 2.5x forward speed. They should intercept the state of the art eventually.