13 ms·
Go Optimization Guide
- parhamn 2y agoNoticed the object pooling doc, had me wondering: are there any plans to make packages like `sync` generic?
- arccy 2y agoeventually: https://github.com/golang/go/issues/71076 https://github.com/golang/go/issues/71076
- roundup 2y agoAdditionally... - https://go101.org/optimizations/101.html https://go101.org/optimizations/101.html - https://github.com/uber-go/guide https://github.com/uber-go/guide I wish this content existed as a model context protocol (MCP) tool to connect to my IDE along w/ local LLM. After 6 months or switching between different language projects, it's challenging to remember all the important things.
- jigneshdarji91 2y agoAdditionally... - https://www.uber.com/en-AU/blog/how-we-saved-70k-cores-across-30-mission-critical-services/ https://www.uber.com/en-AU/blog/how-we-saved-70k-cores-acros... This has saved Uber a lot of money on compute (I'm one of the devs). If your compute fleet is large and has memory to spare (stateless), performing dynamic GOGC tuning to tradeoff higher memory utilization for fewer GC events will save quite a lot of compute.
- TechDebtDevin 2y agoEmbedding those docs in your MCP server takes about 5 seconds with mcp-go's AddResource method https://github.com/mark3labs/mcp-go/blob/main/examples/everything/main.go https://github.com/mark3labs/mcp-go/blob/main/examples/every...
- deleted 2y ago[deleted]
- nopurpose 2y agoEvery perf guide recommends to minimize allocations to reduce GC times, but if you look at pprof of a Go app, GC mark phase is what takes time, not GC sweep. GC mark always starts with known live roots (goroutine stacks, globals, etc) and traverse references from there colouring every pointer. To minimize GC time it is best to avoid _long living_ allocations. Short lived allocations, those which GC mark phase will never reach, has almost neglible effect on GC times. Allocations of any kind have an effect on triggering GC earlier, but in real apps it is almost hopeless to avoid GC, except for very carefully written programs with no dependenciesm, and if GC happens, then reducing GC mark times gives bigger bang for the buck.
- nurettin 2y agoIs it worth making short lived allocations just to please the GC? You might just end up with too many allocations which will slow things down even more.
- aktau 2y agoIt is not. Please see my answer (https://news.ycombinator.com/item?id=43545500 https://news.ycombinator.com/item?id=43545500).
- MarkMarine 2y agoAre you including in this analysis the amount of time/resources it takes to allocate? GC isn't the only thing you want to minimize for when you're making a high performance system.
- nopurpose 2y agoFrom that perspective it boils down to "do less", which is what any perf guide already includes, allocations is just no different from anything else what app do. My comment is more about "reduce allocations to reduce GC pressure" advice seen everywhere. It doesn't tell the whole story. Short lived allocation doesn't introduce any GC pressure: you'll be hard pressed to see GC sweep phase on pprof without zooming. People take this advice, spend time and energy hunting down allocations, just to see that total GC time remained the same after all that effort, because they were focusing on wrong type of allocations.
- ljm 2y agoYou're not really writing 'Go' anymore when you're optimising it, it's defeating the point of the language as a simple but powerful interface over networked services.
- jrockway 2y agoWhy? You have control over the parts where control yields noticeable savings, and the rest just kind of works with reasonable defaults. Taken to the extreme, Go is still nice even with constraints. For example, tinygo is pretty nice for microcontroller projects. You can say upfront that you don't want GC, and just allocate everything at the start of the program (kind of like how DJB writes C programs) and writing the rest of the program is still a pleasant experience.
- ashf023 2y ago100%. I work in Go and use optimizations like the ones in the article, but only in a small percentage of the code. Go has a nice balance where it's not pessimized by default, and you can just write 99% of code without thinking about these optimizations. But having this control in performance critical parts is huge. Some of this stuff is 10x, not +5%. Also, Go has very good built-in support for CPU and memory profiling which pairs perfectly with this.
- emmelaich 2y agoI think you have a point that there's generic advice for optimising: don't. i.e. Make it simple, then measure, then make it fast if necessary. Perhaps all this is understood for readers of the article.
- mariusor 2y agoI think at least some of the patterns shared in the document, using zero-copy, ordering struct properties are all very idiomatic. Writing code in this manner is writing good Go code.
- Cthulhu_ 2y agoWhat do you mean? If you don't want that level of control over e.g. memory allocation, registries, cache lines etc, there's higher level languages than Go you can pick from, e.g. Java / C# / JS.
- jensneuse 2y agoYou can often fool yourself by using sync.Pool. pprof looks great because no allocs in benchmarks but memory usage goes through the roof. It's important to measure real world benefits, if any, and not just synthetic benchmarks.
- makeworld 2y agoWhy would Pool increase memory usage?
- xyproto 2y agoI guess if you allocate more than you need upfront that it could increase memory usage.
- throwaway127482 2y agoI don't get it. The pool uses weak pointers under the hood right? If you allocate too much up front, the stuff you don't need will get garbage collected. It's no worse than doing the same without a pool, right?
- cplli 2y agoWhat the top commenter probably failed to mention, and jensneuse tried to explain is that sync.Pool makes an assumption that the size cost of pooled items are similar. If you are pooling buffers (eg: []byte) or any other type with backing memory which during use can/will grow beyond their initial capacity, can lead to a scenario where backing arrays which have grown to MB capacities are returned by the pool to be used for a few KB, and the KB buffers are returned to high memory jobs which in turn grow the backing arrays to MB and return to the pool. If that's the case, it's usually better to have non-global pools, pool ranges, drop things after a certain capacity, etc.: https://github.com/golang/go/issues/23199 https://github.com/golang/go/issues/23199 https://github.com/golang/go/blob/7e394a2/src/net/http/h2_bundle.go#L998-L1043 https://github.com/golang/go/blob/7e394a2/src/net/http/h2_bu...
- 2y ago
- kevmo314 2y agoZero-copy is totally underrated. Like the site alludes to, Go's interfaces make it reasonably accessible to write zero-copy code but it still needs some careful crafting. The payoff is great though, I've often been surprised by how much time is spent allocating and shuffling memory around.
- jasonthorsness 2y agoI once built a proxy that translated protocol A to protocol B in Go. In many cases, protocol A and B were just wrappers around long UTF-8 or raw bytes content. For large messages, reading the content into a slice then writing that same slice into the outgoing socket (preceded and followed by slices containing the translated bits from A to B) made a significant improvement in performance vs. copying everything over into a new buffer. Go's network interfaces and slices makes this kind of thing particularly simple - I had to do the same thing in Java and it was a lot more awkward.
- jrockway 2y agoGOMEMLIMIT has saved me a number of times. In containerized production, it's nice, because sometimes jobs are ephemeral and don't even do enough allocations to hit the memory limit, so you don't spend any time in GC. But it's saved me the most times in CI where golangci-lint or govulncheck can't complete without running out of memory on a kind-of-large CI machine. Set GOMEMLIMIT and it eventually completes. (I switched to nogo, though, so at least golangci-lint isn't a problem anymore.)
- nikolayasdf123 2y agonice article. good to see statements backed up by Benchmarks right there
- nikolayasdf123 2y agonicely organised. I feel like this could grow into community driven current state-of-the-art of optimisation tips for Go. just need to allow people edit/comment their input easily (preferably in-place). I see there is github repo, but my bet people would not actively add their input/suggestions/research there, it is hidden too far from the content/website itself
- whalesalad 2y agoFor sure. Feels like the broader dev community could use a generic wiki platform like this, where every language or toolkit can have its own section. Not just for performance/optimization, but also for idiomatic ways to use a language in practice.
- deleted 2y ago[deleted]
- EdwardDiego 2y agoHuh, this surprises me about Golang, didn't realise it was so similar to C with struct alignment. https://goperf.dev/01-common-patterns/fields-alignment/#why-alignment-matters https://goperf.dev/01-common-patterns/fields-alignment/#why-...
- Cthulhu_ 2y agoYup, it's a fairly low-level language intended as a replacement to C/C++ but for modern day systems (networked, concurrent, etc). You don't have manual memory management per se but you still need to decide on heap vs stack and consider the hardware.
- jerf 2y ago"you still need to decide on heap vs stack" No, you can't decide on heap vs stack. Go's compiler decides that. You can get feedback about the decision if you pass the right debug flags, and then based on that you may be able to tickle the optimizer into changing its mind based on code changes you make, but it'll always be an optimization decision subject to change without notice in any future versions of Go, just like any other language where you program to the optimizer. If you need that level of control, Go is generally not the right language. However, I would encourage developers to be sure they need that level of control before taking it, and that's not special pleading for Go but special pleading for the entire class of "languages that are pretty fast but don't offer quite that level of control". There's still a lot of programmers running around with very 200x ideas of performance, even programmers who weren't programmers at the time, who must have picked it up by osmosis. (My favorite example to show 200x perf ideas is paginated APIs where the "pages" are generally chosen from the set {25, 50, 100} for "performance reasons". In 2025, those are terribly, terribly small numbers. Presenting that many results to humans makes sense, but my default size for paginating API calls nowadays is closer to 1000, and that's the bottom end, for relatively expensive things. If I have no reason to think it's expensive, tack another order of magnitude on to my minimum.)
- fmstephe 2y agoJust an anecdote from work to back this up. I wrote a system that was taking requests, making another request to a service (that basically wrapped elasticsearch) and then processed the results and returned to the results to the caller. By default the elastic-search results were paginated and defaulted to some small number in the order of 25..100. I increased this steadily upwards beyond 100,000 to the point where every request always returned the entire result in the first page. And it _transformed_ the performance of the service. From one that was unbearably slow for human users to one that _felt_ instantaneous. I had real perf numbers at the time, but now all I have are the impressions. But the lesson on the impact of the overhead of those paginated calls was important. Obviously everything is specific and YMMV, but this something worth having in the back of your mind.
- _345 2y agoAnyone know of a resource like this but for Python 3?
- asicsp 2y agoThis might help: https://pythonspeed.com/datascience/ https://pythonspeed.com/datascience/
- neillyons 2y agoCurious to know what people are building where you need to optimise like this? eg Struct Field Alignment https://goperf.dev/01-common-patterns/fields-alignment/#avoiding-false-sharing-in-concurrent-workloads https://goperf.dev/01-common-patterns/fields-alignment/#avoi...
- dundarious 2y agoFalse sharing is an absolutely classic Concurrency 101 lesson, nothing remarkable about it.
- kubb 2y agoSomething that shouldn’t be written in a GC language.
- piokoch 2y agoI don't think GC has anything to do here, doing manual memory allocation we might hit the same problem.
- Cthulhu_ 2y agoGC is not relevant in this case, it's about whether you can make structs fit in cache lines and CPU registers. Mechanical sympathy is the googleable phrase. GC is a few layers further away.
- devcoder78 2y ago[dead]
- stouset 2y agoChecking out the first example—object pools—I was initially blown away that this is not only possible but it produces no warnings of any kind: pool := sync.Pool{ New: func() any { return 42 } } a := pool.Get() pool.Put("hello") pool.Put(struct{}{}) b := pool.Get() c := pool.Get() d := pool.Get() fmt.Println(a, b, c, d) Of course, the answer is that this API existed before generics so it just takes and returns `any` (née `interface{}`). It just feels as though golang might be strongly typed in principle, but in practice there are APIs left and rigth that escape out of the type system and lose all of the actual benefits of having it in the first place. Is a type system all that helpful if you have to keep turning it off any time you want to do something even slightly interesting? Also I can't help but notice that there's no API to reset values to some initialized default. Shouldn't there be some sort of (perhaps optional) `Clear` callback that resets values back to a sane default, rather than forcing every caller to remember to do so themselves?
- tgv 2y agoYou never programmed in Go, I assume? Then you have to understand that the type of `pool.Get()` is `any`, the wildcard type in Go. It is a type, and if you want the underlying value, you have to get it out by asserting the correct type. This cannot be solved with generics. There's no way in Java, Rust or C++ to express this either, unless it is a pool for a single type, in which case Go generics indeed could handle that as well. But since Go is backwards compatible, this particular construct has to stay. > Also I can't help but notice that there's no API to reset values to some initialized default. That's what the New function does, isn't it? BTW, the code you posted isn't syntactically correct. It needs a comma on the second line.
- eptcyka 2y agoIs there a time in your career where an object pool absolutely had to contain an unbounded set of types? Any time when you would try know at compile time the total set of types a pool should contain?
- zaphodias 2y agoI assume they're referring to the fact that a Pool can hold different types instead of being a collection of items of only one homogeneous type.
- kunley 2y ago"Although the struct Data contains a [1024]int array, which is 4 KB (assuming int is 4 bytes on the architecture used)" Huh,what? I mean, who uses 32b architecture by default?
- donatj 2y agoUnpopular opinion maybe, but sync.Pool is so sharp, dangerous and leaky that I'd avoid using it unless it's your absolute last option. And even then, maybe consider a second server first.
- infogulch 2y agoA new sync/v2 NewPool() is being discussed that eliminates the sharp edges by making it generic: https://github.com/golang/go/issues/71076 https://github.com/golang/go/issues/71076 I haven't personally found it to be problematic; just keep it private, give it a default new func, and be cautious about only putting things in it that you got out.
- nasretdinov 2y agoI think in general people understand that sync.Pool introduces essentially an equivalent of unitialised memory (since objects aren't required to be cleaned up before returning them to the pool), and mostly use it for something like []byte, slicing it like buf[0:0] to avoid accidentally reading someone else's memory. But the instrument itself is really sharp and is indeed kind of last resort
- dennis-tra 2y agoCan someone explain to me why the compiler can’t do struct-field-alignment? This feels like something that can easily be automated.
- CamouflagedKiwi 2y agoBecause the order of fields can be significant. It's very relevant for syscalls, and is observable via the reflect package; it'd be strange if the field order was arbitrarily changed (and might change further between releases). I assume the thinking was that this is pretty easy to optimise if you care, and if it's on by default there'd then have to be some opt-out which there isn't a good mechanism for.
- kbolino 2y agoIn particular, struct field alignment matches C (even without cgo) and so any change to the default would break a lot of code.
- 9rx 2y ago> struct field alignment matches C (even without cgo) The spec defines alignment for numeric types, but that's about it. There is nothing in the spec about struct layout. That is implementation dependent. If you are relying on a particular implementation, you are decidedly in unsafe territory. > so any change to the default would break a lot of code. The compiler can disable optimization on cgo calls automatically and most other places where it matters are via the standard library, so it might not be as much as you think. And if you still have a case where it matters, that is what this is for: https://pkg.go.dev/structs#HostLayout https://pkg.go.dev/structs#HostLayout
- kbolino 2y agoThat's good to know. I'm not making use of this assumption, but the purego package (came from ebitengine) does. It looks like they're aware of HostLayout [1] but I'm not sure how many other people have gotten the memo (HostLayout didn't exist before Go 1.23). [1]: https://github.com/ebitengine/purego/issues/259 https://github.com/ebitengine/purego/issues/259
- black_13 2y ago[dead]
- __turbobrew__ 2y agoCalling mmap “zero copy” is generous. I guess we glaze over the whole page fault thing, or the fact that performance is heavily dependent on how much memory pressure the process is under. This is the same n00b trap that derailed the llama.cpp project last year because people don’t understand how memory maps and paging works, and the tradeoffs.
- inadequatespace 2y agoWhy doesn’t the compiler pack structs for you if it’s as easy as shuffling around based on type?
- greatgib 2y agoBecause the organization of your struct is exactly how the memory have to be organized and that might be important for you. The compiler doesn't know your intended usage so it can't rework the structure at its will. For example, you might take the block of memory and data and send it to another system that will decode it. Or you can take the block of memory and store it in a file or in a hardware device where it means something in this specific order.