6 ms·
A Guide to the Go Garbage Collector
- hsn915 4y agoI wonder if generics would allow custom allocators. I haven't tried it yet but it seems like an Arena/bump allocator for example should be possible now.
- tapirl 4y agoGenerics are totally helpless for runtime things. I would be good for the official runtime to be designed in a plugin way, so that third parties may experiment their own implementations of some aspects of the runtime.
- philosopher1234 4y agoWhy can’t they just fork the runtime to experiment?
- hsn915 4y agofunc Allocate[T](arena *Arena) *T { var bytes = arena.Bump(sizeof(T), alignof(T)) return (*T) bytes } Any reason why this would not work? Maybe you need to cast to unsafe.Pointer or something before returning, but in theory this _should_ work.
- morelisp 4y agoIt could probably be made to work, but would it reduce any GC work compared to any other kind of object pool? The returned value will still be accessible from some stack root and need to be scanned, I think regardless of whether the arena was already scanned - the arena would mark the span of the T itself, but any subfields with pointers would need to be scanned from the T as it would need the type information. And if T has no pointer subfields it would not be scanned anyway. So maybe best case you save a mark per T? The advantage of a bump allocator is being able to throw it all out at once at the end and never have to check it while "in use", and I don't think Go's GC would let you do that.
- hsn915 4y agoRight, the pointers would still be scanned by the GC, so we can't elliminate the cost of the GC, unless we completely turn it off (I think that is possible). But the point of the arena allocator is to make allocations and deallocatins very fast. Say you want to allocate many small objects for a very short amount of time and then get rid of all of them. Doing it on an Arena would save a lot of computation that would be required in a typical heap allocator.
- morelisp 4y ago> But the point of the arena allocator is to make allocations and deallocatins very fast... Doing it on an Arena would save a lot of computation that would be required in a typical heap allocator. Relative to a completely naive allocator yes, but relative to any other kind of pooling (e.g. Go's internal small object pools, or a `sync.Pool`, or just a `make([]T, 1000)`) the advantage of a bump allocator is marginal without the ability to avoid the mark overhead and actually throw it all out at once.
- cube2222 4y agoThis is a really great guide! Nice to have something official and in-depth. I have two tips I can share based on my experience optimizing OctoSQL[0]. First, some applications might have a fairly constant live heap size at any given point in time, but do a lot of allocations (like OctoSQL, where each processed record is a new allocation, but they might be consumed by a very-slowly-growing group by). In that case the GC threshold (which is based on the last live heap size) can be low and result in very frequent garbage collection runs, even though your application is using just megabytes of memory. In that case, using debug.SetGCPercent to modify that threshold at startup to be closer to 10x the live heap size will yield enormous performance benefits, while sacrificing very little memory. Second, even if the CPU profiler tells you the GC is consuming a lot of time, that doesn't mean it's taking it away from your app, if it's single-threaded. `go tool trace` can give you a much better overview of how computationally intensive and problematic the GC really is, even though reading it takes some getting used to. [0]: https://github.com/cube2222/octosql https://github.com/cube2222/octosql
- eatonphil 4y agoI'd love to read more about your experience profiling, how your techniques work.
- cube2222 4y agoThanks, I'll try to whip up an article about it in the not-too-distant future. Though I can tell that the biggest improvement to my profiling flow was adding a `--profile` flag to OctoSQL itself. This way I can easily create CPU/memory/trace profiles of whole OctoSQL command invocations, which makes experiments and debugging on weird inputs much quicker.
- kccqzy 4y ago> Second, even if the CPU profiler tells you the GC is consuming a lot of time, that doesn't mean it's taking it away from your app I have experienced the same issue here. Our load balancer used CPU usage as a proxy for deciding how much traffic should be assigned when performing load balancing. When the app was written in Go, we consistently found that the GC is consuming a lot of CPU time even though all other metrics like request latency were very good, even in the microseconds range. This was the case even when the app was massively parallel with lots of goroutines. But the load balancer kept sloshing traffic around unnecessarily based on its observation that GC is consuming a lot of CPU time.
- deleted 4y ago[deleted]
- deleted 4y ago[deleted]
- erik_seaberg 4y agoHm, I was hoping for a roadmap that would talk about supporting generations and more tuning options.
- morelisp 4y agoWould generational support improve anything given a) 99% of the nursery is probably already on the stack, and b) using generations to inform any kind of compaction / relocation still seems out of the question? `GOMEMLIMIT` described in the document is a new tuning option.
- omginternets 4y agoHas anyone tried "gc_details": true in VSCode? I've just gone through the configuration steps, but I'm not seeing anything obvious. What should I be looking for? EDIT: found it at the top of the file.