4 ms·
The concurrent stop-the-world garbage collector is a performance issue. Additionally, the Plan 9-based compiler is not an optimizing compiler.
by ootachi 15y ago
The concurrent stop-the-world garbage collector is a performance issue. Additionally, the Plan 9-based compiler is not an optimizing compiler.
- dsymonds 15y agoHave you found the GC to be an actual problem? Or are you assuming it is? It's a very different situation to other GC-based languages.
- ootachi 15y agoWhy is it so different? I'd be curious to know. In my experience, a stop-the-world collector that must be threadsafe ends up being a disaster. You hit a brick wall in terms of real-time performance quickly (look at iOS versus Android for an easy example, and Dalvik's GC is much more sophisticated than that of Go these days). I predict that Go is going to have to put a lot of effort into making the GC fast in order to compete with C++ and Java. They're going to need a concurrent, generational garbage collector much like Java's, probably with a server mode and a client mode. It'll have to be heavily tuned and will require man-years of work.
- enneff 15y agoJava requires a powerful, sophisticated garbage collector because it is extremely difficult (and in many cases impossible) to write Java programs that don't generate a lot of garbage. Even many of the core APIs are allocation heavy. It's a pain. Go data structures tend to be much smaller than the Java equivalents, and it is much easier to track down and eliminate unnecessary allocations in Go code. When you have better control over allocations you don't need to lean so hard on the garbage collector. (Some of this is touched on here http://loadcode.blogspot.com.au/2009/12/go-vs-java.html http://loadcode.blogspot.com.au/2009/12/go-vs-java.html) You seem to have a lot of strong opinions (and predictions!) about Go, when you clearly don't have much experience with it.
- ootachi 15y agoGo gives you slightly better control over allocation than Java (in that you have a choice to allocate on the stack and inside other data structures -- but keep in mind that escape analysis can give you this too, see the optimizations in the Jikes RVM). However, in Go, you still have no choice but to allocate on the heap in many instances (for example, when returning a data structure from a function, or when using maps, etc.) You're taking a big bet that programs will be able to use the stack and manually use free lists or whatever (which you can still do in Java!), and that the slow performance of the GC won't be a problem. I wouldn't take that bet. Besides, once you have a stop-the-world multithreaded GC, it has to trace all the roots in the program for correctness. There's no way around that. At that point having all the data on the stack doesn't help you (except for cache locality, but Java's GC already does that via the nursery and copying during major collections). You must still trace every pointer in the object graph, while keeping all threads suspended. You can run the GC less often, but that's not much of a help when interactive performance is at stake. iOS feels so great because the UI is always running at 60 frames per second. Stop-the-world GC can't do that.
- dsymonds 15y agoWe don't bet, we experiment, measure and analyse. So far, I have found memory allocation in Go to be extremely predictable and controllable, and the GC behaviour likewise predictable. I don't know why you claim that "having all the data on the stack doesn't help you". That makes no sense. If data is on the stack, any references starting there disappear as soon as that stack frame is popped, so there's no lingering work for the GC to do; the GC arena is unchanged. Your comments strongly imply that you have no practical experience with Go's GC. I suggest you stop claiming that it has certain performance characteristics or behaviours when you have not experienced it yourself.
- ootachi 15y ago"If data is on the stack, any references starting there disappear as soon as that stack frame is popped, so there's no lingering work for the GC to do" You're only considering the sweep phase. Sweep is the easy part of tracing GC - you can always chuck sweep into a background thread. The mark phase is the problem. Marking always has to trace roots, including stack roots - that's how GC works. Tracing GC never traces dead objects. You can reduce allocation pressure by using the stack, which will make the GC run less often, but my point is that when it does run you're no better off than Java, and quite a bit worse since Java's GC can run in parallel with the mutator and Go's can't. "Your comments strongly imply that you have no practical experience with Go's GC. I suggest you stop claiming that it has certain performance characteristics or behaviours when you have not experienced it yourself." I have experience with GC generally. There's nothing particularly special about Go's GC: it's a standard stop-the-world mark-and-sweep collector for a language that supports a limited form of stack allocation but generally uses heap allocation. The performance characteristics that this form of GC must have are well-known. If you want data, look at the binary-trees benchmark: http://shootout.alioth.debian.org/u32/benchmark.php?test=all&lang=java&lang2=go http://shootout.alioth.debian.org/u32/benchmark.php?test=all... It's mostly a test of GC. Java's GC runs more often (thus the memory use is lower) and yet it's still 4x faster. This is because Java has a generational, concurrent-incremental collector.
- enneff 15y ago"...not an optimizing compiler." ? The gc compiler does perform some optimizations, including inlining across package boundaries. There is also the gccgo compiler which takes full advantage of gcc's many optimizations. Both compilers will be equally supported in Go 1.
- ootachi 15y agoI'm referring to the Plan 9-based compilers, which do not perform the optimizations necessary to make it competitive with modern C and C++ compilers. Here I'm referring to stuff like SSA-based global value numbering, scalar replacement of aggregates, loop-invariant code motion, etc., as well as lower-level stuff like instruction scheduling. gccgo is better here, of course. Unfortunately, the Plan 9-based compilers are the ones most people use...
- enneff 15y agoIt doesn't matter what "most people use". Both compilers are equally accessible with the new go command (http://weekly.golang.org/cmd/go/ http://weekly.golang.org/cmd/go/). People who find their programs faster with gccgo will use gccgo, and vice versa. Not competitive with "modern C and C++ compilers"? There are few real world programs that are improved by code generation micro-optimizations, and even for those that are Go is still competitive. http://blog.golang.org/2011/06/profiling-go-programs.html http://blog.golang.org/2011/06/profiling-go-programs.html
- ootachi 15y ago"There are few real world programs that are improved by code generation micro-optimizations" I don't know how to convince you of the fact that this statement just isn't true. Consider the rasterization software you're using right now in your web browser. Micro-optimizations hugely matter for blitting and tessellation. Consider video decoding (or encoding!). Using SSE instructions instead of going word-by-word or byte-by-byte makes an enormous difference. Read Dark Shikari's blog posts about x264 if you want to see how much good use of SSE matters for video encoding. To name just one example: Modern x86 processors have an optimization that allows the fetch-and-decode step to be skipped for the second and subsequent iterations of small loops. By "small" I mean "really really small" -- something like 16 bytes is the number I've heard. By shaving a byte or two off the instruction encoding to fit a loop from 18 bytes into 16 bytes, the loop performance can double or triple. When that loop is the loop that drives alpha blending on your touch-sensitive mobile app, that can be the difference between an app that feels solid and fluid and an app that feels slow and pokey. People use systems languages for performance. Folks for whom performance is irrelevant are all on dynamic languages, and that's not changing. Systems programmers need a compiler that knows about these micro-optimizations.