4 ms·
I'm not an expert on GC, but I do work on a compiler written in a GC'd language that cares quite a bit about performance. I can share a few guidelines that may
by Locke1689 12y ago
I'm not an expert on GC, but I do work on a compiler written in a GC'd language
that cares quite a bit about performance. I can share a few guidelines that
may prove useful to the author or anyone else having similar trouble.
1) Profile. This does not mean gather a few charts from a primitive profiling
tool that simply tells you which functions have the most time spent in them. You
need a real profiling tool that gives you detailed analysis of where your time
is going. You need CPU stacks, time blocked, GC stacks -- everything you can get
your hands on. Your profile needs to tell you where, what, and how much you're
allocating. It needs to tell you what kind of allocations you are doing and how
long they are living.
2) Once you have your information, you need to figure out what you can do to
improve your performance. To do this, you need to know the characteristics of
your program and all the platforms it runs on. Is your code being JITed? How
long does your program spend in the JIT? Can you mitigate this by doing AOT
compilation, optimization profiles, or producing more JIT-friendly code? For GC,
what kind of GC are you running on? Is in conservative -- do you have to worry
about memory leaks? Is it stop-the-world -- do you have to worry about pauses?
Is it generational -- does the behavior change at inflection points? Do you have
to adapt your allocations to the sections of your program's run time? Does
allocation in one portion affect allocations in another portion, e.g. will your
allocations produce extra memory fragmentation when running on a non-moving GC?
3) You need to own your code and its performance. If you're allocating in a hot
loop and the GC can't keep up, you need to rewrite that section to not allocate.
If you're fragmenting the heap in non-critical paths, you need to stop
allocating there, as well. If you're constantly recreating re-usable ephemeral
objects, you need to look at object pooling. If you need to refactor sections
of your code to avoid allocations and stop-the-world collections, you need to do
that. Blaming the GC will not make your application faster. If you need to
break out a debugger and jump into the GC to see what's taking so much time,
that's what you need to do.
4) After you've optimized, re-evaluate. Are there actionable items you can hand
to the GC devs? Can you point out general points where GC strategy has a known
bad result and a known better implementation? File these on the GC. Maybe you'll
be able to delete some optimizations in future versions. :)
The first version of Roslyn was 30x slower than the old C# compiler written in C++. We're now within 2x on 10-year-old hardware and beating it on newer. Perf optimization works -- but not all programming is a stroll through a park. :)
- seanmcdirmid 12y agoWhy the performance discrepancy between new and old hardware? Do the immutable trees require more RAM that is more likely to be provided on newer hardware? Also, I would suspect going immutable would require more tuning in terms of allocation.
- Locke1689 12y agoYeah -- immutability and high concurrency tend to hurt single threaded performance (especially when compared to optimized C++), but we gain a ton from new highly parallel hardware. Also, GC can do analysis on spare cores.
- seanmcdirmid 12y agoI really should look at the architecture, I'm interested in how to concurrent it is. I use an alternative approach that is much more imperative (using dependency tracing and re-execution to ensure consistency).
- azth 12y ago> We're now within 2x on 10-year-old hardware and beating it on newer. Is it an apples-to-apples comparison though? If you use the same programming approaches in C++ as you're using in Roslyn (immutable structures, etc.), do you expect you would still keep beating it?
- Locke1689 12y agoYup. Because a project that's never done can't beat a project that will be.
- copx 12y ago> Your profile needs to tell you where, what, and how much you're allocating. Which languages offer such advanced profilers, though? I am using Lua (a GCed language) right now, and I am not aware of any Lua profiler which profiles allocation/the GC. I think D's profiler does not feature that either. In fact I have only ever heard of Java having such advanced profiling tools, and I guess Microsoft build the same thing for their version of Java (.NET). But again, which other languages have tools which allow you to profile allocation/GC behavior?