3 ms·
The results for the parallel code are indeed interesting. However, the speedup < 2 with C2 is suspicious. Even on an old dual core / 4 thread system (common wor
by palinkapika 6y ago
The results for the parallel code are indeed interesting. However, the speedup < 2 with C2 is suspicious. Even on an old dual core / 4 thread system (common worker pool should default to 3 workers), I would expect it to be above 2 for this kind of workload.
The last time I looked into Graal (over a year ago), I ran into something I could not make sense of.
We were designing an API with methods similar to Arrays.setAll(double[], IntToDoubleFunction) which we expected to be called with large input arrays from many places in the code base.
The performance of the code generated by C2 dropped as one would expect once you were using more that two functions (lambda functions were no longer inlined and invoked for each element). For simple arithmetic operations the runtime increased by around 5x.
The performance of the code generated by Graal was competitive with C2's code for mono- and bimorphic scenarios for a much larger number of input functions. I don't recall the exact number, but I believe it was around 25. However, once we hit that threshold, performance tanked. I believe the difference was closer to 15x than than C2's 5x.
We never figured out what caused this.
Do you have an idea? Or, since this is probably outdated data, do you know what the expected behavior is nowadays? Is there maybe any documentation on this?
- gergo_barany 6y ago> However, the speedup < 2 with C2 is suspicious. Even on an old dual core / 4 thread system (common worker pool should default to 3 workers), I would expect it to be above 2 for this kind of workload. The speedup I found on C2 is 2.547 / 1.719 ~= 1.5, the OP's speedup (also on C2) was 0.71 s / 0.38 s ~= 1.9. Not the same factor, but not above 2 either. The benchmark is quite small, so the overhead of setting up the parallel computation might matter. The JDK version might matter as well. > Graal was competitive with C2's code for mono- and bimorphic scenarios for a much larger number of input functions. I don't recall the exact number, but I believe it was around 25. If we are talking about a call like Arrays.setAll(array, function) and you have one or two candidates for function, then you are mono- or bimorphic. If you have 25 candidates for function, you are highly polymorphic. I don't understand what you mean by a mono- or bimorphic scenario with up to 25 candidates. Performance will usually decrease as you transition from a call site with few candidates to one with many candidates. I agree that a factor of 15x looks steep, but it's hard to say more without context. The inliners (Community and Enterprise use different ones) are constantly being tweaked, though that's not my area of expertise. It's very possible that this behaves differently than it did more than a year ago. If you have a concrete JMH benchmark you can link to and are willing to post to the GraalVM Slack (see signup link from https://www.graalvm.org/community/ https://www.graalvm.org/community/), that would be great. Alternatively, you could drop me a line at <first . last at oracle dot com>, and I could have a look at a benchmark, without promising anything.
- palinkapika 6y ago> If we are talking about a call like Arrays.setAll(array, function) and you have one or two candidates for function, then you are mono- or bimorphic. If you have 25 candidates for function, you are highly polymorphic. I don't understand what you mean by a mono- or bimorphic scenario with up to 25 candidates. Sorry, I phrased that poorly. What I was trying to say was that the moment we introduced a 3rd candidate the performance with C2 decreased by ~5x. Whereas with Graal the performance stayed virtually the same when introducing a 3rd or 4th candidate. And it was performing well. But it did decrease at some point after adding more and more candidates and then the performance hit was much bigger. I will reach out to my colleagues still working on this next week (long weekend ahead where I live). If we can reproduce this with a current Graal release, I'll share the benchmark – always happy to learn something new and if it is 'just' why out benchmark is broken. :)
- eggsnbacon1 6y agoI really appreciate the in-depth replies on HN. Something Reddit lost ages ago. I gotta say your posts have introduced me to a lot of dark magic in the JVM that isn't really documented anywhere unless you count source code. Much appreciated!