6 ms·
You might already know this, but there is one potential caveat with APIs like that when it comes to performance or at least measuring their performance. The hot
by palinkapika 6y ago
You might already know this, but there is one potential caveat with APIs like that when it comes to performance or at least measuring their performance. The hot loop (the code that actually iterates over your data points) does not live in your code base. But performance often depends on what the JIT compiler makes out of that loop. If the API is used at several locations in your program, the compiler might not be able to generate code that is optimal for your call-site and inputs. Instead, it will generate generalized code that works for all inputs but might be slower.
However, when writing benchmarks, there is often no other code around to force the JIT compiler to generate such generalized code.
The following code demonstrates this. If the parameter warmup is set, I invoke the forEach methods with different inputs first (they do the same but are different methods in the Java byte code). The purpose is to force the compiler to generate generalized code:
@Param({"true", "false"})
public boolean warmup;
@Setup
public void setup() {
if (!warmup) {
Interval.oneTo(20).forEach(
(IntProcedure) i -> Interval.oneTo(1_000_000).forEach(
(IntProcedure) j -> Integer.toString(j)));
Interval.oneTo(20).forEach(
(IntProcedure) i -> Interval.oneTo(1_000_000).forEach(
(IntProcedure) j -> Integer.toString(j)));
Interval.oneTo(20).forEach(
(IntProcedure) i -> Interval.oneTo(1_000_000).forEach(
(IntProcedure) j -> Integer.toString(j)));
}
}
@Benchmark
public void coollectionsBlackhole(Blackhole blackhole) {
Interval.oneTo(20).forEach(
(IntProcedure) i -> Interval.oneTo(1_000_000).forEach(
(IntProcedure)j -> blackhole.consume(Integer.toString(j))));
}
@Benchmark
public void collectionsDeadCode() {
Interval.oneTo(20).forEach(
(IntProcedure) i -> Interval.oneTo(1_000_000).forEach(
(IntProcedure)j -> Integer.toString(j)));
}
And the for loop implementations for reference:
@Benchmark
public void loopBlackhole(Blackhole blackHole) {
for (int j = 0; j < 20; j++) {
for (int i = 0; i < 1_000_000; i++) {
blackHole.consume(Integer.toString(i));
}
}
}
@Benchmark
public void loopDeadCode() {
for (int j = 0; j < 20; j++) {
for (int i = 0; i < 1_000_000; i++) {
Integer.toString(i);
}
}
}
@Benchmark
public void loopBlackholeOnly(Blackhole hole) {
for (int j = 0; j < 20; j++) {
for (int i = 0; i < 1_000_000; i++) {
hole.consume(0xCAFEBABE);
}
}
}
On my desktop machine, this gives me the following results (Java 11/Hotspot/C2):
Benchmark (warmup) Mode Cnt Score Error Units
collectionsDeadCode true avgt 99 162.482 ± 1.615 ms/op
collectionsDeadCode false avgt 99 218.816 ± 4.217 ms/op
coollectionsBlackhole true avgt 99 235.122 ± 1.362 ms/op
coollectionsBlackhole false avgt 99 270.192 ± 1.627 ms/op
loopBlackhole true avgt 99 207.214 ± 1.162 ms/op
loopBlackhole false avgt 99 206.711 ± 0.932 ms/op
loopBlackholeOnly true avgt 99 74.774 ± 0.180 ms/op
loopBlackholeOnly false avgt 99 74.359 ± 0.180 ms/op
loopDeadCode true avgt 99 143.394 ± 0.900 ms/op
loopDeadCode false avgt 99 142.654 ± 0.795 ms/op
All results are for single threaded code. Now, the difference is not huge, but significant. Overall it looks like the invocation of toString(int) is not removed and accounts for most of the runtime.
Just to be clear: I am not saying one should stay way from the stream APIs. As soon as the work per item is more than just a few arithmetic operations, there is a good chance the difference in runtime is negligible. But when doing numerical work (aggregations etc.), a simple loop might be the better option.
Finally, these differences are of course compiler-dependent. For example, C2 might behave differently than Graal and who knows what the future brings.
- palinkapika 6y agoAnd of course I screwed up the warmup parameter...
- eggsnbacon1 6y agoI didn't think about this!! That could a big performance hit to pay for nice API's. I think inlining might save us here? I'm not sure what the metrics are around inlining lambdas... Does it look at the size of the code wrapping the lambda or the size of the lambda block itself to decide inlining eligibility? If it only considers the size of the wrapper, maybe the small forEach method will always get inlined with the block it calls? EDIT: according to this https://stackoverflow.com/questions/44161545/can-hotspot-inline-lambda-function-calls https://stackoverflow.com/questions/44161545/can-hotspot-inl... HotSpot will inline lambda method calls as long as depth doesn't exceed maxinlinelevel (9 by default). Assuming the API designers made their methods small enough to be inlined, they should all be flattened out. Interestingly, this maxinlinelevel default was increased to 15 in Java 14! https://bugs.openjdk.java.net/browse/JDK-8234863 https://bugs.openjdk.java.net/browse/JDK-8234863 . Probably because more people are using streaming/reactive API's?
- palinkapika 6y agoYes, this is strongly related to inlining. Without the warmup code C2 most likely inlines your lambdas into the code generated for the for each method. At least I have seen this behavior with similar APIs (you can use a tool called JITWatch to visualize compiled and inlined methods). However, this does not scale. Last time I checked C2 inlines at most two implementations at a polymorphic call-site (in this case the line that calls your lambda). If you pass in more lambda functions, it won’t inline them and might de-optimize existing compiled code to remove previously inlined functions (there are also cases where one implementation that is heavily used gets inlined and the others will be invoked via a function call). Graal does not seem to have the same limit but when I tested it two years ago it had others averaging out to the same performance. Thus, increasing the max inline level does not necessarily help (you can already do that in Java 8 and 11) for this particular problem. What you would want is that the compiler inlines the for each method and the local lambdas into the methods using the API (the benchmark methods in our case), i.e., you want the compiler to copy the hot loop somewhere higher up in the call stack and then optimize it. But apparently this is easier said than done. But again, this is probably only relevant for workloads where there is very little work to do for each item.