5 ms·
The entry to the article suggests compiling for your architecture and the n deps several other things. The article benchmarks - Ofast which is poorly named. It
by eadler 7y ago
The entry to the article suggests compiling for your architecture and the n deps several other things.
The article benchmarks - Ofast which is poorly named. It's really -Obroken-by-design. It'll be "faster" but completely break applications.
It also suggests using omit frame pointer which destroys debugability.
-march and -mtune are the parts that the the article title and intro actually suggest. While possible, I see no evidence that this matters. As I understand it the arch that Java is compiled with is not the same as the one that gets used for JIT compiling.
- earenndil 7y ago> [-Ofast will] be "faster" but completely break applications Except for the jvm, apparently. And everything else I've tried it with. > It also suggests using omit frame pointer which destroys debugability. Which is completely useless except for jvm developers. > As I understand it the arch that Java is compiled with is not the same as the one that gets used for JIT compiling. The performance of the compiler itself matters, not just the performance of the generated code, because, since it's a JIT, compiler code continues to run.
- rjsw 7y agoThe Hotspot JIT reads the CPU configuration at runtime to choose which optimizations are best or which instruction extensions are available.
- chrisseaton 7y agoThat doesn't help the performance of the runtime code, which is C++ or AOT-compiled Java.
- pron 7y agoRight, but the OpenJDK VM (HotSpot) uses three JITs -- C1, C2 and Graal -- and two of them are written in C++, so C++ compiler flags could affect the performance of the JIT compilers, although not of the code they generate. Because the performance of the emitted code is far more important than the performance of the compiler, I doubt that will make a difference, but there are other important parts of the OpenJDK VM that are written in C++ and whose performance might be affected, most notably the GCs.
- msbarnett 7y agoThat's great for the Java code being JITed, but does nothing for the JIT compiler itself, or the garbage collectors, or any other runtime components that are written in C or C++ and compiled (and optimized) ahead of time.
- simpsond 7y agoYeah, I suspect a good chunk of the performance improvement is related to GC performance improvement. The netty benchmark with pooled buffers compared to unpooled buffers hints at that (although I don't know for sure).
- paulddraper 7y ago> omit frame pointer which destroys debugability. IIRC frame pointers were necessary for a Linux flame graph tool I used on the JVM.
- umanwizard 7y agoBased on `perf`? If so, try passing `--call-graph dwarf` in your `perf record` invocation. Then things should work fine despite lack of frame pointers.
- edoceo 7y agoI run other apps (in production) with omit-frame-ponter. We know it kills debug. Some narrow cases it makes sense. Like many optimization "tricks"
- gpderetta 7y agoWait, omit-frame-pointer has been the default on GCC x86_64 for a while now, even in debug build. The debugger is supposed to use unwind tables to to traverse the stack. Am I missing something?
- chrisseaton 7y ago> It also suggests using omit frame pointer which destroys debugability. Well, yeah it's an article about performance optimisations, and disabling this debug mechanism increases performance (or maybe it does - I haven't measured it myself.) If you need to debug, don't optimise for performance this aggressively. Seems a reasonable tradeoff?
- alfalfasprout 7y agoIt can completely break applications. Particularly those that require a particular level of floating point precision (in which case a fixed precision library is actually a better choice and those that use abundant aliasing (which plenty of modern static analyzers can help with). But for the most part I've been using -Ofast for years in a wide variety of applications with no issues.