40 ms·
Nice. I'm always slightly disappointed in the amount of optimization and intrinsics in the JDK for these kinds of fundamental and frequently-used methods, despi
by dtech 3y ago
Nice. I'm always slightly disappointed in the amount of optimization and intrinsics in the JDK for these kinds of fundamental and frequently-used methods, despite this always being one of the main arguments for JIT.
- saagarjha 3y agoWhy is it disappointing?
- WJW 3y agoNot GP, but in almost every introductory text to JIT compilers I've read there is a section called "advantages and disadvantages" and it almost always mentions that JIT compilers in theory have more information available than ahead-of-time compilers (ie all the information AOT compilers have and then also runtime information) and can use that information to safely do optimizations that AOT compilers cannot. But then when you look at the actual code of JIT compilers, such JIT-only optimizations seem extremely rare. The JVM, surely one of the platforms that have had more than enough dollars and top-level CS talent thrown at it, even after ~20 years it still apparently lacks many optimizations that you'd expect to be in there if you've only read the introductory texts about it. The linked MR about AVX optimizations for sorting is one such example IMO. AVX512 was first announced 10 years ago, and (not being deeply into JVM development myself) I might have assumed its use would be more prevalent in the JIT output than actually seems to be the case.
- nraynaud 3y agoI believe by having deeper inlining, unrolling and specialization in a JIT you can gain more performance than by static analysis. You squeeze more performance out of stupider algorithms.
- WJW 3y agoOf course, and dead code elimination can also be improved upon (compared to AOT compilers) because some inputs like command line flags will never change during the execution of the program. This means you can better predict which branches will be taken. These are optimizations that work for almost any program. OTOH, another way in which JITs "should" generate better-performing code is by tailoring their output to the platform on which the program is currently being run. With AVX being quite prevalent on the server-grade CPUs on which many big JVM programs run on, I don't think it would be unreasonable to expect the JVM to have more support for AVX512 in its code generator than it apparently does. Is low hanging fruit like 10x speed improvements in sorting something you'd expect out of a very mature platform like the JVM? I don't mean to harp on the JVM devs here, JIT development is a Very Hard Problem. It's just that I can understand why GGP is disappointed, JITs in general don't seem to quite deliver on the excitement they generated when they were new.
- kaba0 3y agoI don’t know — having almost native speed with all the benefits of a fat runtime like attaching a debugger to a prod process to check the number of objects, or streaming runtime logs without almost any overhead sounds like they do deliver.
- kllrnohj 3y ago> having almost native speed But it doesn't unless your "almost" is very generous. Java is pretty consistently 2-10x slower than the major performance-focused AOT offerings (C, C++, Rust) Now maybe you call 2x "almost", but let's phrase it in terms of CPU performance over time. That's equivalent to 10 years of CPU hardware advancements. To me that's a lot of overhead. Depending on who is paying for the CPU time vs. the developer time it's regularly a cost worth paying, but at the same time don't pretend it's "almost native speed", either. It is a cost and a rather significant one at that. Just, so are engineers. They also aren't cheap.
- pjmlp 3y agoOn the other hand .NET is much closer, because it supports value types and CLR was designed to be targed by C++ as well, so it supports most of the crazy C and C++ stuff (not counting UB related ones). Likewise if you code in C, C++ or Rust with allocations everywhere, bad algorithms or data structures, being AOT won't help.
- dtech 3y agoUnfortunately, Hotspot does not do all that much in that regards either in my experience. Graal is better but lags in language support and usage.
- kllrnohj 3y agoAOTs consistently do much more inlining & unrolling than JITs do. JITs often have preset heuristics for figuring out where to split the code that's practical to implement rather than being the most performant possible. After all, the code needs to hot-swapped in without much disruption. And then similarly JITs need to optimize for compilation speed which limits the amount of optimization passes it'll do, because the JIT result is itself a "hot path." JITs do regularly have a lot more specialization optimizations, though, but is that really because it's a JIT instead of an AOT or is it more because JIT'd languages just often tend to also be more dynamic ones as well?
- gergo_barany 3y ago> AOTs consistently do much more inlining & unrolling than JITs do. Nonsense. What evidence do you have for this claim? > JITs often have preset heuristics for figuring out where to split the code that's practical to implement rather than being the most performant possible. After all, the code needs to hot-swapped in without much disruption. You seem to be talking about JITs without on-stack replacement. So not state of the art high performance JITs.
- deleted 3y ago[deleted]
- dtech 3y agoYou might not be me, but formulated my point extremely well.
- kaba0 3y agoAutovectorization is a famously difficult optimization and runtime information doesn’t make it any easier. That is more useful for optimistic assumptions, like these pointers won’t be zero, only a single instance exists for this interface, etc. Nonetheless, you might find the Graal compiler doing a better job at autovectorization than C2.
- MrBuddyCasino 3y agoOne would think such relatively low-hanging fruit (dispatching to an existing avx512 lib) would have been picked by now, considering the massive effort to improve the JVM in other areas (virtual threads, value objects, C interop, Graal, novel GCs). Maybe Array.sort() isn’t that frequently used, as data sorting is often done by the database?
- carpenecopinum 3y agoThe big issue here (provided that I'm reading the PR correctly) is that it's purely for arrays of the primitive number types. Whereas most real applications (that I've seen) will be sorting objects by some (potentially computed) property, be it an ID, a timestamp or a name. For all of these cases, the linked pull request won't be doing anything useful. The most useful application for this (outside of getting nicer numbers on a benchmark) that I see is computing the median/quantiles of some property on a bunch of objects that aren't already sorted by the interesting property.
- marginalia_nu 3y agoYou still see quite a lot of primitive number sorting in library code. If you want your Java code to go fast, you typically stick to primitives in arrays. Most Java application code doesn't, but it typically uses libraries that do. A sorted array is a priority queue, a binary search tree, etc.
- marginalia_nu 3y agoI think the number of applications that will see a measurable speed-up from this improvements is very small. If you have an application that is truly bottlenecked the performance of number-sorting in any measurable way, then you most probably didn't write it in Java. It's not really a number crunching language for a variety of reasons.
- xmcqdpt2 3y agoJava high-performance programming would be significantly improved by generics over primitive types. Writing performant code in Java feels like writing Fortran or C. You end up using libraries like fastutil, which is "generic code" templated by C preprocessor macros, https://fastutil.di.unimi.it/ https://fastutil.di.unimi.it/