4 ms·
For the limited scope of rectangular loop nests without loop carried dependencies, my Julia library LoopVectorization.jl is able to vectorize and achieve often
by celrod 6y ago
For the limited scope of rectangular loop nests without loop carried dependencies, my Julia library LoopVectorization.jl is able to vectorize and achieve often significantly better performance than Clang, GCC, or the Intel compilers; see the different examples here: https://chriselrod.github.io/LoopVectorization.jl/latest/examples/matrix_multiplication/#Matrix-Multiplication-1 https://chriselrod.github.io/LoopVectorization.jl/latest/exa...
However, that's only one component of vectorization. It's still relying on LLVM to do lots of clean up, among other things.
I feel like there's a lot of reverence surrounding compilers and their capabilities, but at least when it comes to SIMD, it's not hard to find examples where a human can do much better, or can at least come up with algorithms that will do better for that and related problems.
I think saying compilers produce better asm than a programmer is like saying statistical libraries like `GLM.jl` are better at doing statistics than statisticians. True in that you should let the routines crunch the numbers. The skill is in deciding which routines/analysis to run, how to organize them, and -- when necessary -- when to develop your own.
- gnufx 6y agoI don't know what to make of results without any idea how to reproduce them. For what it's worth, BLIS' generic C DGEMM code gave around 60% of the hand-coded kernel on Haswell with GCC (8, from memory) on large matrices without any real effort to speed it up. Obviously that's not a simple triply-nested loop, since just vectorizing is useless for large dimensions, and the generic code does get well auto-vectorized. There are results from Polly/LLVM showing it do well against tuned GEMM. However, Polly apparently pattern-matches the loop nest and replaces it with a Goto BLAS-like implementation, and I don't know whether it's now in clang/flang. I also don't know how well the straight Pluto optimization does relative to BLAS, and whether that's what original Polly uses; it's a pity GCC's version isn't useful.
- celrod 6y agoI should include more detailed instructions. Dependencies include: eigen, gcc, clang, and (harder) the Intel compilers, as well as a recent version of Julia. If you have those, you should be able to instantiate the Manifest: https://github.com/chriselrod/LoopVectorization.jl/tree/master/benchmark https://github.com/chriselrod/LoopVectorization.jl/tree/mast... and then `include("/path/to/LoopVectorization/benchmark/driver.jl")` to run the benchmarks. There are other examples. For example, for code as simple as the dot products, clang shows some pretty bad performance characteristics while gcc needed a few infrequently used compiler options to do well there. Image filtering / convolutions without compile-time known kernel sizes are another example that showed dramatic speedup. When talking about optimized gemm, LoopVectorization currently only does register tiling. Therefore its matmul performance will continue to degrade beyond 200x200 or so on my deskop. An optimized BLAS implementation would still need to take care of packing arrays to fit into cache, multithreading, etc. Maybe I should add pluto. I did test Polly before, but it took several minutes to compile the simple c file, and the only benchmark that seemed to improve was the gemm without any transposes. And it still didn't perform particularly well there -- worse than the Intel compilers, IIRC. PRs are welcome.