3 ms·
Writing inline assembly to get SIMD performance is likely to cost immense time of an architecture expert and doesn't scale. So my point stands: If a compiler c
by sddfd 9y ago
Writing inline assembly to get SIMD performance is likely to cost immense time of an architecture expert and doesn't scale.
So my point stands: If a compiler can't produce vectorised code, the compiler needs to be improved.
Spending time on improving the compiler is sustainably spent time. Spending time on programming SIMD in assembly by hand likely is not.
I believe you that the vectorization support at the moment is not good enough, and I completely understand that not everyone can spend time on improving the compiler.
However, it seems you know exactly what a compiler should be able to do and what feedback from the compiler, or annotations for the compiler, would be helpful for your use-cases where optimization somehow didn't figure out what to vectorize.
I was just pointing out that communicating with compiler people and getting them to improve the compiler is likely a better move than asking for better support for inline assembly.
- dagss 9y agoThe problem is not communication. To quote Paul Graham in a keynote from some years back: People have been waiting for "sufficiently smart compilers" to cut programmers out of the loop for 30 years. It is a neat idea in theory. In practice it is just very hard. You need to design the algorithm for the hardware, so it is sort of an AI problem. Programmers do fill a role usually in coming up with the best algorithms... The point is simply that C doesn't have the concepts you want to program with. Vectorization only gets you so far, there are many other things you can do with AVX/SSE. It doesn't need to be ASM, it could be supersets of C instead (like CUDA). Intrinsic functions for AVX is what I use and I don't see the problem with using those. It sort of is such a superset of C. In practice a few numerical computations (BLAS, FFTs, etc) are reused a lot by many. It is worth it to write those few libraries in assembly (or at least intrinsics). For the rest we just have to live with a small performance penalty. I am just saying "if you want to go the extra mile to get high performance, you can beat compilers". In most scenarios of course it makes most sense to not bother.
- sddfd 9y agoI get your point. However, I don't think it is useful to consider a compiler to be a replacement for programmers. Compilers are tools for programmers. The more time a compiler saves a programmer, the better it is. I think there is a middle ground between writing inline assembly and fully automatic vectorization that would be less time intensive than manual vectorization and more predictable and available earlier than fully automatic vectorization. I wonder what would have to be done to find it and provide support for it in GCC/LLVM.
- corysama 9y agoSo... SIMD intrinsics? You also need to take into account that a huge chunk of the work in SIMD is the need to rearrange your data to be more amenable to the CPU. C++ compilers are very restricted in what they can do for you there.
- deleted 9y ago[deleted]
- coldtea 9y ago>The problem is not communication. To quote Paul Graham in a keynote from some years back: People have been waiting for "sufficiently smart compilers" to cut programmers out of the loop for 30 years. It is a neat idea in theory. In practice it is just very hard. I don't see how it's "very hard" in the general case. Compilers have been increasingly cutting programmers out of the loop for 30 years now. It's the reason we program apps and even games in C++ and not in hand-rolled assembly as we did in the 80s, and why we can use crazily high level languages like JS and still get within an order of magnitude of C with a modern JIT.