3 ms·
Half the point of using SIMD intrinsics rather than embedding assembly is that register allocation and instruction scheduling are performed by the compiler. Usi
by nikic 8y ago
Half the point of using SIMD intrinsics rather than embedding assembly is that register allocation and instruction scheduling are performed by the compiler. Using intrinsics certainly does not guarantee that exactly the corresponding instructions will be emitted in exactly the given order.
- Const-me 8y ago> Using intrinsics certainly does not guarantee that exactly the corresponding instructions will be emitted in exactly the given order. I hope they will fix the compilers to add such guarantees. Currently, every time compilers try to mess with manually-written intrinsics, they decrease performance. See this bug about LLVM failing to emit the exact instructions, decreasing the performance: https://bugs.llvm.org/show_bug.cgi?id=26491 https://bugs.llvm.org/show_bug.cgi?id=26491 See this bug about VC++ failing to emit them in the given order, again decreasing performance: https://developercommunity.visualstudio.com/content/problem/350355/the-compiler-optimized-a-single-very-fast-instruct.html https://developercommunity.visualstudio.com/content/problem/...
- nikic 8y agoAs usual, things aren't quite that simple. While it certainly can happen that the compiler can pessimize carefully crafted SIMD code by applying undesirable transformations, simply not optimizing intrinsics would lead to other problems. If you consider SIMD code that spans more than just one small function, then it is quite useful if the compiler can SROA (e.g. you pass around an array of vectors instead of having 8 arguments to each function, but still want them to end up in registers eventually), LICM (you're calling a function that needs a constant vector in a loop) or otherwise optimize your code. If you want the compiler to leave alone your intrinsics, link in an assembly file.
- Const-me 8y ago> it is quite useful if the compiler can SROA __attribute__((always_inline)) / __forceinline usually help. > LICM (you're calling a function that needs a constant vector in a loop) I can calculate that vector outside of the loop. > link in an assembly file. Way more complex, I need to write outer scalar code in assembly too, need to manually allocate registers, also C++ templates are sometimes very useful for that kind of code. In 99% of cases intrinsics are good enough for me, but I would love the compilers to leave alone my intrinsics.
- marmaduke 8y agoCan’t just isolate those in a compilation unit with -O0?
- Const-me 8y agoInteresting idea, will try next time. I usually want to optimize the scalar code outside of the manually-vectorized body of the loops. A function call to that external compilation unit will be slower than inlining I have when everything is in the same unit. However, it could be the call overhead is small enough, obviously need to profile.
- comex 8y agoOn GCC-like compilers, you could just use inline assembly. That definitely won’t get rewritten into some other instruction sequence, and it can handle things like register allocation and loads/stores for you. Downsides include that the compiler won’t be able to estimate instruction timings, the ease of screwing up the input/output notation, and that MSVC doesn’t support inline assembly on x64 at all.