3 ms·
It's a shame that SIMD is still a dark art. I've looked at writing a few simple algorithms with it but have to do it in my own time as it'll be difficult to jus
by secondcoming 2y ago
It's a shame that SIMD is still a dark art. I've looked at writing a few simple algorithms with it but have to do it in my own time as it'll be difficult to justify it with my employer. I do know that gcc is generally terrible at auto-vectorising code, clang is much better but far from perfect. Using intrinsics directly will just lead to code that's unmaintainable by others not versed in the dark art. Even wrappers over intrinsics don't help much here. I feel there's a lot of efficiency being left on the table because these instructions aren't being used more.
- Sesse__ 2y agoThe problem is that the different SIMD instruction sets are genuinely... different. The basics of “8-bit unsigned add” and similar are possible to abstract over, but for a lot of cases, you may have to switch your entire algorithm around between different CPUs to get reasonable performance (or even gain over the scalar code at all). There's no way a compiler or SIMD abstraction library will do that for you.
- Miniminix 2y agoRe: SIMD Suggest you look at the Julia Language, a high-level but still capable of C-like speed. It has built in support for SIMD (and GPU) processing. Julia is designed to support Scientific Computing, with a growing library spanning different domains. https://docs.julialang.org/en/v1/ https://docs.julialang.org/en/v1/
- deleted 2y ago[deleted]