3 ms·
Why not tools like https://github.com/ispc https://github.com/ispc? This seems really close to the metal, either to have a non-negligible maintenance cost or n
by xpuente 5y ago
Why not tools like https://github.com/ispc https://github.com/ispc?
This seems really close to the metal, either to have a non-negligible maintenance cost or not being able to fully exploit the hardware at use.
- Scaevolus 5y agoSwitching compilers is often too high-risk, but there are header-only libraries that get you most of the same benefits with normal C++ and wrappers around the intrinsics: https://github.com/richgel999/CppSPMD_Fast https://github.com/richgel999/CppSPMD_Fast
- vgatherps 5y agoI've used ISPC before, as well as enoki (sort of like ISPC-in-c++), and found that they have a lot of sharp performance edges. My experience with both was that as I moved away from the super classic SIMD cases, the more I ran into crazy compiler cliffs where tiny tweaks would blow up the codegen. In each case I gave up, reimplemented what I wanted directly in c++ (the second time using anger fog's wonderful vector class library), and easily got the results I wanted without a ton of finagling the compiler and libraries.
- bjourne 5y agoIt doesn't always emit optimal SIMD code. Plus, when you get the hang of it, writing your own SIMD library is fairly simple so you don't need a tool for it. C++ templates and operator overloading really shines here. For example, you can write sqrt(x*y+z) and have the the template system select the most optimal SIMD intrinsics depending on whether x, y, and z are int, float, int16, float8, double4, etc.
- janwas 5y ago+1 to intrinsics or wrappers giving us more control over performance. > Plus, when you get the hang of it, writing your own SIMD library is fairly simple hm.. it's indeed easy to start, but maintaining https://github.com/google/highway https://github.com/google/highway (supports clang/gcc/MSVC, x86/ARM/RiscV) is quite time-consuming, especially working around compiler bugs.
- Const-me 5y agoHarder to use. That’s another language which requires that special compiler from Intel. The intrinsics are already supported in all modern C and C++ compilers, with little to no project setup. For many practical problems, the ISPC’s abstraction is not a good fit. It’s good for linear algebra with long vectors and large matrices, but SIMD is useful for many other things besides that. A toy problem: compute count of spaces in a 4 GB-long buffer in memory. I’m pretty sure manually written SSE2 or AVX2 code (inner loop doing _mm_cmpeq_epi8 and _mm_sub_epi8, outer one doing _mm_sad_epu8 and _mm_add_epi64) gonna be faster than ISPC-made version.
- mattpharr 5y ago> It’s good for linear algebra with long vectors and large matrices, but SIMD is useful for many other things besides that The main goal in ispc's design was to support SPMD (single program multiple data) programming, which is more general than pure SIMD. Handling the relatively easy cases of (dense) linear algebra that are easily expressed in SIMD wasn't a focus as it's pretty easy to do in other ways. Rather, ispc is focused on making it easy to write code with divergent control flow over the vector lanes. This is especially painful to do in intrinsics, especially in the presence of nested divergent control flow. If you don't have that, you might as well use explicit SIMD, though perhaps via something like Eigen in order to avoid all of the ugliness of manual use of intrinsics. > I’m pretty sure manually written SSE2 or AVX2 code (inner loop doing _mm_cmpeq_epi8 and _mm_sub_epi8, outer one doing _mm_sad_epu8 and _mm_add_epi64) ispc is focused on 32-byte datatypes, so I'm sure that is true. I suspect it would be a more pleasant experience than intrinsics for a reduction operation of that sort over 32-bit datatypes, however.
- Const-me 5y ago> This is especially painful to do in intrinsics Depends on use case, but yes, can be complicated due to lack of support in hardware. I’ve heard AVX512 fixed that to an extent, but I don’t have experience with that tech. > perhaps via something like Eigen I do, but sometimes I can outperform it substantially. It’s optimized for large vectors. In some cases, intrinsics can be faster, and in my line of work I encounter a lot of these cases. Very small matrices like 3x3 and 4x4 fit completely in registers. Larger square matrices of size like 8 or 24, and tall matrices with small fixed count of columns, don’t fit there but a complete row does, saving a lot of RAM latency when dealing with them. > to avoid all of the ugliness of manual use of intrinsics I don’t believe they are ugly; I think they just have a steep learning curve. > I suspect it would be a more pleasant experience than intrinsics for a reduction operation of that sort over 32-bit datatypes Here’s an example how to compute FP32 dot product with intrinsics: https://stackoverflow.com/a/59495197/126995 https://stackoverflow.com/a/59495197/126995 I have doubts the ISPC’s reduction gonna result in similar code. Even clang’s automatic vectorizer (which I have a high opinion of) is not doing that kind of stuff with multiple independent accumulators.