5 ms·
> Have you used ISPC No professional kernel writer uses Auto-vectorization. > I feel it's a bit ridiculous that in this day and age you have to write SIMD cod
by almostgotcaught 1y ago
> Have you used ISPC
No professional kernel writer uses Auto-vectorization.
> I feel it's a bit ridiculous that in this day and age you have to write SIMD code by hand
You feel it's ridiculous because you've been sold a myth/lie (abstraction). In reality the details have always mattered.
- CyberDildonics 1y agoISPC is a lot different from C++ compiler auto vectorization and it works extremely well. Have you tried it or not? If so where does it actually fall down? It warns you when doing slow stuff like gathers and scatters.
- camel-cdr 1y agoIs it a lot different from autovec with #pragma omp simd? I played around with ISPC a bit and it didn't seem much different.
- CyberDildonics 1y agoI don't have any experience with that aspect of openmp. When you use ISPC are you using the varying and uniform keywords? You can write something that is almost C but it basically forced to vectorize. If you both are vectorizing the same thing there might not be much difference. If both are not vectorizing there might not be much difference, but with ISPC you can easily make sure that it does use vectorization and the best instruction set for your CPU.
- rerdavies 1y agoI'm hard pressed to think of a kernel function that would benefit from auto-vectorization.
- aa-jv 1y agoA good example is a vector addition kernel, which is simple, embarrassingly parallel, and well-suited for SIMD: void vector_add(float *a, float *b, float *c, int n) { for (int i = 0; i < n; i++) { c[i] = a[i] + b[i]; } }
- rerdavies 1y agoOK. And which kernel function does that? Serious question.
- rerdavies 1y agoOk. Miscommunication. Different kernels. I was actually asking "which Linux kernel functions are going to benefit from auto-vectorization. There are not a lot of Linux kernel functions that take arguments that are two arrays of floating point values. Lots of string stuff that typically doesn't vectorize well (except in freakish cases), perhaps. But very very few functions that take arrays of floats as arguments. Yes, graphics libraries if you want to count those as kernel functions; but those primarily concern themselves with passing arrays of floats to co-processors, to be vectorized by a GPU (different problem). As an interesting point of reference, the last time I did Windows kernel development (which was admittedly not recently), code running in ring 0 was not allowed to access SIMD registers because they weren't saved and loaded during kernel-code context switches. Not sure if that's the case, but it probably is. Context switches are a LOT faster if you don't have to save and load SIMD registers.
- Const-me 1y agoIndeed, automatic vectorizers do such simple things pretty reliably these days. However, if you build your software from kernels like that, you leave a lot of performance on the table. For example, each core of my Zen 4 CPU at base frequency can add FP32 numbers with AVX1 or AVX-512 at 268 GB/sec, which results in 806 GB/sec total bandwidth for your kernel with two inputs and 1 output. However, dual-channel DDR5 memory in my computer can only deliver 83 GB/sec bandwidth shared across all CPU cores. That’s an order of magnitude difference for a single threaded program, and almost 2 orders of magnitude difference when computing something on the complete CPU. Even worse, the difference between compute and memory widens over time. The next generation Zen 5 CPUs can add numbers twice as fast per cycle if using AVX-512. For this reason, ideally you want your kernels to do much more work with the numbers loaded from memory. That’s why efficient compute kernels are often way more complicated than the for loop in your example. Sadly, seems modern compilers can only reliably autovectorize very simple loops.