4 ms·
A few years back I decided to entertain myself by testing how smart today's smart compilers really are when it comes to auto-vectorization. I had this small an
by YSFEJ4SWJUVU6 9y ago
A few years back I decided to entertain myself by testing how smart today's smart compilers really are when it comes to auto-vectorization.
I had this small and simple C application I'd written years earlier that tried to find inputs whose corresponding MD5 hashes started with certain bytes. It was a good base because it was obviously vectorizable.
At first enabling the vectorizer didn't result in any changes to the binaries. I then correctly guessed that (potentially) calling printf function inside a hot loop might confuse it. After slight refactoring I got the compiler to output SSE instructions, which resulted in a nice 2.5× testing speed over the original (incidentally even without auto-vectorizing the refactored code resulted in faster binaries, which is not all that surprising).
Anyway, I also rewrote the application to use intrinsics. I hadn't used them before myself, but it didn't really take much time at all familiarize myself with them and write the code, and it was indeed quite a bit faster than what the compiler was capable of with resulting binary having 14× speed compared to the original, or over 5× compared to what the compiler could achieve without explicit hints from intrinsics.
Edit: added back a few words I had accidentally removed when rearranging sentences, causing a confusing incomplete sentence. Corrected comparing figures like for like.
- faragon 9y agoThat matches my experience. With auto-vectorization you get some speed-up helping the compiler (not always obvious, often requiring +1 increments, etc.), but for full speed you need to do handwritten SIMD intrinsics. I would like to have at least 50% of the optimal by the compiler, without intrinsics (and using intrinsics for the most critical code).
- foota 9y agoWould be interesting to see the difference in the compiled assembly.
- deleted 9y ago[deleted]
- evincarofautumn 9y agoI’ve had similar experiences. If I want vectorised code, I just write it myself using intrinsics or assembly. It’s fine if the compiler can autovectorise something I didn’t feel like doing by hand, but I’m not going to rely on heuristic voodoo to get the machine code I want for a hot loop. I wouldn’t mind a slightly nicer wrapper API for the intrinsics, though, something like glsl-sse2[1]. And that’s more or less what I’m planning to do in a programming language I’m working on, actually—if you use a SIMD-compatible array type, the compiler will try to keep it in a vector register, and some operations will be faster (e.g., “+” on two Float32^4 values will compile to an addps) but it’s up to the programmer to use the instructions they actually want, or tell the compiler with a macro “please vectorise this loop or warn me about why you can’t”. [1] https://github.com/LiraNuna/glsl-sse2 https://github.com/LiraNuna/glsl-sse2
- t0rakka 9y agohttps://github.com/t0rakka/mango https://github.com/t0rakka/mango