3 ms·
Does anyone know any more sources that would be useful to learn more about SIMD intrinsics and how to use them? I've tried a few times (with moderate success)
by dialecticDolt 6y ago
Does anyone know any more sources that would be useful to learn more about SIMD intrinsics and how to use them?
I've tried a few times (with moderate success) to parse through source-code using them but I've never found a good introduction.
- QShift 6y agoHandmade Hero is one of the best learning resources on SIMD (and, in my opinion, programming in general). Here's the episode where SIMD is properly introduced: https://guide.handmadehero.org/code/day115 https://guide.handmadehero.org/code/day115
- Const-me 6y agoIf you’re interested in AMD64 SIMD, deep inside that article on SO there’s a link to my earlier article on the subject: http://const.me/articles/simd/simd.pdf http://const.me/articles/simd/simd.pdf I also made offline docs, there: https://github.com/Const-me/IntelIntrinsics https://github.com/Const-me/IntelIntrinsics
- dragontamer 6y agoHonestly, most SIMD intrinsics are pretty simple and straightforward to use. You can probably get 90% of the way there by simply browsing Intel's intrinsic reference manual. https://software.intel.com/sites/landingpage/IntrinsicsGuide/ https://software.intel.com/sites/landingpage/IntrinsicsGuide... Intel's intrinsics guide has an approximate clock-count on all instructions. There are a few obscure details to memorize, but if you already understand super-scalar, pipelined, and OoO operations of modern CPUs, its all straight-forward. --------- The last 10% of using SIMD successfully is obscure and eclectic. Its not "difficult", its just a whole bunch of tricks that you'll probably never figure out on your own. I'm talking about pshufb magic, prefix sum / scan patterns, gather/scatter patterns, and the like. As complicated as pshufb is, the Intel intrinsic documentation does a great job precisely describing the operation: https://software.intel.com/sites/landingpage/IntrinsicsGuide/#text=pshufb&expand=5194,5153,5153 https://software.intel.com/sites/landingpage/IntrinsicsGuide... _mm_shuffle_epi8 gives you a register-to-register (no RAM touched, faster than even L1 cache) 16-byte lookup table. An alternative viewpoint: it gives you a register-to-register arbitrary gather command. I've seen so many tricks from this one instruction alone. --------- Another straightforward guide is Intel's optimization manual. Its straightforward, but a lot of material to read: https://www.intel.com/content/dam/www/public/us/en/documents/manuals/64-ia-32-architectures-optimization-manual.pdf https://www.intel.com/content/dam/www/public/us/en/documents... Chapters 4, 5, and 6 deal with SIMD coding.
- OnlyOneCannolo 6y agoThis [1] is from a MOOC on programming high-performance dense linear algebra, which is part of a series [2] of MOOCs on linear algebra. It's nice because there are free online books, videos, and code. Also it's self-paced and you don't have to register. [1] http://www.cs.utexas.edu/users/flame/laff/pfhp/week2-optimizing-the-micro-kernel.html http://www.cs.utexas.edu/users/flame/laff/pfhp/week2-optimiz... [2] http://ulaff.net/ http://ulaff.net/