3 ms·
Interesting, I started out thinking along these lines, but once I figured out I could use PEXT, I just went with that. I think this approach needs some tweaks,
by zwegner 8y ago
Interesting, I started out thinking along these lines, but once I figured out I could use PEXT, I just went with that.
I think this approach needs some tweaks, though. Mainly that the vpermb at the end is the inverse of what we want--the bytes at dense indices get spread out to the sparse indices (it works analogously to gather, but we want scatter). I can't think of a way around this right now...
That said, it's an interesting approach. I think the PEXTs would be the bottleneck in my code (looks like there's only one execution unit for them, whereas there's two for the VPADDs), and finding a way to parallelize all the VPADDs could lead to a nice speedup.
- dragontamer 8y agoYou're right. I did a brief look through AVX512 instructions to look for a solution, and unfortunatley, it seems we both may have been overthinking this. vpcompressb more or less does the job in one instruction. Agner Fog doesn't have a latency listed however. --------- My search methodology was basically this: https://software.intel.com/sites/landingpage/IntrinsicsGuide/#cats=Swizzle&expand=1219,1227,1229&text=__m512i https://software.intel.com/sites/landingpage/IntrinsicsGuide... Search for __m512i (integer-based ZMM registers), with the category "swizzle" (which includes permute, insert, and other such instructions). I figure any potential AVX512 instruction would be a "Swizzle" style instruction. ------- Note: I originally responded to the wrong location in this thread. I copy/pasted my text to here, which is where I originally intended to respond.
- nkurz 8y agoDo you recall which machines VPCOMPRESSB works on? I think it's next generation Icelake? Or is it there already on Cannonlake? And along the same lines, is there a good general way of looking this up? Coincidentally, searching for this, I found Geoff Langdale's blog post where in addition to describing VPCOMPRESSB as 'dynamiting the trout stream', he also describes something very close to zwegner's PEXT approach: https://branchfree.org/2018/05/22/bits-to-indexes-in-bmi2-and-avx-512/ https://branchfree.org/2018/05/22/bits-to-indexes-in-bmi2-an...
- BeeOnRope 8y agoIt's not in Cannonlake (nor the W variants). The D and Q versions are in SKX though and they are 4L2T IIRC. You need a CPU with VBMI2 for the B variant, can't remember off the top of my head if Icelake has that.
- zwegner 8y agoOh sweet! That's an awesome instruction. I'd imagine that would be useful for lots of things. I believe I've seen vcompressd before, but totally forgot about it. Unfortunately it looks like the byte-wise version is part of AVX512-VBMI2, which won't be out until Ice Lake...
- wmu 8y agoYou might have seen vcompessd in context of sorting; I used it for partition part in qsort.
- zwegner 8y agoIt actually would've been during my time at Intel working on the graphics stack for Larrabee, in the 2010-2011 timeframe--vcompressd was part of LRBNI. I was mainly doing infrastructure/compiler/optimization type work, and not much graphics stuff, so I can't recall using the instruction personally, but pretty sure it was used in various places around the stack.