3 ms·
Hey, thanks for the comment! I did actually take a look at the aes instructions and came to the conclusion that they are in fact faster than the hash I used, b
by jesse__ 1y ago
Hey, thanks for the comment! I did actually take a look at the aes instructions and came to the conclusion that they are in fact faster than the hash I used, but I think I'd decided I would have to swizzle the data in a way that was a pain because of how the aes mixdown works (ie it mixes across lanes, so I would have to change the output pattern, if that makeshifts sense)
Maybe I'll dust it off one day and try again. That seems like it could be an easy win.
- dragontamer 1y ago> but I think I'd decided I would have to swizzle the data in a way that was a pain because of how the aes mixdown works (ie it mixes across lanes, so I would have to change the output pattern, if that makeshifts sense) 100% agree. I solved this with a 64-bit x2 SIMD add instruction. State += 0x0305071113171923, which ensures a 1-bit 'carry bit' dependency as well so we have (barely) enough data mixing for lots of cool entropy effects. Because this is an odd number (bottom bit is 1), it cycles every 2^64, which should be a sufficient cycle length for most simulations. That 1-bit difference was enough to then pass PractRand and BigCrush. Don't swizzle the bits. Just add a number across all 128-bits (as 2x 64-bit adds) and bam. We get a lot of lovely RNG properties thanks to AES mixing. It's not 'purely' aesenc. I did a few little tidbits that fixed all the problems of AES data mixing. ------ The real fun part is that the latency/dependency limitation on my code is this Add instruction. The AES stuff is done in parallel later and thus easily parallelizes to modern 4x512-bit AES as is available on Zen5. (Maybe the compilers won't see it yet, but it's bloody obvious for humans to see it IMO). IE: the critical path of my code is: simd-add state, 0x030507..... State gets SSA'd by the out of order system on the processors and thus future iterations of the RNG loop can execute in parallel.
- jesse__ 1y agoThat sounds really great. I'm away from the office for the week but when I'm back I'll take a closer look and maybe squeeze some more juice out of it :)