3 ms·
I found a figure of 0.4 cpb for xoshiro256++. AES-128-CTR using AES-NI runs at over 10 GB/s per core on a 3900X. That's approximately the same throughput. Of co
by blattimwind 7y ago
I found a figure of 0.4 cpb for xoshiro256++. AES-128-CTR using AES-NI runs at over 10 GB/s per core on a 3900X. That's approximately the same throughput. Of course, performance embedded in a real simulation may differ wildly, I expect xoshiro is much, much friendlier to frequent invocations generating a block each than AES, so to reach that throughput with AES you'd likely need a not-so-small buffer, which eats into L1 and L2 cache budgets.
- pps43 7y agoI suspect that cpb is without using vector operations. AVX-512 lets you run 8 generators in parallel, and you can do even better with CUDA.
- justincormack 7y agoThere are parallel AES ops on new CPUs too though. Not sure if that figure is already using those.
- zokier 7y ago.25 cpb (1cycle/32bit) for xorshift and pcg with avx512 in here: https://lemire.me/blog/2018/06/07/vectorizing-random-number-generators-for-greater-speed-pcg-and-xorshift128-avx-512-edition/ https://lemire.me/blog/2018/06/07/vectorizing-random-number-...