3 ms·
If you need the speed of this PRNG, you probably don't want the overhead of calling a library function. Instead you would implement it in place with AVX-512 or
by pps43 7y ago
If you need the speed of this PRNG, you probably don't want the overhead of calling a library function. Instead you would implement it in place with AVX-512 or equivalent for the platform you're running on.
- nn3 7y agoFunction calls are very cheap on modern CPUs. The same ILP argument the author makes applies in most cases. And both vectorization and inlining works fine with a function with modern tool chains that do LTO.