3 ms·
I'm in a hurry right now, but I'll link to it later. It's on a i7-4750HQ, it was about 10% faster as measured with rdtsc and looping it a couple of million time
by Coding_Cat 11y ago
I'm in a hurry right now, but I'll link to it later. It's on a i7-4750HQ, it was about 10% faster as measured with rdtsc and looping it a couple of million times.
Granted, Intel's implementation was for 8x8 (floats), perhaps that makes a difference in the instruction pipelining. I'll see if it does later.