4 ms·
What scaling are you referring to when talking about "linear"? Did you really get 2x absolute performance going from 128-bit to 256-bit SIMD (regardless of the
by slizard 9y ago
What scaling are you referring to when talking about "linear"? Did you really get 2x absolute performance going from 128-bit to 256-bit SIMD (regardless of the uarch, e.g. both with SBE and HSW with the former having a relatively poor cache performance). I'd be surprised, but if your code is in the >>O(10) flops/byte regime [1] and especially friendly instruction stream too. Otherwise, I'm skeptical.
Putting aside the "scaled almost linearly" statement, I'm not surprised that AVX-512 did not give the expected benefits ootb, you suddenly need double the amount of data loaded into the registers for every 512-bit instruction. You'd also quite like want to make sure masking is used effectively. [2]
[1] http://people.eecs.berkeley.edu/~kubitron/cs258/lectures/lec12-Merrimac.pdf http://people.eecs.berkeley.edu/~kubitron/cs258/lectures/lec...
[2] https://software.intel.com/en-us/node/523777 https://software.intel.com/en-us/node/523777