3 ms·
> the rep prefixed instructions for string operations ... are actually preferred over a hand-written vectorized loop on Ivy Bridge and up (see [1] section 3.7.
by cfallin 10y ago
> the rep prefixed instructions for string operations
... are actually preferred over a hand-written vectorized loop on Ivy Bridge and up (see [1] section 3.7.7, "Enhanced REP MOVSB and STOSB operation (ERMSB)"). It's indicated by a CPUID feature flag bit (edit: grep for "erms" in /proc/cpuinfo to see this).
The reason is that microcode knows more about the dcache microachitecture, load/store units, special features (weak ordering with fence at end? [2]), etc., than you do, and can be specially optimized for the particular design. There's a slight cost to transitioning to microcode and back (to ordinary hardware-decoded instructions), of course, so for small operations it might not be a win.
The section cited above shows ERMSB as ~break-even vs. 128 bit AVX on Ivy Bridge from 128 bytes up to 2KB, and about 2% faster above that.
[1] Intel 64 and IA-32 Architectures Optimization Reference Manual (order #248966-026), April 2012. http://www.intel.com/content/dam/doc/manual/64-ia-32-architectures-optimization-manual.pdf http://www.intel.com/content/dam/doc/manual/64-ia-32-archite...
[2] http://stackoverflow.com/questions/33480999/how-can-the-rep-stosb-instruction-execute-faster-than-the-equivalent-loop http://stackoverflow.com/questions/33480999/how-can-the-rep-...
- olegolegovich 10y agoDo you have benchmarks, supporting "~break-even vs. 128 bit AVX on Ivy Bridge from 128 bytes up to 2KB" as I have not found it to be the case at least on Haskell (my benchmarking code is @ https://bitbucket.org/olegoandreev/scratch/src/dd7ab9008c59c8254dd5bdb4becae20457214c3f/memcpy-test/?at=default https://bitbucket.org/olegoandreev/scratch/src/dd7ab9008c59c...).
- cfallin 10y agoI was just quoting the cited PDF ("the section cited above shows..."), which the author(s) did on Ivy Bridge. I haven't benchmarked it myself. It would be interesting to see how the most common cores today do on this...
- sounds 10y agoAnd by not having a group of hand-optimized special cases, the REP MOVSB and REP STOSB can be inlined saving instruction cache.