9 ms·
To define "new" more precisely, there is an "Enhanced REP MOVSB" flag in cpuid that tells you you can use those. It is set from Ivy Bridge and Zen 3 up. It's s
by floatboth 5y ago
To define "new" more precisely, there is an "Enhanced REP MOVSB" flag in cpuid that tells you you can use those. It is set from Ivy Bridge and Zen 3 up.
It's still not always faster: https://stackoverflow.com/questions/43343231/enhanced-rep-movsb-for-memcpy https://stackoverflow.com/questions/43343231/enhanced-rep-mo...
- ckennelly 5y agoAs mentioned in that Stack Overflow post, though, things change again with FSRM (Fast Short Rep Mov). While there are still startup costs, the overhead of calling a function (especially via a PLT) and incurring instruction cache misses is hard to demonstrate in a microbenchmark, while rep movsb encodes more compactly than many flavors of call. In an actual application though, the "slower" but smaller implementation can often win (https://research.google/pubs/pub50338.pdf https://research.google/pubs/pub50338.pdf and https://research.google/pubs/pub48320.pdf https://research.google/pubs/pub48320.pdf)