3 ms·
I don't understand "ERMS" and "FSRM" and there seems to be nothing good on google about it. Are these just CPUID flags that tell you that you can use a rep mov
by atesti 3y ago
I don't understand "ERMS" and "FSRM" and there seems to be nothing good on google about it.
Are these just CPUID flags that tell you that you can use a rep movsb for maximum performance instead of optimized SSE memcpy implementations? Or is it a special encoding/prefix for rep movsb to make it faster? In case of the later, why would that be necessary? How does one make use of fsrm?
- tommiegannert 3y agoFound this [1], which also links to the Intel Optimization Manual [2]. Seems like ERMS was a cheaper replacement for AVX and FSRM was a better version, for shorter blocks. > Cheapest versions of later processors - Kaby Lake Celeron and Pentium, released in 2017, don't have AVX that could have been used for fast memory copy, but still have the Enhanced REP MOVSB. And some of Intel's mobile and low-power architectures released in 2018 and onwards, which were not based on SkyLake, copy about twice more bytes per CPU cycle with REP MOVSB than previous generations of microarchitectures. > Enhanced REP MOVSB (ERMSB) before the Ice Lake microarchitecture with Fast Short REP MOV (FSRM) was only faster than AVX copy or general-use register copy if the block size is at least 256 bytes. For the blocks below 64 bytes, it was much slower, because there is a high internal startup in ERMSB - about 35 cycles. The FSRM feature intended blocks before 128 bytes also be quick. [1] https://stackoverflow.com/a/43837564 https://stackoverflow.com/a/43837564 [2] http://www.intel.com/content/dam/www/public/us/en/documents/manuals/64-ia-32-architectures-optimization-manual.pdf http://www.intel.com/content/dam/www/public/us/en/documents/...
- ithkuil 3y agoFSRM is just the name of a cpu optimization that affects existing code. Choosing an optimal instruction choice and scheduling can be done statically during compile time or dynamically (via chosing one of several library functions at runtime, or jitting). In order to be able to detect which is the optimal instruction scheduling at runtime you need to know the actual CPU. You could have a table of all cpu models or you could just ask your OS whether the CPU you run on has that optimization implemented. Linux had to be patched so that it can _report_ that a CPU does implement that optimization. https://www.phoronix.com/news/Intel-5.6-FSRM-Memmove https://www.phoronix.com/news/Intel-5.6-FSRM-Memmove
- rwmj 3y agoThe flags just tell you that, on this CPU, rep movsb is fast so you don't need to use an SSE/AVX-optimized implementation.