3 ms·
Interestingly the std::array::fill member function is identical in the case of int or char, I suppose because there's only one overload of fill and it has to ta
by thestoicattack 7y ago
Interestingly the std::array::fill member function is identical in the case of int or char, I suppose because there's only one overload of fill and it has to take the element type. No idea if the generated stosq is as fast as built-in memset: https://godbolt.org/z/4iYGup https://godbolt.org/z/4iYGup
- leeter 7y agoIIRC it's very CPU dependent. I remember running into a similar issue with memcpy vs. memmove where the latter was actually faster on some CPUs because it used stosb instead of AVX/SSE and some intel CPUs could make that really go fast. TL;DR: benchmark for your specific hardware if it really matters.
- BeeOnRope 7y agoRecent glibc seems to use 'rep stosb' for largish regions of memset. At least for the numbers I give in this post, the ~30 bytes/cycle is actually coming from rep stosb inside memset. The q variants (as opposed to b) are a bit of a grey area: Intel has this REP ERMBS thing [1] that promises fast performance for rep movsb and rep movsb specifically, but not for the w, d or q variants. However, I think all Intel hardware that has implemented it has implemented the d and q variants just as fast. It would be good to verify it though... --- [1] https://stackoverflow.com/q/43343231 https://stackoverflow.com/q/43343231