3 ms·
On my system the naive approach + mmap + gcc-7.2 (-O3 -march=native) gets 0.126s if the file is in cache, so 8GB/s. The same code on memory instead of a file g
by tpolzer 9y ago
On my system the naive approach + mmap + gcc-7.2 (-O3 -march=native) gets 0.126s if the file is in cache, so 8GB/s.
The same code on memory instead of a file gives me around 14GB/s. Maximum I can achieve with threads and/or intrinsics is 18GB/s (theoretical maximum for my dual channel memory is 20GB/s). This is pretty much a memory bounded problem.
I assume the author was using an older gcc version (as mentioned, autovectorization is not really a solved problem), and a lot of time in this scenario is going into file I/O:
real 0m0.126s
user 0m0.071s
sys 0m0.056s