3 ms·
I did some experimentation last night as well. I suspected a lot of the cost came from the unmapping the files and the required invalidations and TLB shootdowns
by prattmic 10y ago
I did some experimentation last night as well. I suspected a lot of the cost came from the unmapping the files and the required invalidations and TLB shootdowns required to do so.
I made rg simply not munmap files when it was done with them (I made this drop do nothing: https://github.com/danburkert/memmap-rs/blob/master/src/unix.rs#L165 https://github.com/danburkert/memmap-rs/blob/master/src/unix...)
Searching for PM_RESUME in the Linux source gave me these results:
--no-mmap: ~400ms
--mmap (with munmap): ~750ms
--mmap (without munmap): ~550ms
So eliding munmap made a big difference, but it was still not enough to beat out reading the files. perf shows that the mmap syscall itself is just too expensive (this is --mmap (with munmap)):
Children Self Command Shared Object Symbol
- 81.88% 0.00% rg rg [.] __rust_try
- __rust_try
- 50.57% std::panicking::try::call::ha112cda315d6c57d
- 47.73% rg::Worker::search_mmap::h5179a76c63e344d0
- 23.91% __GI___munmap
6.08% smp_call_function_many
3.14% rwsem_spin_on_owner
1.86% native_queued_spin_lock_slowpath
0.94% osq_lock
0.67% native_write_msr_safe
0.52% unmap_page_range
+ 21.41% _$LT$rg..search_buffer..BufferSearcher$LT$$u27$a$C$$u20$W$GT$$GT$::run::hd0f8b2830716be0c
0.80% memchr
- 17.99% __mmap64
5.20% rwsem_down_write_failed
1.96% rwsem_spin_on_owner
0.77% osq_lock
0.59% native_queued_spin_lock_slowpath
+ 6.79% 0x1080d
1.72% __GI___libc_close
1.27% __memcpy_sse2_unaligned
0.98% __fxstat64
0.56% __GI___ioctl