3 ms·
Something I ~never see mentioned in mmap discussions is that it completely bypasses the expensive syscall transition for read calls, which has only grown rapidl
by ComputerGuru 3y ago
Something I ~never see mentioned in mmap discussions is that it completely bypasses the expensive syscall transition for read calls, which has only grown rapidly in cost thanks to all the speculative execution mitigations.
In benchmarking I found an mmap approach to be integral to achieving maximum performance in IO-heavy workloads on modern OSes [0]. (Note that it still is a win even with mitigations disabled.)
[0]: https://neosmart.net/blog/using-simd-acceleration-in-rust-to-create-the-worlds-fastest-tac/ https://neosmart.net/blog/using-simd-acceleration-in-rust-to...
- lvlabguy 3y agoThen you have page faults to deal with.
- deleted 3y ago[deleted]
- hnlmorg 3y agoUnless I’m completely missing your point, and to be fair I might be because it’s late in Europe and I should have gone to bed hours ago, read speeds due to bypassing expensive syscall overhead is an often cited part of the mmap discussion.
- ComputerGuru 3y agoSorry, I meant specifically the post-mitigations aspect. It really does change the calculus.
- lathiat 3y agoYou can likely be solve that with io_uring
- hinkley 3y agoIf it's not solved by io_uring, then what is io_uring for?
- hyc_symas 3y agoYeah, we knew that LMDB would be unaffected. Other folks demonstrated that too. https://www.pugetsystems.com/labs/hpc/Intel-CPU-flaw-kernel-patch-effects---GPU-compute-Tensorflow-Caffe-and-LMDB-database-creation-1093/ https://www.pugetsystems.com/labs/hpc/Intel-CPU-flaw-kernel-...