22 ms·
This micro benchmark measures the direct cost of system calls. But flushing the TLB and trashing the caches carries an indirect cost too. According to some ben
by exDM69 3y ago
This micro benchmark measures the direct cost of system calls. But flushing the TLB and trashing the caches carries an indirect cost too.
According to some benchmarks I've seen (sorry no link), a system call in the middle of a memory heavy inner loop can impact performance for 100 us before performance is back to steady state. This is 100x longer than the system call alone. Of course there is work being done, but at a reduced throughput. This is in line with my practical experience from working with performance sensitive code.
These things are difficult to benchmark reliably and draw actionable conclusions from the results.
This is not a criticism to the author of this article, the article clearly describes the methodology of the benchmarks and does not suggest any wrong conclusions from the data.
- sapiogram 3y ago> According to some benchmarks I've seen (sorry no link), a system call in the middle of a memory heavy inner loop can impact performance for 100 us before performance is back to steady state. This is 100x longer than the system call alone. Do you some intuition of what causes this? I don't have how much experience with this kind of work, but 100µs is enough time to do hundreds of random memory accesses. How can a single syscall do so much damage to the cache?
- Karellen 3y agoAs the GP mentioned, it's TLB cache invalidation that can be the problem. Reloading the mappings from virtual (per-process) addresses to physical addresses after returning from the kernel can cause delays (on some architectures), even if most of the physical memory is still in the cache and valid. https://en.wikipedia.org/wiki/Translation_lookaside_buffer https://en.wikipedia.org/wiki/Translation_lookaside_buffer (Also, worth pointing out that they didn't claim a delay of 100µs, just that some (presumably, much smaller, on the order of ns?) delays can show up up to 100µs later before "steady state" is fully restored.)
- gpderetta 3y agoAs syscalls run more code and access more data they can increase TLB pressure in the same way the increase general cache pressure, but syscalls per-se (outside of things like munmap) don't typically invalidate the TLB as the mapping is not changed on an user-space kernel-space transiation, only the access right (there were exceptions like the brief period when people where running with full 4GB user address space on 32bits cpu).
- panzi 3y ago> but syscalls per-se (outside of things like munmap) don't typically invalidate the TLB as the mapping is not changed on an user-space kernel-space transiation, only the access right Isn't that exactly what changed for the Meltdown/Spectre mitigations? That it is invalidated now?
- gpderetta 3y agoI think you are referring to KTPI, but a) it might not be needed on recent CPUs that had meltdown-like issues patched, and b) in any case on CPUs that can tag pages with process identifiers (most of them these days), it can avoid TLB flushes most of the time. But I don't claim any specific knowledge of on these mitigations.
- exDM69 3y agoTLB Flushing and dcache/icache evicting hot cache lines in favor of the instructions and data needed by kernel mode. It takes a while of normal operation until these are populated again with the hot data, until then the system overall throughput is reduced.
- sujayakar 3y agoIt's a bit old and doesn't include recent microarchitectural changes, but Section 2 of the FlexSC paper from 2010 (https://www.usenix.org/legacy/event/osdi10/tech/full_papers/Soares.pdf https://www.usenix.org/legacy/event/osdi10/tech/full_papers/...) has a detailed discussion of these indirect effects. I especially like how they quantify the indirect effects by measuring user code's IPC after the syscall.
- exDM69 3y agoI think that's the benchmarks I allude to in the GP post. Table 1 on page 3 is absolute gold, it quantifies the indirect costs by listing the number of cache lines and TLB entries evicted. The numbers are much larger than I remembered. According to the table, the simplest syscall tested (stat) will evict 32 icache lines (L1), a few hundred dcache lines (L1), hundreds of L2 lines and thousands of L3 lines, and about twenty TLB entries. After returning from said syscalls, you'll pay a cache miss for every line evicted. Also worth noting that inside the syscall, the instructions per clock (IPC) is less than 0.5. When the CPU is happy, you generally see IPC figures around 2 to 3.
- cb321 3y agoYeah.. FlexSC / Soares is my favorite paper from OSDI 2010. The system call batching with "looped" multi-call they mention there relates to the roughly 30 line (not actually looping) assembly language in my other comment here (https://news.ycombinator.com/item?id=39189135 https://news.ycombinator.com/item?id=39189135) and in a few ways pre-saged io_uring work. Anyway, a 20-line example of a program written against said interpreter is https://github.com/c-blake/batch/blob/1201eefc92da9121405b793cc571e3d09643eb0c/examples/total.c#L9 https://github.com/c-blake/batch/blob/1201eefc92da9121405b79... but that only needs the wdcpy fake syscall not the conditional jump forward (although that could/should be added if the open can succeed but the mmap can fail and you want the close clean-up also included in the batch, etc., etc.). I believe Cassyopia (also mentioned in Soares) hoped to be able to analyze code in user-space with compiler techniques to automagically generate such programs, but I don't know that the work ever got beyond a HotOS paper (i.e. the kinda hopes & dreams stage) and it was never clear how fancy the anticipated batches being. The Xen/VMware multi-calls Soares2010 also mentions do not seem to have inline copy/jumps, though I'd be pretty surprised if that little kernel module is the only example of it.