4 ms·
If my process has async/await and uses epoll, wouldn't this have about the same performance as io_uring? E.g. with io_uring, the event triggers the kernel to r
by justsomeuser 5y ago
If my process has async/await and uses epoll, wouldn't this have about the same performance as io_uring?
E.g. with io_uring, the event triggers the kernel to read into a buffer.
With async/await epoll wakes up my process which does the syscall to read the file.
In both cases you still need to read from the device and get the data to the user process?
- Misdicorl 5y agoImagine your process instead of getting woken up to make a syscall gets woken up with a pointer to a filled data buffer.
- justsomeuser 5y agoBut some process (kernel or user space) still needs to spend the computers finite resources to read it? If the work happens in the kernel process or user process it still costs the same?
- Misdicorl 5y agoYes, the syscall work still happens and the raw work has not changed. The overhead has dropped dramatically though since you've eliminated at least 2 context switches per data fill cycle(and probably more).
- wtallis 5y agoThe transition from userspace to the kernel and back takes a similar amount of time to actually doing a small read from the fastest SSDs, or issuing a write to most relatively fast SSDs. So avoiding one extra syscall is a meaningful performance improvement even if you can't always eliminate a memcpy operation.
- Veserv 5y agoThat sounds unlikely. syscall hardware overhead (entry + exit) on a modern x86 is only on the order of 100 ns which is approximately main memory access latency. I am not familiar with Linux internals or what the fastest SSDs are capable of these days, but I am fairly sure that for your statement to be true Linux would need to be adding 1 to 2 orders of magnitude in software overhead. This occurs in the context switch pathway due to scheduling decisions, but it is fairly unlikely it occurs in the syscall pathway unless they are doing something horribly wrong.
- vlovich123 5y agoMy understanding is that's the naiive measurement of the cost of just the syscall operation (i.e. if you measure issue to kernel is executing). Does this actually account for the performance loss of cache innefficiency? If I'm not mistaken at a minimum the CPU needs to flush various caches to enter the kernel, fill them up as the kernel is executing, & then repopulate them when executing back in userspace. In that case (even if it's not a full flush), you have a hard to measure slowdown on the code processing the request in the kernel & in userspace after the syscall because the locality assumptions that caches rely on are invalidated. With an io_uring model, since there's no context switches, temporal & spatial locality should provide an outsized benefit beyond just removing the syscall itself. Additionally, as noted elsewhere, you can chain syscalls pretty deeply so that the entire operation occurs in the kernel & never schedules your process. This also benefits spatial & temporal locality AND removes the cost of needing to schedule the process in the first place.
- wtallis 5y agoI overstated things. Sorry. A syscall that does literally nothing can still have a latency as observed from userspace that is vastly faster than the hardware latency of a fast SSD, though Spectre and friends have slowed this down a bit. However, when comparing the latency as observed from userspace of syscalls that actually do significant work and cause stuff to happen, such as shepherding an IO request through the various layers to get down to the actually submitting commands to the hardware, then you usually do start talking about latencies that are on a similar order of magnitude to fast SSDs. In particular, if you're trying to minimize IO latency, the kernel can no longer afford to put a CPU core to sleep or switch to run another task while waiting for the IRQ signalling that the SSD is done. The interrupt latency followed by the context switch back to the kernel thread that handles the IO completion is a substantial delay compared to spinning in that thread until the SSD responds.
- dundarious 5y agoNot necessarily, as you can use io_uring in a zero-copy way, avoiding copies of network/disk data from kernel space to user space.
- Matthias247 5y agoThe answer is "go ahead and benchmark". There are some theoretical advantages to uring, like having to do less syscalls and have less userspace <-> kernelspace transitions. In practice implementation differences can offset that advantage, or it just might not make any difference for the application at all since its not the hotspot. I think for socket IO a variety of people did some synthetic benchmarks for epoll vs uring, and got all kinds of results from either one being a bit faster to both being roughly the same.
- yxhuvud 5y agoIt also depends heavily on your kernel version. There are really wide swings.