3 ms·
I remember that when io_uring was in it's early stages several pointed to kqueue (also nt'wait_for_multiple_something or iocp and solaris event ports). By comin
by marcodiego 2y ago
I remember that when io_uring was in it's early stages several pointed to kqueue (also nt'wait_for_multiple_something or iocp and solaris event ports). By coming later, I think io_uring was able to better fit modern reality and also avoid problems from previous implementations.
Hope security issues with it are solved and usage becomes mostly transparent through userspace libs, it looks like a high performance strategy for our current computing hardware.
- petesergeant 2y ago> I think io_uring was able to better fit modern reality Genuinely would be fascinated to understand what that means
- nesarkvechnep 2y agoI'm wondering what has fundamentally changed in computing between the creation of kqueue and io_uring so the latter is able to better fit modern reality.
- vacuity 2y agoI don't know that there were relevant changes between the advent of each, so much as kqueue just didn't have existing features in mind. I assume GP is referring to the ring buffer design and/or completion-based processing, as part of the ability to batch syscall processing. This is reminiscent of external work like FlexSC and can be viewed as mechanical sympathy.
- deleted 2y ago[deleted]
- mananaysiempre 2y agoI don’t know which things here are actually relevant to the differences between the two, but of course there have been changes. Core counts are much higher. You can stripe some enterprise SSDs and get bandwidth within an order of magnitude or so of your RAM. Yet clocks aren’t that much higher, and user-supervisor transitions are comparatively much more expensive. There’s a reason Lemire’s talk on simdjson is called “Data engineering at the speed of your disk”.
- adrian_b 2y agoI have not looked at the io_uring implementation to see if it really has improvements from this point of view, but something that has changed during the quarter of century from the design of kqueue until now is that currently it has become much more important than before to minimize the number of context switches caused by system calls in the programs that desire to reach a high performance for input/output. The reason is that with faster CPU cores, more cores sharing the memory, bigger CPU core states and relatively slower memory in comparison with the CPU cores, the time wasted by a context switch has become relatively greater in comparison with the time used by a CPU core to do useful work. So hopefully, implementing some I/O task using io_uring should require less system calls than when using kqueue or epoll. According to the io_uring man page, this should be true.
- oasisaimlessly 2y agoIn addition to the reasons you listed, context switches have also been significantly slowed down by Meltdown/Spectre speculative execution vulnerability mitigations.
- muststopmyths 2y agoI haven't looked into io_uring except superficially, but Windows implemented Registered I/O in Windows 8 circa 2011. this is the basically the same programming paradigm used in io_uring, except it is sockets-only. Talk here [1] speaks to the modern reality of 14 years ago :-). Since kqueue seems very similar to IOCP in paradigm, I guess some of the overheads are similar and hence a ring-buffer-based I/O system would be more performant. It's worth noting that NVME storage also seems to use a similar I/O pattern as RIO, so I assume we're "closer to the hardware" in this way. 1. https://learn.microsoft.com/en-us/shows/build-build2011/sac-593t https://learn.microsoft.com/en-us/shows/build-build2011/sac-...
- JackSlateur 2y agoio_uring is not limited to networking/sockets io_uring is not even limited to IO
- muststopmyths 2y agoI was talking about RIO
- JackSlateur 2y agoAnd I was talking about the modern reality enabled by iouring. Not really 14yo stuff, to be honest.
- p_ing 2y agoMicrosoft implemented I/O rings in an update to Windows 10 with some differences and it is largely a copy-and-paste of io_uring. It's important to note that the NT kernel was built to leverage async I/O throughout. It was part of the original design documents and not an after-thought. https://learn.microsoft.com/en-us/windows/win32/api/ioringapi/ https://learn.microsoft.com/en-us/windows/win32/api/ioringap... https://windows-internals.com/i-o-rings-when-one-i-o-operation-is-not-enough/ https://windows-internals.com/i-o-rings-when-one-i-o-operati... https://windows-internals.com/ioring-vs-io_uring-a-comparison-of-windows-and-linux-implementations/ https://windows-internals.com/ioring-vs-io_uring-a-compariso... https://www.cs.fsu.edu/~zwang/files/cop4610/Fall2016/windows.pdf https://www.cs.fsu.edu/~zwang/files/cop4610/Fall2016/windows...
- Asmod4n 2y agosyscalls which go from user to kernelspace got more expensive after the mitigations for vulnerabilities in intel and amd CPUs. io_uring solved that.
- Someone 2y ago> io_uring was able to better fit modern reality and also avoid problems from previous implementations. Learning from the past can indeed lead to better designs. > Hope security issues with it are solved So, if I understand that correctly, it also introduced new problems? If so, do we know those new issues are solvable without inventing yet another API? If not, is it really better or just having a different set of issues?
- zamalek 2y agoHere's a recent vulnerability: https://blog.exodusintel.com/2024/03/27/mind-the-patch-gap-exploiting-an-io_uring-vulnerability-in-ubuntu/ https://blog.exodusintel.com/2024/03/27/mind-the-patch-gap-e... It's just your typical multithreading woes (in an unsafe language). One presumes that these problems will be ironed out eventually. Unfortunately it seems as though various hardened distros turn it off for this reason (source was a HN comment I read a while back).