6 ms·
I think this might be the best explanation I've read of why io_uring should be better than epoll since it effectively collapses the 'tell me when this is ready'
by tele_ski 5y ago
I think this might be the best explanation I've read of why io_uring should be better than epoll since it effectively collapses the 'tell me when this is ready' with the 'do action' part. That was the really enlightening part for me.
I have to say though, the name io_uring seems unfortunate and I think the author touches on this in the article... the name is really an implementation detail but io_uring's true purpose is a generic asynchronous syscall facility that is currently tailored towards i/o. syscall_queue or async_queue or something else...? A descriptive api name and not an implementation detail would probably go a long way in helping the feature be easier to understand. Even window's IOCP seems infinitely better named than 'uring'.
- eloff 5y agoI've seen mixed results so far. In theory it should perform better than epoll, but I'm not sure it's quite there yet. The maintainer of uWebSockets tried it with an earlier version and it was slower. Where it really shines is disk IO because we don't have an epoll equivalent there. I imagine it would also be great at network requests that go to or from disk in a simple way because you can chain the syscalls in theory.
- zxzax 5y agoThe main benefit to me has been in cases that previously required a thread pool on top of epoll, it should be safe to get rid of the thread pool now and only use io_uring. Socket I/O of course doesn't really need a thread pool in a lot of cases, but disk I/O does.
- eloff 5y agoI dug up the thread with the benchmarks: https://github.com/axboe/liburing/issues/189 https://github.com/axboe/liburing/issues/189 io_uring does win on more recent versions, but it's not like it blows epoll out of the water - it's an incremental improvement. Specifically for websockets where you have a lot of idle connections, you don't want a bunch of buffers waiting for reads for long periods of time - this is why Go doesn't perform well for websocket servers as calling Read() on a socket requires tying up both a buffer and a goroutine stack. I haven't looked into how registering buffers works with io_uring, if it's 1-to-1 mapping or if you can pass a pool of N buffers that can be used for reads on M sockets where M > N. The details matter for that specific case. Again where io_uring really shines is file IO because there are no good solutions there currently.
- pydry 5y agoI'm still confused coz this is exactly what I always thought the difference between epoll and select was. "what if, instead of the kernel telling us when something is ready for an action to be taken so that we can take it, we tell the kernel what action to we want to take, and it will do it when the conditions become right." The difference between select and epoll was that select would keep checking in until the conditions were right while epoll would send you a message. That was gamechanging. - I'm not really sure why this is seen as such a fundamental change. It's changed from the kernel triggering a callback to... a callback.
- simcop2387 5y agoIo uring can in theory be built to subscribe to any syscall (though it hasn't yet). I don't believe epoll can do things like stat, opening files, closing files, and syncing though.
- coder543 5y agoepoll: tell me when any of these descriptors are ready, then I'll issue another syscall to actually read from that descriptor into a buffer. io_uring: when any of these descriptors are ready, read into any one of these buffers I've preallocated for you, then let me know when it is done. Instead of waking up a process just so it can do the work of calling back into the kernel to have the kernel fill a buffer, io_uring skips that extra syscall altogether. Taking things to the next level, io_uring allows you to chain operations together. You can tell it to read from one socket and write the results into a different socket or directly to a file, and it can do that without waking your process pointlessly at any intermediate stage. A nearby comment also mentioned opening files, and that's cool too. You could issue an entire command sequence to io_uring, then your program can work on other stuff and check on it later, or just go to sleep until everything is done. You could tell the kernel that you want it to open a connection, write a particular buffer that you prepared for it into that connection, then open a specific file on disk, read the response into that file, close the file, then send a prepared buffer as a response to the connection, close the connection, then let you know that it is all done. You just have to prepare two buffers on the frontend, issue the commands (which could require either 1 or 0 syscalls, depending on how you're using io_uring), then do whatever you want. You can even have numerous command sequences under kernel control in parallel, you don't have to issue them one at a time and wait on them to finish before you can issue the next one. With epoll, you have to do every individual step along the way yourself, which involves syscalls, context switches, and potentially more code complexity. Then you realize that epoll doesn't even support file I/O, so you have to mix multiple approaches together to even approximate what io_uring is doing. (Note: I've been looking for an excuse to use io_uring, so I've read a ton about it, but I don't have any practical experience with it yet. But everything I wrote above should be accurate.)
- infogulch 5y agosyscall_uring would be my preference.
- pkghost 5y agoW/r/t a more descriptive name, I disagree—though until fairly recently I would have agreed. I would guess that the desire for something more "descriptive" reflects the fact that you are not in the weeds with io_uring (et al), and as such a name that's tied to specifics of the terrain (io, urings) feels esoteric and unfamiliar. However, to anyone who is an immediate consumer of io_uring or its compatriots, "io" obviously implies "syscall", but is better than "syscall", because it's more specific; since io_uring doens't do anything other than io-related syscalls (there are other kinds of syscalls), naming it "syscall_" would make it harder for its immediate audience to remember what it does. Similarly, "uring" will be familiar to most of the immediate audience, and is better than "queue", because it also communicates some specific features (or performance characteristics? idk, I'm also not in the weeds) of the API that the more generic "_queue" would not. So, while I agree that the name is mildly inscrutable to us distant onlookers, I think it's the right name, and indeed reflects a wise pattern in naming concepts in complex systems. The less ambiguity you introduce at each layer of indirection or reference, the better. I recently did a toy project that has some files named in a similar fashion: `docker/cmd`, which is what the CMD directive in my Dockerfile points at, and `systemd/start`, which is what the ExecStart line of my systemd service file points at. They're mildly inscrutable if you're unfamiliar with either docker or systemd, as they don't really say much about what they do, but this is a naming pattern that I can port to just about any project, and at the same time stop spending energy remembering a unique name for the app's entry point, or the systemd script. Some abstract observations: - naming for grokkability-a-first-glance is at odds with naming for utility-over-time; the former is necessarily more ambiguous - naming for utility over time seems like obviously the better default naming strategy; find a nice spot in your readme for onboarding metaphors and make sure the primary consumers of your name don't have to work harder than necessary to make sense of it - if you find a name inscrutable, perhaps you're just missing some context
- wtallis 5y agoThe io specificity is expected to be a temporary situation, and that part of the name may end up being an anachronism in a few years once a usefully large subset of another category of syscalls has been added to io_uring. The shared ring buffers aspect definitely is an implementation detail, but one that does explain why performance is better than other async IO methods (and also avoids misleading people into thinking that it has something to do with the async/await paradigm). If the BSDs hadn't already claimed the name, it would probably have been fine to call this kqueue or something like that.
- hawski 5y agoNow I wonder if my idea of having a muxcall or a batchcall as I thought about it a few years ago is something similar to io_uring, but on a lesser scale and without eBPF goodies. My idea was to have a syscall like this: struct batchvec { unsigned long batchv_callnr; unsigned long long batchv_argmask; }; asmlinkage long sys_batchcall(struct batchvec *batchv, int batchvcnt, long args[16], unsigned flags); You were supposed to give in a batchvec a sequence of system call numbers and a little mapping to arguments you provided in args. batchv_argmask is a long long - 64 bit type, this mask is divided to 4 bit fields, every field can address a long from args table. AFAIR Linux syscalls have up to 6 arguments. 6 fields for arguments and one for return value, that gives 7 fields - 28 bits and now I don't remember why I thought I need a long long. It would go like this pseudo code: int i = 0; for(; i < batchvcnt; i++) { args[batchv[i].argmask[6]] = sys_call_table[batchv[i].callnr](args[batchv[i].argmask[0]], args[batchv[i].argmask[1]], args[batchv[i].argmask[2]], args[batchv[i].argmask[3]], args[batchv[i].argmask[4]], args[batchv[i].argmask[5]]); if(args[batchv[i].argmask[6]] < 0) { break; } } return i; It would return a number of successfully run syscalls. It would stop on first failed one. The user would have to pick up the error code out of args table. I would be interested to know why it wouldn't work. I started implementing it against Linux 4.15.12, but never went to test it. I have some code, but I don't believe it is my last version of the attempt.
- R0b0t1 5y agoI am hoping io_uring goes this way (maybe just call it uring). It would make people design APIs with a thought to how they could be async'd.
- moonchild 5y agoFwiw posix specifies readv and writev already; you may want to take a look at those. It's not really the same, though; readv and writev are still synchronous APIs, they just do more at once.
- hawski 5y agoI'm aware of readv/writev. This was supposed to cover much more, because it is with arbitrary syscalls. Open a file and if successful read from it - all with a single syscall. Another example: socket, setsockopt, bind, listen, accept in a single syscall.
- deleted 5y ago[deleted]
- matheusmoreira 5y ago> io_uring's true purpose is a generic asynchronous syscall facility Exactly. It's such an awesome design. I wonder how many system calls will be supported in the future.
- yxhuvud 5y agoWell, they have space for 255 ops :)