4 ms·
I don't think we concluded that there were any fundamentally unsafe aspects of io_uring. We decided to look at it because it's an interface of great interest to
by staticassertion 5y ago
I don't think we concluded that there were any fundamentally unsafe aspects of io_uring. We decided to look at it because it's an interface of great interest to us as a company, and we suspected that the combination of: new code, performance oriented code, concurrency oriented code, would be a great place to find some bugs.
Whenever we are interested in adopting new technologies we do a security review, so this naturally came out of that. We'll be posting more posts on other areas of interest for us.
- wahern 5y agoio_uring relies on a pool of independent kernel threads performing operations on buffers (and other resources) provided by non-privileged userspace processes. While conceptually simple, it's a sharp departure from the standard userspace/kernel syscall and process model. It was inevitable that it would stumble over little nooks & crannies of the kernel that silently made risky assumptions dependent on the standard model. > I don't think we concluded that there were any fundamentally unsafe aspects of io_uring. Does io_uring still have the trap that a ring context initialized before a process drops privileges with setuid can still dispatch operations on root-privileged kernel worker threads? That's a nasty problem, partly related to the fact that Linux has no process-global UID--every thread's UID (effective and saved, plus GID, supplementary GIDs, etc) has to be managed separately per thread, which requires herculean hacks in libc to provide POSIX setuid semantics, which except for very specialized, Linux-specific software (e.g. runc) is all that most people care about or even consider. IIRC there was a related issue when passing a ring context to a different process altogether, but it was already fixed or at least mitigated. There's some irony in io_uring both being so performant and becoming popular; it has microkernel written all over, which is an approach that Linux (and Linus) notoriously ridiculed so many years ago as requiring interfaces that were both too slow and too complicated. Except, oddly, instead of where a proper microkernel would preserve and sharpen privilege boundaries (including capability objects, VM isolation, etc), io_uring hacks around them and reduces their effectiveness.
- 10000truths 5y ago> io_uring relies on a pool of independent kernel threads performing operations on buffers (and other resources) provided by non-privileged userspace processes. While conceptually simple, it's a sharp departure from the standard userspace/kernel syscall and process model. It was inevitable that it would stumble over little nooks & crannies of the kernel that silently made risky assumptions dependent on the standard model. The problem here has nothing to do with kernel threads reading user-space data asynchronously. The problem is that user-provided struct field in a system call could be interpreted as a kernel-space address and operated on, and one of the kernel functions missed a check for that overload. > Does io_uring still have the trap that a ring context initialized before a process drops privileges with setuid can still dispatch operations on root-privileged kernel worker threads? That's a nasty problem, partly related to the fact that Linux has no process-global UID--every thread's UID (effective and saved, plus GID, supplementary GIDs, etc) has to be managed separately per thread, which requires herculean hacks in libc to provide POSIX setuid semantics, which except for very specialized, Linux-specific software (e.g. runc) is all that most people care about or even consider. IIRC there was a related issue when passing a ring context to a different process altogether, but it was already fixed or at least mitigated. Those permissions semantics are by design, even in POSIX. A file’s associated permissions are determined at the time of creation, not at the time of access. It is expected that one can open a root-owned resource as root, drop privileges, and still access the resource. And if you think about it in the context of io_uring, there isn’t really any other sensible way to do it - there’s no way for the kernel to determine which task submitted an SQE because it’s just a write in a memory address space that may be shared by any number of tasks.
- hinkley 5y agoContainers look an awful lot like user processes for microkernels as well. At this point I’m wondering if Tannenbaum will live long enough to say “I told you so.”
- yxhuvud 5y ago> io_uring relies on a pool of independent kernel threads performing operations on buffers Do note that this setup was rewritten at kernel version 5.13 or something like that, and the current model seem to be some sort of hybrid variant of kernel and userspace threads. From what I gather it was a huge improvement compared to how it was before.