2 ms·
"Spin up this worker" is trivial: one memory allocation, probably from a free list, to allocate a thread stack, setting a couple of pointers, and a function ca
by memefrog 3y ago
"Spin up this worker" is trivial: one memory allocation, probably from a free list, to allocate a thread stack, setting a couple of pointers, and a function call. Doing any kind of I/O is much heavier-weight. Thread stacks don't have to even be very large. It doesn't expend any additional resources to have a userspace thread blocked on a socket. You can just have a pointer to the thread as the userdata of the SQE.
It doesn't make any sense to say 'use the time to do useful work instead'. That's what you do when you block: in a threaded model, when you block you do 'io_uring_get_sqe', 'io_uring_prep_read' (or whatever), 'io_uring_set_userdata' 'io_uring_submit' and then you switch contexts (with <ucontext.h> or some equivalent) to another thread that is ready. When you have no ready threads you just call 'io_uring_wait_cqe', and for each completion you do 'io_uring_get_userdata', get a pointer to the thread, record the result of the operation and append the thread to a runqueue to be resumed.
From the perspective of the thread, you write
int e = read(fd, buf, sizeof buf);
if (e < 0) err(1, "read: %s", strerror(-e));
and it all feels completely sequential. Underneath it is stackful coroutines.