3 ms·
> They're not that expensive; something like 1000 cycles. When you factor in all the atomic operations, which take 10-25x more cycles than normal operations, re
by fstrthnscnd 5y ago
> They're not that expensive; something like 1000 cycles. When you factor in all the atomic operations, which take 10-25x more cycles than normal operations, required to post and pull from shared memory queues, you're not actually saving much at all.
I thought that io_uring buffers weren't shared? Or do you mean in a multithreaded app?
- wahern 5y agoAn io_uring context is a thread pool, except the worker threads are kernel threads. The userland application posts an operation request (poll, open, read/write, etc) to a special mmap'd buffer that is dequeued by a kernel thread dedicated to that context. That kernel thread then either performs the operation itself, or if it's a potentially blocking operation (i.e. file read) passes it to another worker thread. The results are then posted back to another shared buffer-based message queue. (I'm ignoring what happens when the shared buffers are filled, which has its own implications--good and bad--for performance.) It's basically no different than a typical multi-threaded userland application using a ring buffer message queue, except the worker threads operate in kernel context and can access unpublished kernel APIs. But like all multi-threaded designs, the performance choke points are always the places where you need to either obtain a lock or use atomic operations. The problem with locks are obvious. But atomics operations absolutely destroy the performance of pipelined CPU architectures, both directly and indirectly (e.g. if working data is being passed to a thread scheduled on a different core). Ever wonder why software transactional memory never really succeeded despite the hype? Because a lock often only requires a single atomic operation, whereas lockless algorithms often rely on a whole series of atomic operations. In practice lockless algorithms, while they scale well asymptotically, have horrible absolute performance. I'm not saying that io_uring's operation queue is intrinsically slower. My point is that 1) context switches, especially in Linux, aren't that slow and 2) all the atomic memory operations have real costs. So the relative benefit of the message passing and worker thread architecture (which is a fixed cost in io_uring, much like a context switch for a syscall) isn't as great as you'd think. It's almost de minimis from the perspective of whole-application architecture.
- eru 5y ago> Ever wonder why software transactional memory never really succeeded despite the hype? You can implement software transactional memory with locks just fine. Are you talking about hardware transactional memory perhaps? (In some sense, database transactions are pretty similar to software transactional memory. And that points to a different possible reason why software transactional memory (STM) hasn't taken off: STM works really well in a language like SQL or Haskell with carefully restricted side-effects. But bolting STM onto an impure language is asking for trouble.) PS Your main point about locks vs atomic primitive still stands regardless.