4 ms·
Ignoring all the other problems with this article that have been pointed out around definitions, it also claims that lockless is slow anyway. Without giving lit
by DannyBee 1y ago
Ignoring all the other problems with this article that have been pointed out around definitions, it also claims that lockless is slow anyway. Without giving literally any data.
Good news: it's not slow
Let's take rust, which has an oddly thriving ecosystem of lockless mpsc/mpmc/etc queues, and lots of benchmarks of them in lots of configurations.
The fastest ones easily do at least 30 million elements/second in most configurations. The "slowest" around 5-10.
So the fastest is doing 33ns per element and the slowest is 100ns.
Which is probably why the article offers no data. It's not actually slow. It's actually really fast when done well.
- habibur 1y agoI read "slow" and "fast" in the article as comparative term. "slower than" what the writer has seen in other cases. Absolute ms doesn't prove much unless you put it in comparison with other best.
- DannyBee 1y agoWithout any data it's impossible to even tell what this means or whether it matters. IE even as a comparitive term it's useless without real data. Theoretically better (which is questionable at best in this case) doesn't mean anything useful if it doesn't actually matter in practice, or is always worse in pracftice.
- pclmulqdq 1y agoI have always found MPMC ring buffers to have worse tails than pointer-based MPMC queues. These are very unfashionable (especially in rust) but can be made wait-free with a single atomic for readers and writers. The SPSC ring buffer is perfect, but adding either more producers or more consumers makes the whole thing have terrible pathologies. Reading the article again, this lockless bag idea also needs two atomic writes on every operation.
- pkhuong 1y agoSPMC ring buffer or SPMC "disruptor" aren't that bad. Multiple producers in a general ring buffer definitely introduce a lot of issues.
- pclmulqdq 1y agoThe "disruptor" pattern has some significant problems with tail performance (around pathological OS scheduling) if you aren't running one thread per core the way LMAX designed it.
- SUPERCILEX 1y agoAs noted by other commenters, the point I was trying to get across is that the way we implement lockless channels is suboptimal and could be made faster from a theoretical standpoint. In my benchmarks[1], the average processing time for an element is 250ns with 4 producers and 4 consumers contending heavily. That's terrible! Even if your numbers are correct, 100ns is a bit faster than two round trips to RAM while 33ns is about three round trips to L3 and ~100x slower than spamming a core with add operations. That's slow. [1]: https://github.com/SUPERCILEX/lockness/blob/master/bags/benches/atomic_bag.rs https://github.com/SUPERCILEX/lockness/blob/master/bags/benc... $ cargo bench -- 8_threads/std_mpmc
- DannyBee 1y agoIt's only terrible if it actually matters? Also it doesn't look like you are using any of the optimized lockless implementations like crossfire/etc, so not sure what this actually proves?
- koakuma-chan 1y agoWhy would lockless be slower? Is it ever slower?