4 ms·
As noted by other commenters, the point I was trying to get across is that the way we implement lockless channels is suboptimal and could be made faster from a
by SUPERCILEX 1y ago
As noted by other commenters, the point I was trying to get across is that the way we implement lockless channels is suboptimal and could be made faster from a theoretical standpoint.
In my benchmarks[1], the average processing time for an element is 250ns with 4 producers and 4 consumers contending heavily. That's terrible! Even if your numbers are correct, 100ns is a bit faster than two round trips to RAM while 33ns is about three round trips to L3 and ~100x slower than spamming a core with add operations. That's slow.
[1]: https://github.com/SUPERCILEX/lockness/blob/master/bags/benches/atomic_bag.rs https://github.com/SUPERCILEX/lockness/blob/master/bags/benc...
$ cargo bench -- 8_threads/std_mpmc
- DannyBee 1y agoIt's only terrible if it actually matters? Also it doesn't look like you are using any of the optimized lockless implementations like crossfire/etc, so not sure what this actually proves?