4 ms·
Nanosecond? I don't think so... Unless by having nanosecond latency you mean that it has a latency that is representable in nanoseconds... but that's not what i
by dicroce 11y ago
Nanosecond? I don't think so... Unless by having nanosecond latency you mean that it has a latency that is representable in nanoseconds... but that's not what it means to me. Tens or hundreds of microseconds I would buy, but not nanosecond.
- vardump 11y agoIf you busy poll, getting to 50 ns latency from producer to consumer is easy, 20 ns is possible. The best I've gotten with a bare bones optimized test was about 16 ns. You tend to get best results with high frequency dual core CPUs with hyper-threading disabled. Bigger issue is that the reader needs to busy poll, consuming 100% CPU time on one CPU core. Maybe energy consumption could be reduced by using monitor/mwait - not sure if it's possible from user mode or only in kernel. Another issues in this implementation is use of compare-and-swap. For this purpose, fetch-and-add (LOCK XADD on x86) would be more efficient. Multiple writer contention is much worse with CAS. Record size is also fixed, you can't have multiple different message sizes. Overhead per message is 12 bytes, two ints that can be either 0 or 1 and an int called Metadata. Seeing there's commit and rollback fields, both 1 byte, 2 byte overhead for commit and rollback should have been enough. You're going to get false sharing whether it's an int or a byte.
- orf 11y ago> For this purpose, fetch-and-add would be more efficient. Multiple writer contention is much worse with CAS. Why, out of interest?
- vardump 11y agoWhen there's heavy amount of updates, CAS loop updates will start to fail, so it needs to be retried the more the more there is contention. Fetch-and-add is totally wait-free, much better upper bound. It always succeeds on first try.
- jberryman 11y agoI can confirm. If you're curious I wrote a fast concurrent (single process) queue library for Haskell based around fetch-and-add: https://github.com/jberryman/unagi-chan https://github.com/jberryman/unagi-chan Interested to check out your work! I'd love to be able to extend my library to support IPC in some way.
- dicroce 11y agoUgh. Polling? You better have serious throughput to justify that....
- vardump 11y agoYes, polling. If kernel gets involved, we're talking about microseconds.
- rdtsc 11y ago> You better have serious throughput to justify that.... It is often about latency not just throughput. Although sometimes they go hand in hand. For example you can achieve pretty high throughput if you take the whole network stack outside the kernel and talk directly to the network card. http://www.intel.com/content/dam/www/public/us/en/documents/presentation/dpdk-packet-processing-ia-overview-presentation.pdf http://www.intel.com/content/dam/www/public/us/en/documents/... But I've seen this spin polling done when latency needed to be optimized.
- jbooth 11y agoJust to throw it out there, I thought communication across 2 cores on different sockets is generally regarded as impossible in less than 100ns regardless of language? If you're staying in L3 on a single socket, I could believe it, maybe, but at a certain point you're tuning the benchmark rather than the application.
- vitalyd 11y agoFor low latency messaging like this you'd typically want to avoid cross socket data transfers. However, IIRC cacheline snooping across sockets is a bit quicker (Intel QPI at least) than remote dram access so if the lines stay in cache the communication may be under 100ns. Also, depending on access pattern, hardware prefetch may hide some of the latency.
- kasey_junk 11y agoDo you have a particular reason to claim IPC isn't possessible sub micro? Cause I've written/used systems that performed IPC measured in 10s of nanos to the best of my ability to measure it.
- dicroce 11y agoI hadn't considered someone would use polling like this. Profoundly wasteful of the CPU unless you have a ridiculous throughput.
- kasey_junk 11y agoOr low latency requirements.
- ma2rten 11y agoMaybe in that case you should use a single thread and avoid ipc altogether?
- kasey_junk 11y agoFor sure if you can avoid it. But sometimes you can't. In those cases you need inter-thread communication to be as fast as possible. A reason inter-process communication is interesting, especially in java, is that there is a reasonably common pattern of isolating low latency code in 1 JVM and less sensitive code in another to prevent system wide pauses from bleeding across. This can be super painful for development iteration though so one thing that is interesting about this library is that it seems to make that configuration time concern.
- vitalyd 11y agoIt may also be beneficial to move some processing to a different core (perhaps even different socket) to avoid execution resource contention (e.g. d/i-cache, execution ports, etc).
- 11y ago
- tlrobinson 11y ago(for those not so good at math, 1ns = 1 clock cycle at 1GHz)