4 ms·
That is hardly the state of the art: they are basically sacrificing all abstractions that are rightly so required in the name of speed. This is no different t
by diffserv 8y ago
That is hardly the state of the art: they are basically sacrificing all abstractions that are rightly so required in the name of speed. This is no different than using vanilla DPDK with no congestion and flow control and being able to process 20 mil packets per core (better than the numbers in the paper). Getting 75Gbps per core using RDMA is hardly hard or new.
You can't use this abstraction for anything ... really unless you are on a completely lossless fabric that has enough capacity to avoid congestion.
Yes, Google/MS/FB are going towards such fabrics but the hard part is not building the RPC abstraction suggested in this paper---the hard part is getting to that fabric.
Just to put things into perspective, if you had a quantum computer you could do all sort of crazy stuff with it. You could "parallelize loops! and be super fast" is what this paper is suggesting (I specifically chose parallelizing loops cause there is nothing new in it).
- shereadsthenews 8y agoWell this sets a lower bound on the state of the art that is in any case orders of magnitude higher than what many people suffer from in industry practice.
- diffserv 8y agoJust to restate what I said---you cannot use this in its current state in production (or industry), ever. And by the time it becomes useful, it becomes a natural solution because the infrastructure supports it. Every person that works on kernel knows the overhead that comes with abstractions. Everybody that has worked with DPDK knows that you can get 20Mpps+ on a single core. All they have done is to "frame" the usage and the term RPC differently ... i.e., it's all story telling and no real meat :) Look at the previous set of publications by the same author: e.g., Achieve a Billion Requests Per Second Throughput on a Single Key-Value Store Server Platform, etc. they are all based on a single assumption that if you don't implement X and Y in your stack you can get better performance. Of course, you can. If you use your calculator to only compute 2+2 you might as well hardcode 4 as the output of your calculator. They are not setting a lower bound. They are hardcoding and bypassing the parts they don't find useful.
- anujkaliaitd 8y agoeRPC author chiming-in again. I agree that our code is not production-ready, but an industry team at Intel is trying eRPC out: https://github.com/daq-db https://github.com/daq-db. I am curious to know what features you believe we omitted in our ISCA 15 paper. Our key-value store (https://github.com/efficient/mica2 https://github.com/efficient/mica2) supports a memcached-like API over UDP.
- anujkaliaitd 8y agoI'm the main author of eRPC, and I wanted to clarify some things. Your comment suggests (please correct me if you meant differently) that (a) eRPC does not perform congestion control, and (b) eRPC requires a lossless fabric. In fact, eRPC implements congestion control, and it works well in a lossy network. Those are the two main contributions of the paper. We get 75 Gbps with only UDP/Ethernet packet I/O, without RDMA support. eRPC implements transport-layer functionality atop a fast packet I/O engine like DPDK, so comparing eRPC to DPDK isn't apples-to-apples.
- diffserv 8y agoHey, Anuj. Your repo is really nice for an academic paper. Thank you for that. It's rare to see a "networked system's" repositories that has readable code. I mainly checked large-tput example: A few questions— 1) For your 75Gbps, what percentage of the payload of the RPC do you touch? I.e., what portion of the message is used on that core? More directly, say you have a service that can sustain 100kQPS, if they switch to eRPC, what can they expect? Asked differently, what is the base overhead of today's RPC libraries? Especially ones that bypass kernel. 2) The congestion and flow control is debatable, and their efficacy is up for debate. Especially in a DC setting. Can you claim that eRPC would work for any types of the workload in a DC setting? How would it play out with other connections? At the end of the day, if you are forced to play nice, you may eventually add up branches in your code. Your fast path gets split depending on the connection type, etc. Is that something that you think is preventable? 3) How do you distribute the load across different cores at 75Gbps? How does the CPU ring, contention, etc. come into play? I.e., can you do useful work with that 75Gbps? or should I just read it as a "wow" number? Asking a different question, if I have a for loop that can do 10 billion loops per second and by just adding a function that drops down to 10k loops per second, why would I care about that 10 billion iterations? 4) You claim that it works well in a lossy network, yet your goodput drops to 18~2.5Gbps at 10^-4/10^-3 packet loss---I am still assuming the library is still flooding the network at 75Gbps. How does this play out in scale? All in all, I do appreciate your work. My issue is that academic people like to make big claims, especially in an academic setting. People in the industry are aware of fast-paths. Kernel networking stack uses fast-paths rigorously. Sure it is heavy and it comes with a lot of bulk, but you can as easily cut it down.