5 ms·
So the question is why hasn't an optimized RPC implementation emerged for the data center that avoids the headaches that come with layering over TCP. Turns out
by cjsplat 4y ago
So the question is why hasn't an optimized RPC implementation emerged for the data center that avoids the headaches that come with layering over TCP.
Turns out integrating with the control plane is hard at data center scale.
When everything is working, an optimized protocol stack is great and everyone is happy.
When congestion happens (fabric, kernel, memory bus, cache footprint, etc), and your optimized app causes a legacy production app to behave strangely, all hell breaks loose.
For example, a distributed file system SRE sees some strange tail latency effects and doesn't know that a cluster is shared with a non-TCP app and of course the optimized stack doesn't respond to congestion the way TCP would.
Worst case, the SREs notice this just after an update is pushed in their app, and they spend a random amount of time trying to figure out if the update is what changed the behavior.
Best case, they know there are non-TCP apps in the area, and they ping the SRE for that stack and say "my app is behaving strange, please hit your big red off button so I can see if that fixes the problem."
Someplace in between those options, the SRE is using TCP aware network debugging tools to try to figure out why this day is different from yesterday, or this cluster is different from the one where everything is working fine.
Regardless, you get an unhappy SRE. But of course you never get just one unhappy SRE.
So you need to generate a lot of value got justify their pain.