4 ms·
I feel the main limitation over here is hardware optimisation support. With TCP you have the congestion control algo baked in hardware, tcp segment and checksu
by kd913 3y ago
I feel the main limitation over here is hardware optimisation support.
With TCP you have the congestion control algo baked in hardware, tcp segment and checksum offload. You can pass things directly to the NIC for massive latency, bandwidth and offloading processing away from the cpu.
A properly tuned system with a user space networking stack with ctpio, hardware offload and proper tuning beats the pants of QUIC for latency.
It is possible to get some of the same benefits I guess with GSO. In any case, the slow hardware support here I suspect is a bottleneck. You may not get much benefit given more layers are binary/encrypted and not visible to hardware.
The above is not relevant for hyperscalers like Google I imagine where they can make the hardware and the sheer amount of customers offer bandwidth benefits.
- jabart 3y agoThe internet itself is not properly tuned and distance is an issue. On a local lan, sure a TCP tuned system will do amazing. On the public internet where who knows where your packet is going it's an optimal but not tuned system with latency TCP wasn't designed for. Also a modern CPU with SIMD/AVX can handle a lot of traffic in user space. UDP also has a checksum in it's packet header.
- kd913 3y agoFor hyperscalers in the cloud and low latency finance, tcp is likely still king. You do not want things running on CPU as that is compute that could be sold to someone else. This needs an easy way to offload to a specialised hardware accelerator.
- jabart 3y agohttps://aws.amazon.com/blogs/aws/new-http-3-support-for-amazon-cloudfront/ https://aws.amazon.com/blogs/aws/new-http-3-support-for-amaz... Seems AWS is doing just fine with it. They also have built their own custom chipsets though I don't see a mention of it offloading QUIC to those chips.
- jeffbee 3y agoAWS in some ways wasn't really ready for something like QUIC to emerge. Nitro networking does, or did, penalize all UDP flows. Which obviously is not great for QUIC.
- belthesar 3y agoHyperscalers in the cloud and low latency finance will drive the innovation to make HW offload a thing sooner. For the rest of us normal folk, we can likely live with CPU managed packets until that happens, and then take advantage of the optimization benefits when they make it to us. I think it's fair to call this out, but my initial read, and the reason why I'm responding, is because it seems like you're saying that it's non-viable.
- mannyv 3y agoUh, the fact is that the internet works and has been working pretty well for the last few decades. While you may have issues with certain aspects of TCP and its behavior over WAN, your issues have no practical significance in real life.
- jcelerier 3y ago> pretty well Doing a lot of heavy lifting right there. My experience (across many countries, travels, etc) is that internet (and things in general) just barely work to a barely tolerable level. Not a day where a coworker doesn't get some internet issue, some sudden disconnection, some weird TURN route killing latency..
- 01HNNWZ0MV43FF 3y agoI believe you but as a networking noob, could someone tell me how segmentation and old checksum algorithms are significant compared to TLS overhead? Is it because TLS is hardware accelerated with AES instructions or something already?
- kd913 3y agoTLS offload also exists and is normally implemented on NIC too. Basic goal here is not to process this on CPU as that is slow and compute that could be used for user apps/customers.
- jeffbee 3y agoIn my experience, congestion control in hardware would be the very last thing I would want. Everything needs to be pushed as far toward the edges of the system as possible. This is what quic offers.
- kd913 3y agoFor what? It has been done for a while now. The fastest, lowest latency mechanisms will always offload to accelerator cards in hardware.
- jeffbee 3y agoThe NIC does not have the information required to make machine-wide optimal decisions about congestion control. It doesn't know what flows are critical at the application level. NICs also can't know whether you want to optimize median or tail latency for a given flow. Only the application knows and that's where the congestion control algorithm needs to run.
- b112 3y agoWith TCP you have the congestion control algo baked in hardware, tcp segment and checksum offload. You can pass things directly to the NIC for massive latency, bandwidth and offloading processing away from the cpu. A lot of NICs are essentially just as software modems, with the CPU handling this anyhow. Even server hardware sometimes has these lame NIC chips on them. Be careful to rid the specs of the hardware you buy, otherwise it will indeed be your CPU doing all that work.
- wmf 3y agoWith TCP you have the congestion control algo baked in hardware This part isn't correct, and you wouldn't want e.g. NewReno baked in to your NIC and preventing you from using CUBIC or BBR. It's true that TCP benefits more from NIC offloads than QUIC but most places (besides Netflix) aren't driving enough WAN traffic per server to matter.