11 ms·
It's time to replace TCP in the datacenter (2023)
- wmf 2y agoPrevious discussions: Homa, a transport protocol to replace TCP for low-latency RPC in data centers https://news.ycombinator.com/item?id=28204808 https://news.ycombinator.com/item?id=28204808 Linux implementation of Homa https://news.ycombinator.com/item?id=28440542 https://news.ycombinator.com/item?id=28440542
- unsnap_biceps 2y agoThe original paper was discussed previously at https://news.ycombinator.com/item?id=33401480 https://news.ycombinator.com/item?id=33401480
- parasubvert 2y agoand here, from an LWN analysis. https://news.ycombinator.com/item?id=33538649 https://news.ycombinator.com/item?id=33538649
- ksec 2y agoHoma: A Receiver-Driven Low-Latency Transport Protocol Using Network Priorities https://people.csail.mit.edu/alizadeh/papers/homa-sigcomm18.pdf https://people.csail.mit.edu/alizadeh/papers/homa-sigcomm18....
- 7e 2y agoTCP was replaced in the data centers of certain FAANG companies years before this paper.
- wmf 2y agoIf they keep it secret they don't get credit for it.
- andrewflnr 2y agoHow do you figure? The right decision is the right decision, even if you don't tell people. (granting, for the sake of argument, that it is the right decision)
- wmf 2y agoYeah, you get the benefit of secret tech (in this case faster networking) but people shouldn't give social credit for it because that creates incentives to lie. And, sadly, tech adoption runs entirely on social proof.
- deleted 2y ago[deleted]
- andrewflnr 2y agoWe're not really trying to allocate social credit here, not as our main goal anyway. We're evaluating raw effectiveness of the tech. So if they made an effective decision, we give them credit for, uh, making an effective decision. You don't have to love them for it.
- michaelt 2y agoWhen you are an outsider it's wise to take such claims with a grain of salt, because as the "secret" made its way to you, recounted by one person to another to another, there might have been an exaggeration, over-simplification or misunderstanding. It's easy to imagine how, in the hands of tech journalists and youtubers optimising for clicks, "Google likes QUIC" and "Some ML clusters use infiniband" could get distorted into several faangs and the complete elimination of TCP.
- andrewflnr 2y agoTrue as far as it goes, but "they didn't actually do it" is a different story from "they did it secretly". The two claims exclude each other, so you can't really compare them in the context of "credit".
- bushbaba 2y ago*minority of the fangs.
- albert_e 2y agoCurious ... replaced with what, I would like to know.
- parasubvert 2y agoHTTP/3 aka QUIC (UDP).
- JoshTriplett 2y agoOr SRD, in AWS: https://aws.amazon.com/blogs/hpc/in-the-search-for-performance-theres-more-than-one-way-to-build-a-network/ https://aws.amazon.com/blogs/hpc/in-the-search-for-performan...
- dradra67 2y agohttps://research.google/pubs/snap-a-microkernel-approach-to-host-networking/ https://research.google/pubs/snap-a-microkernel-approach-to-...
- akira2501 2y ago> If Homa becomes widely deployed, I hypothesize that core congestion will cease to exist as a significant networking problem, as long as the core is not systemically overloaded. Yep. Sure; but, what happens when it becomes overloaded? > Homa manages congestion from the receiver, not the sender. [...] but the remaining scheduled packets may only be sent in response to grants from the receiver I hypothesize it will not be a great day when you do become "systemically" overloaded.
- andrewflnr 2y agoWill it be a worse day than it would be with TCP? Either way, the only solution is to add more hardware, unless I'm misunderstanding the term "systemically overloaded".
- bayindirh 2y agoI think so. If your core saturates, you add more capacity to your core switch. In HOMA, you need to add more receivers, but if you can't add them because the core can't handle more ports? Ehrm. Looks like core saturation all over again. Edit: Just popped to my mind. What prevents maliciously reducing "receive quotas" on compromised receivers to saturate an otherwise capable core? Looks like it's a very low bar for a very high impact DOS attack. Ouch.
- klysm 2y agoThis is designed for in data center use, so the security tradeoff is probably worth it
- bayindirh 2y agoNope. Tending a datacenter close to two decades, I can say that putting people behind NATs and in jails/containers/VMs doesn't prevent security incidents all the time. With all the bandwidth, processing power and free cooling, a server is always a good target, and sometimes people will come equipped with 0-days or exploits which are very, very fresh. Been there, seen and cleaned that mess. I mean, reinstallation is 10 minutes, but the event is always ugly.
- yesbut 2y agoAnother thing not worth investing time into for the rest of our careers. TCP will be around for decades to come.
- t-writescode 2y agoTrue! And chances are, if you're developing website software or video game software, you'll never think about these sorts of things, it'll just be a dumb pipe for you, still. And that's okay! But there are other sorts of computer people than website writers and business application devs, and they're some of the people this would be interesting for!
- imtringued 2y ago>True! And chances are, if you're developing website software or video game software, you'll never think about these sorts of things, it'll just be a dumb pipe for you, still. Wrong. I've experienced most of the complaints in the paper when developing multiplayer video games. These days I simply use websockets instead of raw TCP because it is not worth the effort and yet you still have to do manual heartbeats.
- UltraSane 2y agoI wonder why Fibre Channel isn't used as a replacement for TCP in the datacenter. It is a very robust L3 protocol. It was designed to connect block storage devices to servers while making the OS think they are directly connected. OSs do NOT tolerate dropped data when reading and writing to block devices and so Fibre Channel has a extremely robust Token Bucket algorithm. The algo prevents congestion by allowing receivers to control how much data senders can send. I have worked with a lot of VMware clusters that use FC to connect servers to storage arrays and it has ALWAYS worked perfectly.
- Sebb767 2y ago> I wonder why Fibre Channel isn't used as a replacement for TCP in the datacenter But it is often used for block storage in datacenters. Using it for anything else is going to be hard, as it is incompatible with TCP. The problem with not using TCP is the same thing HOMA will face - anything already speaks TCP, nearly all potential hires know TCP and most problems you have with TCP have been solved by smart engineers already. Hardware is also easily available. Once you drop all those advantages, either your scale or your gains need to be massive to make that investment worth it, which is why TCP replacements are so rare outside of FAANG.
- ksec 2y agoI wonder if there are any work on making something similar ( conceptually ) to TCP, super / sub set of TCP while offering 50-80% benefits of HOMA. I guess I am old. Everytime I see new tech that wants to be hyped, completely throw out everything that is widely supported and working for 80-90% of uses cases, not battle tested and may be conceptually complex I will simply pass.
- Sebb767 2y agoIf you have a sufficiently stable network and/or known failure cases, you can already tune TCP quite a bit with nodelay, large congestion windows etc.. There's also QUIC, which basically is a modern implementation of TCP on top of UDP (with some trade-offs chosen with HTTP in mind). Once you stray too far, you'll loose the ability to use off-the-shelve hardware, though, at which point you'll quickly hit the point of diminishing returns - especially when simply upgrading the speed of the network hardware is usually a cheap alternative.
- bmitc 2y agoUnrelated to this article, are there any reasons to use TCP/IP over WebSockets? The latter is such a clean, message-based interface that I don't see a reason to use TCP/IP.
- tacitusarc 2y agoWebsockets is a layer on top of TCP/IP.
- bmitc 2y agoYes, I know that WebSockets layer over TCP/IP. But that both misses the point and is part of the point. The reason that I ask is that WebSockets seem to almost always be used in the context of web applications. TCP/IP still seems to dominate control communications between hardware. But why not WebSockets? Almost everyone ends up building a message framing protocol on top of TCP/IP, so why not just use WebSockets which has bi-directional message framing built-in? I'm just not seeing why WebSockets aren't as ubiquitous as TCP/IP and only seem to be relegated to web applications.
- j16sdiz 2y agoWebSocket is fairly inefficient protocol. and it needs to deal with the upgrade from HTTP. and you still need to implement you app specific protocol. This is adding complexity without additional benefit It make sense only if you have an websocket based stack and don't want to maintain a second protocol.
- imtringued 2y agoYou can easily build a JSON based RPC protocol in a few minutes using WebSockets and be done. With raw TCP you're going to be spending a week doing something millions of other developers have done again and again in your own custom bespoke way that nobody else will understand. Your second point is very dismissive. You're inserting random application requirements that the vast majority of application developers don't care about and then you claim that only in this situation do WebSockets make sense when in reality the vast majority of developers only use WebSockets and your suggestion involves the second unwanted protocol (e.g. the horror that is protobuffers and gRPC).
- slt2021 2y agothe problem with trying to replace TCP only inside DC, is because TCP will still be used outside DC. Networking Engineering is already convoluted and troublesome as it is right now, using only tcp stack. When you start using homa inside, but TCP from outside things will break, because a lot of DC requests are created as a response for an inbound request from outside DC (like a client trying to send RPC request). I cannot imagine trying to troubleshoot hybrid problems at the intersection of tcp and homa, its gonna be a nightmare. Plus I don't understand why create a a new L4 transport protocol for a specific L7 application (RPC)? This seems like a suboptimal choice, because RPC of today could be replaced with something completely different, like RDMA over Ethernet for AI workloads or transfer of large streams like training data/AI model state. I think tuning TCP stack in the kernel, adding more configuration knobs for TCP, switching from stream(tcp) to packet (udp) protocols where it is warranted, will give more incremental benefits. One major thing author missed is security applications, these are considered table stakes: 1. encryption in transit: handshake/negotiation 2. ability to intercept and do traffic inspection for enterprise security purposes 3. resistance to attacks like flood 4. security of sockets in containerized Linux environment
- nicman23 2y agoonly thing homa makes sense is when there is no external tcp to the peers or at least not on the same context ie for roce
- slt2021 2y ago1. add software defined network, where transport and signaling is done by vendor-specific underlay, possibly across multiple redundant uplinks 2. term "external" is really vague as modern networks have blended boundaries. Things like availability zone, region make dc-dc connection irrelevant, because at any point of time you will be required to failover to another AZ/DC/region. 3. when I think of inter-Datacenter, I can only think of Ethernet. That's really it. Even in Ethernet, what you think of a peer and existing in your same subnet, could be a different DC, again due to software-defined network.
- jayd16 2y ago
- runlaszlorun 2y agoFor those who might not have noticed, the author is John Ousterhout- best known for TCL/Tk as well as the Raft consensus protocol among others.
- signa11 2y agoand more recently (?) the book : “a philosophy of software design”, highly recommended !
- stiray 2y agoHow long did we need to support ipv6? Is it supported yet and more widely in use than the ipv4, like in mobile networks where everything is stashed behind NAT and ipv4 kept? Another protocol, something completely new? Good luck with that, i would rather bet on global warming to put us out of our misery (/s)... https://imgs.xkcd.com/comics/standards.png https://imgs.xkcd.com/comics/standards.png
- detaro 2y agoMobile networks especially are widely IPv6, with IPv4 being translated/tunneled where still needed. (End-user connections in general skew IPv6 in many places - it's observable how traffic patterns shift with people being at work vs at home. Corporate networks without IPv6 leading to more IPv4 traffic during the day, in the evening IPv6 from consumer connections takes over)
- stiray 2y agoAndroid: Settings -> About (just checked mine, 10...*), check your IP. We have 3 providers in our country, all 3 are using ipv4 "lan" for phone connectivity, behind NAT and I am observing this situation around most of EU (Germany, Austria, Portugal, Italy, Spain, France, various providers).
- freetanga 2y agoSo, back to the mainframe and SNA in the data centers?
- wmf 2y agoIf Rosenblum can get an award for rediscovering mainframe virtualization, why not give Ousterhout an award for rediscovering SNA? (SNA was before my time so I have no idea if Homa is similar or not.)
- parasubvert 2y agoThis has already been done at scale with HTTP/3 (QUIC), it's just not widely distributed beyond the largest sites & most popular web browsers. gRPC for example is still on multiplexed TCP via HTTP/2, which is "good enough" for many. Though it doesn't really replace TCP, it's just that the predominant requirements have changed (as Ousterhout points out). Bruce Davie has a series of articles on this: https://systemsapproach.substack.com/p/quic-is-not-a-tcp-replacement https://systemsapproach.substack.com/p/quic-is-not-a-tcp-rep... Also see Ivan Pepelnjak's commentary (he disagrees with Ousterhout): https://blog.ipspace.net/2023/01/data-center-tcp-replacement/ https://blog.ipspace.net/2023/01/data-center-tcp-replacement...
- jpgvm 2y agoPlus in a modern DC you can trivially convert it to be essentially lossless with DCSP/PFC and ECN which both work perfectly with any UDP based protocol (and is why NVMeOF, FCoE and RoCEv2 all woke so well today). ECN isn't a necessity unless you need truly lossless network but the rest should get you pretty far as long as you are reasonably careful about communication patterns and blocking ratio of spine/core. For anything that really needs the lowest possible latency at the cost of all other considerations there is still always Infiniband.
- wbl 2y agoQUIC is not trying to solve the same problem as Ousterhout is. End user networks very different from datacenter.
- parasubvert 2y agoHow so? The same RPC-oriented L7 protocols are largely in use, just with a lot more east/west communications.
- aseipp 2y agoQUIC and Homa are not remotely similar and have completely different design constraints. I have no idea why people keep bringing up QUIC in this thread other than "It's another thing that isn't TCP." Yes, many things are not-TCP. The details are what matter.
- dveeden2 2y agoWasn't something like HOMA already tried with SCTP?
- iforgotpassword 2y agoAnd QUIC. And that thing tesla presented recently, with custom silicon even. And as usual, hardware gets faster, better and cheaper over the next years and suddenly the problem isn't a problem anymore - if it even ever was for the vast majority of applications. We only recently got a new fleet of compute nodes with 100gbit NICs. The previous one only had 10, plus omnipath. We're going ethernet only this time. I remember when saturating 10gbit/s was a challenge. This time around, reaching line speed with tcp, the server didn't even break a sweat. No jumbo frames, no fiddling with tunables. And that actually was while testing with 4 years old xeon boxes, not even the final hw. Again, I can see how there are use cases that benefit from even lower latency, but thats a niche compared to all DC business, and I'd assume you might just want rdma in that case, instead of optimizing on top of ethernet or IP.
- silisili 2y agoThis is a solid answer, as someone on the ground. TCP is not the bogeyman people point it out to be. It's the poison apple where some folks are looking for low hanging fruit.
- GoblinSlayer 2y ago> For many years, RDMA NICs could cache the state for only a few hundred connections; if the number of active connections exceeded the cache size, information had to be shuffled between host memory and the NIC, with a considerable loss in performance. A massively parallel task? Sounds like something doable with GPGPU.
- mafuy 2y agoThis has nothing to do with computing, it is about memory access.
- dboreham 2y agoNow the information has to be shuffled between the NIC and host memory and the GPU.
- GoblinSlayer 2y agoJust plug ethernet cable into the graphics card, then you need to shuffle memory between GPGPU and the wire.
- Woodi 2y agoYou want to replace TCP becouse it is bad ? Then give better "connected" protocol over raw IP and other raw network topologies. Use it. Done. Don't mess with another IP -> UDP -> something
- tsimionescu 2y agoWithin a data center? Maybe. On the Internet? You try offering a consumer service over SCTP first, and that is decades old by this point.
- indolering 2y agoSo token ring?
- pif 2y ago> Although Homa is not API-compatible with TCP, IPv6 anyone? People must start to understand that "Because this is the way it is" is a valid, actually extremely valid, answer to any question like "Why don't we just switch technology A with technology B?" Despite all the shortcomings of the old technology, and the advantages of the new one, inertia _is_ a factor, and you must accept that most users will simply even refuse to acknowledge the problem you want to describe. For you your solution to get any traction, it must deliver value right now, in the current ecosystem. Otherwise, it's doomed to fail by being ignored over and over.
- bamboozled 2y agoAlso there needs to be a big push to educate people on <new thing>. I know TCP very well, it would need quite a lot of incentive for me to drop that for something I don't yet understand as well. TCP was highly beneficial which is why we all adopted it in the first place, whatever is to replace it needs to be at least that beneficial...which will be a tall order.
- nine_k 2y agoI'd say that usually it's not about the balance of advantages but the balance of pain. You go through the pains of switching to a new and unfamiliar solution if your current solution is giving you even more pain. If you don't feel much pain, you can and should stay with your current solution. If it's not broken, or not broken badly enough, don't fix it by radical surgery.
- stonemetal12 2y ago> inertia _is_ a factor why would inertia be a factor? If I want to support protocol ABC in my data center, then I buy hardware that supports protocol ABC, including the ability to down shift to TCP when data leaves the data center. We aren't talking about the internet at large so there is no need to coordinate support with different organizations with different needs. Google, could mandate you can't buy a router or firewall that doesn't support IPv6. Then their entire datacenter would be IPv6 internally. The only time to convert to IPv4, would be if the local ISP doesn't support v6.
- ironhaven 2y ago
- kmeisthax 2y agoDumb question: why was it decided to only provide an unreliable datagram protocol in standard IP transit?
- michaelt 2y agoBecause when you're sending a signal down a wire or through the air, fundamentally the communication medium only provides "Send it, maybe it arrives" At any time, the receiver could lose power. Or a burst of interference could disrupt the radio link. Or a backhoe could slice through the cable. Or many other things. IP merely reflects this physical reality.
- kmeisthax 2y agoOk, but why does TCP exist, then? If we could make streams reliable in the late 1970s why didn't we apply that to datagrams as well?
- mhandley 2y agoIt's already happening. For the more demanding workloads such as AI training, RDMA has been the norm for a while, either over Infiniband or Ethernet, with Ethernet gaining ground more recently. RoCE is pretty flawed though for reasons Ousterhout mentions, plus others, so a lot of work has been happening on new protocols to be implemented in hardware in next-gen high performance NICs. The Ultra Ethernet Transport specs aren't public yet so I can only quote the public whitepaper [0]: "The UEC transport protocol advances beyond the status quo by providing the following: ● An open protocol specification designed from the start to run over IP and Ethernet ● Multipath, packet-spraying delivery that fully utilizes the AI network without causing congestion or head-of-line blocking, eliminating the need for centralized load-balancing algorithms and route controllers ● Incast management mechanisms that control fan-in on the final link to the destination host with minimal drop ● Efficient rate control algorithms that allow the transport to quickly ramp to wire-rate while not causing performance loss for competing flows ● APIs for out-of-order packet delivery with optional in-order completion of messages, maximizing concurrency in the network and application, and minimizing message latency ● Scale for networks of the future, with support for 1,000,000 endpoints ● Performance and optimal network utilization without requiring congestion algorithm parameter tuning specific to the network and workloads ● Designed to achieve wire-rate performance on commodity hardware at 800G, 1.6T and faster Ethernet networks of the future" You can think of it as the love-child of NDP [2] (including support for packet trimming in Ethernet switches [1]) and something similar to Swift [3] (also see [1]). I don't know if UET itself will be what wins, but my point is the industry is taking the problems seriously and innovating pretty rapidly right now. Disclaimer: in a previous life I was the editor of the UEC Congestion Control spec. [0] https://ultraethernet.org/wp-content/uploads/sites/20/2023/10/23.07.12-UEC-1.0-Overview-FINAL-WITH-LOGO.pdf https://ultraethernet.org/wp-content/uploads/sites/20/2023/1... [1] https://ultraethernet.org/ultra-ethernet-specification-update/ https://ultraethernet.org/ultra-ethernet-specification-updat... [2] https://ccronline.sigcomm.org/wp-content/uploads/2019/10/acmdl19-343.pdf https://ccronline.sigcomm.org/wp-content/uploads/2019/10/acm... [3] https://research.google/pubs/swift-delay-is-simple-and-effective-for-congestion-control-in-the-datacenter/ https://research.google/pubs/swift-delay-is-simple-and-effec...
- rwmj 2y agoOn a related topic, has anyone had luck deploying TCP fastopen in a data center? Did it make any difference? In theory for shortlived TCP connections, fastopen ought to be a win. It's very easy to implement in Linux (just a couple of lines of code in each client & server, and a sysctl knob). And the main concern about fastopen is middleboxes, but in a data center you can control what middleboxes are used. In practice I found in my testing that it caused strange issues, especially where one side was using older Linux kernels. The issues included not being able to make a TCP connection, and hangs. And when I got it working and benchmarked it, I didn't notice any performance difference at all.
- ezekiel68 2y agoI feel like there had ought to be a corollary to Betteridge's law which gets invoked whenever any blog, vlog, paper, or news headline that begins with "It's Time to..." But the new law doesn't simply negate the assertion. It comes back with: "Or else, what?" If this somehow catches on, I recommend the moniker "Valor's Law".
- deleted 2y ago[deleted]
- KaiserPro 2y agoTLDR: No, not its not. HOMA is great, but not good enough to justify a wholesale ripout of TCP in the "datacentre" Sure a lot of traffic is message oriented, but TCP is just a medium to transport those messages. Moreover its trivial to do external requests with TCP because its supported. There is not a need to have HOMA terminators at the edge of each datacentre to make sure that external RPC can be done. The author assumes that the main bottleneck to performance is TCP in a datacentre. Thats just not the case, in my datacentre, the main bottleneck is that 100gigs point to point isnt enough.
- javier_e06 2y agoIn your data center usually there is a collection of switches from different vendors purchased through the years. Vendor A tries to outdo vendor B with some magic sauce that promise higher bandwidth. With open standards. To avoid vendor lock-ins. The risk averse manager knows the equipment might need to be re-used or re-sold elsewhere. Want to try something new? Plus: Who is ready to debug-maintain the new fancy standard?
- cletus 2y agoNetwork protocls are slow to change. Just look at IPv6 adoption. Some of this is for good reason. Some isn't. Because of everything from threat reduction to lack of imagination equipment at every step of the process will tend to throw away anything that looks weird, a process somebody coined as ossification. You'll be surprised how long-lasting some of these things are. Story time: I worked on Google Fiber years ago. One of the things I worked on was on services to support the TV product. Now if you know anything about video delivery over IP you know you have lots of choices. There are also layers like the protocls, the container format and the transport protocol. The TV product, for whatever reason, used a transport protocol called MPEG2-TS (Transport Streams). What is that? It's a CBR (constant bit rate) protocol that stuffs 7 188 byte payloads into a single UDP packet that was (IPv4) multicast. Why 7? Well because 7 payloads (plus headers) was under 1500 bytes and you start to run into problems with any IP network once you have larger packets than that (ie an MTU of 1500 or 1536 is pretty standard). This is a big issue with high bandwidth NICs such that you have things like Jumbo frames to increase throughput and decrease CPU overhead but support is sketchy on a hetergenous network. Why 188 byte payloads? For compatibility with Asynchronous Transfer Mode ("ATM"), a long-dead fixed-packet size protocol (53 byte packets including 48 bytes of payload IIRC; I'm not sure how you get from 48 to 188 because 4x48=192) designed for fiber networks. I kind of thought of it was Fiber Channel 2.0. I'm not sure that's correct however. But my point is that this was an entirely owned and operated Google network and it still had 20-30+ year old decisions impacting its architecture. Back to Homa, three thoughts: 1. Focusing on at-least once delivery instead of at-most once delivery seems like a good goal. It allows you to send the same packet twice. Plus you're worried about data offset, not ACKing each specific packet; 2. Priority never seems to work out. Like this has been tried. IP has an urgent bit. You have QoS on even consumer routers. If you're saying it's fine to discard a packet then what happens to that data if the receiver is still expecting it? It's well-intentioned but I suspect it just won't work in practice, like it never has previously; 3. Lack of connections also means lack of a standard model for encryption (ie SSL). Yes, encryption still matters inside a data center on purely internal connections; 4. QUIC (HTTP3) has become the de-facto standard for this sort of thing, although it's largely implementing your own connections in userspace over UDP; and 5. A ton of hardware has been built to optimize TCP and offload as much as possible from the CPU (eg checksumming packets). You see this effect with QUIC. It has significantly higher CPU overhad per payload byte than TCP does. Now maybe it'll catch up over time. It may also change as QUIC gets migrated into the Linux kernel (which is an ongoing project) and other OSs.
- efitz 2y agoI wonder which problem is bigger- modifying all the things to work with IPv6 only or modifying all the things to work with (something-yet-to-be-standardized-that-isn’t -TCP)?
- stego-tech 2y agoAs others have already hit upon, the problem forever lies in standardization of whatever is intended to replace TCP in the data center, or the lack thereof. You’re basically looking for a protocol supported in hardware from endpoint to endpoint, including in firewalls, switches, routers, load balancers, traffic shapers, proxies, etcetera - a very tall order indeed. Then, to add to that very expensive list of criteria, you also need the talent to support it - engineers who know it just as thoroughly as the traditional TCP/IP stack and ethernet frames, but now with the added specialty of data center tuning. Then you also need the software to support and understand it, which is up to each vendor and out of your control - unless you wrap/encapsulate it in TCP/IP anyway, in which case there goes all the nifty features you wanted in such a standard. By the time all of the proverbial planets align, all but the most niche or cutting edge customer is looking at a project the total cost of which could fund 400G endpoint bandwidth with the associated backhaul and infrastructure to support it - twice over. It’s the problem of diminishing returns against the problem of entrenchment: nobody is saying modern TCP is great for the kinds of datacenter workloads we’re building today, but the cost of solving those problems is prohibitively expensive for all but the most entrenched public cloud providers out there, and they’re not likely to “share their work” as it were. Even if they do (e.g., Google with QUIC), the broad vibe I get is that folks aren’t likely to trust those offerings as lacking in ulterior motives.
- klysm 2y agoIf anybody is gonna do it, it's gonna be someone like amazon that vertically integrates through most of the hardware
- stego-tech 2y agoThat’s my point: TCP in the datacenter remains a 1% problem, in the sense that only 1% of customers actually have this as a problem, and only 1% of those have the ability to invest in a solution. At that point, market conditions incentivize protecting their work and selling it to others (e.g., Public Cloud Service Providers) as opposed to releasing it into the wild as its own product line for general purchase (e.g., Cisco). It’s also why their solutions aren’t likely to ever see widespread adoption, as they built their solution for their infrastructure and their needs, not a mass market.
- ghaff 2y agoThere was a ton of effort and money especially during the dot-com boom to rethink datacenter communications. A fair number of things did happen under the covers--offload engines and the like. And InfiniBand persevered in HPC, albeit as a pretty pale shadow of what its main proponents hoped for--including Intel and seemingly half the startups in Austin.
- tails4e 2y agoThe cost of standards is very high, probably second to the cost of no standards! Joking aside, I've seen this first thang when using things like ethernet/tcp to transfer huge amounts of data in hardware. The final encapsulation of the data is simple, but there are so many layers on top, and it adds huge overhead. Then stanrds have modes, and even if you use a subset the hardware must usually support all to be compliant, adding much most cost in hardware. A clean room design could save a lot of hardware power and area, but the loss of compatibility and interop would cost later in software.. hard problem to solve for sure.
- gafferongames 2y agoGame developers have been avoiding TCP for decades now. It's great to finally see datacenters catching up.
- gafferongames 2y agoDownvote away but it's the truth :)
- tonetegeatinst 2y agoMany I am misunderstanding something about the issue, but isn't DCTCP a standard? See the rfc here: https://www.rfc-editor.org/rfc/rfc8257 https://www.rfc-editor.org/rfc/rfc8257 The DCTCP seems like its not a silver bullet to the issue, but it does seem to address some of the pitfalls of TCP in a HPC or data center environment. Iv even spun up some vm's and used some old hardware to play with it to see how smooth it is and what hurdles might exist, but that was so long ago and stuff has undoubtedly changed.
- wmf 2y agoHoma has much lower latency than DCTCP.
- cryptonector 2y agoVarious RDMA protocols were all about that. Where are they now?
- SoftTalker 2y agoInfiniband is still widely used in data centers, RDMA and IPoIB. Intel tried Omnipath but that died quickly (I don't know specifically why).
- ltbarcly3 2y agoJust using IP or UDP packets lets you implement something like Homa in userspace. What's the advantage of making it a kernel level protocol?