4 ms·
I'll go out on a limb: inside datacenter on your own hardware, you can safely ignore low-level pedantry and mostly ignore “weird networks” and use TCP as two-wa
by hamilyon2 2y ago
I'll go out on a limb: inside datacenter on your own hardware, you can safely ignore low-level pedantry and mostly ignore “weird networks” and use TCP as two-way Unix pipe.
“Mostly” because you still care about bandwidth limits and packet RPS limits and latency of course.
- toast0 2y agoI wouldn't, unless you've got a really solid understanding of your datacenter network and it's 100% good all the time. Which is unlikely, from my experience as a server person. If you've got dirty optics between two switches, now you're getting packet loss and TCP rears its head. Hopefully it's not an issue now, but diagnosing microbursting[1] was lots of fun, and really wigs TCP out. I've also run into 'fabric congestion'. My true favorite though is when you've got 2x aggregation on servers, and 4x aggregation for top of rack switches to spine switches, so there's 8 paths in each direction between two servers in adjacent racks, and only one path (sometimes in only one direction) is only running at 99.9%. That's a real PITA to track down unless you have visibility into switching metrics. [1] https://en.m.wikipedia.org/wiki/Micro-bursting_(networking) https://en.m.wikipedia.org/wiki/Micro-bursting_(networking)
- macintux 2y agoI was always suspicious about self-hosted high availability solutions (typically just diagrams, not yet implemented) that included redundant switches. Given how generally reliable switches are, I was inclined to believe that a misconfiguration or flaky network cable on one switch was more likely to cause a downtime (or significant degradation) than an outright switch failure, so adding another switch was doubling the chances of trouble and, as you note, making it harder to troubleshoot.
- toast0 2y agoIt kind of depends. You do get some weird stuff to debug, and more connections = more likely that one of them is broken. Otoh, if you ever do any scheduled maintenance on your switches (which is likely if they're doing anything fancy), having properly setup redundancy means you can announce a likely brief loss of redundancy, rather than a likely brief full loss of connectivity. If you have the right knobs, you can gracefully fail out the switch under maintenance and everything goes smoothly. Of course, sometimes you reboot the redundant switch and it confuses the other one and servers lose connectivity anyway.
- to11mtm 2y agoAgreed. Having done Akka.NET Remote/Cluster setups in prod that survived multiple 'new to the org' categories of DC Failures at their level of scale/capacity [0] there's a lot to account for if you want to keep everything happy and visible [1][2][3] [0] - Cut fiber between DCs, Rack failures due to IO-ish type issues, bad switches... at least 2 out of 3. [1] - The upshot was we were able to survive all of the scenarios in at worst a degraded state, Once or twice we needed a restart. [2] - We also had enough metrics going on that we could detect DC/server outages about as quickly as whoever actually was monitoring the failing subsystem. [3] - But here's the funny rub. An APM tool was the Achilles heel for both our Akka Links, as well as our SQLServer connections. Once they installed an 'agent' we more frequently had to do a 'full cycle' to clean things up after an outage, or even an MSSQL Server reboot. After I left the shop I got confirmation that yes, the APM module was the problem.
- toast0 2y ago> We also had enough metrics going on that we could detect DC/server outages about as quickly as whoever actually was monitoring the failing subsystem. Yeah, my Erlang clustering experience was that we (the customer) were the monitoring system for the DC/managed hosting provider. Although, by the time we left there, they would have outage notifications before we put in tickets.