5 ms·
Who's their network switch vendor? I'm not a networking expert, but boy - it sure seems like their switch vendor has screwed some things up royally. Or perhap
by dos1 14y ago
Who's their network switch vendor? I'm not a networking expert, but boy - it sure seems like their switch vendor has screwed some things up royally. Or perhaps this is common with all complex network topologies regardless of hardware vendor?
EDIT: I would just like to say, along with others, I greatly enjoy their postmortems and I feel as though I learn something every time. Kudos to them for being forthright. I host my personal and professional projects with them and am supremely confident that my data is as safe with them as it is with anyone.
- rdl 14y agoThey are hosted at Rackspace, so I think the default is "Vendor C". I'm curious if they'd even be able to support an Arista network. Although this would probably not be a problem with that kind of network.
- sounds 14y agoThey don't use Cisco for their Aggregation switches.
- rdl 14y agoWhen did that change? They did when I last checked (which, admittedly, was Some Years Ago.)
- nixgeek 14y agoRecently.
- rdl 14y agoTo what, some other vendor for price? (oh, I see you are with GitHub. Somehow I suspect you did not push for this downgrade.) (followup to sounds: I'm not particularly pro cisco except that all-cisco lets you avoid interop problems. Since at least some of this is cisco, there's a good argument for all cisco. I've actually used HP and Dell successfully for certain things, but only because I kept it as simple as possible. In the past using other stuff was necessary at the high end, too, because cisco didn't provide 10GE very well, etc. What I am against is cheaping out on your ~few aggregation switches, which I've seen other people do, when you're already spending good money on everything else. Maybe it's different once you get to Google or Facebook scale, but I've never been responsible for quite that many switches, and an extra $500-$1000 per 24 ports isn't going to kill you. I actually prefer non-cisco substantially for routing and for security products.)
- sounds 14y agoI'm not with GitHub. Please don't take this the wrong way – can you explain some of your experiences with Cisco? I sense you're pretty loyal to Cisco. I've had the exact opposite experience, getting rock solid uptime and no-compromises performance from commodity hardware. My experience with Cisco is typified by the IOS command line: hard to use, even harder to hire someone who makes it easy, and every minute you feel the pull on your wallet. The one place Cisco seems to shine is at the boardroom table when negotiating contract renewals. :) Again, I don't mean to offend.
- winthrowe 14y agoNot the person you replied to, but in my experience the other place Cisco shines (relatively speaking) is in hellish environments. This isn't so much applicable to datacenter usage, but to serve some of our satellite offices reqires placing equipment in places that I'd rather not have to have any hardware; in these situations cisco has been much more reliable than the HP switches we migrated away from.
- rdl 14y agoIt's also way easier to find people with experience using Cisco equipment in "moderately complex" environments than HP/Dell/Arista/Juniper/Force10/Foundry/Extreme/etc.
- thaumaturgy 14y agoGitHub had a switch-related outage earlier in December, too: https://github.com/blog/1346-network-problems-last-friday https://github.com/blog/1346-network-problems-last-friday I've seen some really wonky behavior in Cisco equipment even in a really small ISP network, but I have no idea if this is normal or not. Maybe one of the very few people that's responsible for a network of GitHub's size could chime in; I'd be interested too.
- nixgeek 14y agoAdapting a legacy network to deal with growing pains is hard. Anyone who says different is probably lying, or hasn't administered a network outside of 'bedroom scale' previously. Whilst everyone prefers to work with a blank canvas that is a rare opportunity, and one which still has constraints like budget limitations or "We need this by Wednesday!". All vendors have their foibles. No vendor is perfect. Datasheets are rarely 100% accurate, behaviour fluctuates based on what code train/version you're running on devices. Network vendors are like everyone else - humans - so just like everyone else they break things occasionally too. One of the best ways to mitigate this as demonstrated by some of the largest players is to have N+1 physical sites and to be able to control traffic distribution to each of them, including taking one entirely offline without your customers noticing that it's happened. It's also extremely hard to do when you start looking at how CAP theorem applies to your given application, and not just the infrastructure changes required to network it all together, but the application changes required to not have it just explode on you.
- 23david 14y agoAgreed that technical debt is a pain to deal with. But it's definitely possible to migrate and upgrade live systems without significant technical risk. The company just needs to make the decision to invest the proper resources and decide to have executive-level focus on Operations. From your comments it sounds like this was a failure of management to properly plan out these changes. Lack of resources shouldn't be an issue here, since clearly github's $100 Million+ in funding is enough to build a reliable HA system with adequate redundancies and hire the people who are qualified to run it. There just isn't anyone to point fingers at, including vendor screw-ups. All of those items would be addressed in a proper risk assessment plan.
- sounds 14y agoThis is likely the reason the large cloud companies all do testing that actively causes outages. I'm over-simplifying on purpose here: this is something that requires a lot of thought. At first glance that seems foolish, but to quote you, "complex network topologies" are very prone to falling over badly. Since they all seem to be a one-off custom setup these days, how can you be sure it won't fall over? Here are the testing approaches I know about: 1. Netflix Chaos Monkey: http://techblog.netflix.com/2012/07/chaos-monkey-released-into-wild.html http://techblog.netflix.com/2012/07/chaos-monkey-released-in... - but that doesn't mean Netflix has it all together. They still have outages. 2. Google does it. They have a team that goes around unplugging network cables and monitoring how fast the engineers can find and fix the problem. I can't dig it up but it was only a few months ago - hey, Google, your search engine can't find an article about you. :)
- tonfa 14y agohttp://www.wired.com/wiredenterprise/2012/10/ff-inside-google-data-center/all/ http://www.wired.com/wiredenterprise/2012/10/ff-inside-googl... (search for DiRT).
- madsushi 14y agoThat was my question too. My guess is that it's Arista, based on the MLAG, ISSU, and 'agent' references. There's a chance it could be Cisco or Juniper, but I don't think so. The author of their post-mortems does a very good job of removing any vendor-specific language that would give it away.
- regularfry 14y agoIt doesn't sound to me like the switch upgrade was the problem here. Yes, it was the trigger, but the massive failure was caused by the failover setup. The learning seems to be that if you're going to have automated failover, it must go in one direction only. That means you're guaranteed never to get flaps or mexican standoffs, and determining which node was active is a (relatively) simple question of "did a failover happen?" I'm sure there's a downside to this, but I can't see one that outweighs the gains from having much simpler failure modes.
- lspruel 14y agoAs others have mentioned, it is indeed Arista: https://twitter.com/markimbriaco/status/276416483939213312 https://twitter.com/markimbriaco/status/276416483939213312 I think they should be more forward about naming and shaming, but I understand why they'd rather not. Personally I have horrible experience with Nortel/Avaya and would recommend anyone against using their equipment for core switching. The least worst, again in my experience, seem to be Cisco and Juniper.