8 ms·
This is a great write-up, but one thing I don't understand is why the effect of withdrawing the BGP prefixes was instantaneous (if I understand that correctly),
by throwdecro 5y ago
This is a great write-up, but one thing I don't understand is why the effect of withdrawing the BGP prefixes was instantaneous (if I understand that correctly), but it's taking hours (so far) to re-announce the prefixes. Why would it take so long to flip the switch back the other way?
- kube-system 5y agoGiven my experience with DNS issues, I am guessing that they are running into dependencies along the way that assume/require DNS be available to function.
- theshadowknows 5y agoYeah as part of my job I often have to work with our DNS team to provision say a subdomain or get some domain verified. They’ve got like…three people…trying to service thousands of teams across the enterprise. I do not envy their job at all.
- EricE 5y agoYour guys need BP Diamond IP http://www.diamondipam.com http://www.diamondipam.com It's an IP address management (IPAM) solution that also just happens to be a fantastic, federated (if you want) DNS management system too. Indeed a previous org I worked at bought it strictly to tame the DNS beast - local sys admins could control DNS for their subnets but not affect anything else. If we wanted, we could have had approval processes on top of the change requests - the system supported that too. I think the security teams finally woke up to the IP address management functionality and were slowly starting to integrate that into the rest of the infrastructure - but I was leaving around then. It was a fantastic system. One of the best hierarchical role-based access control systems in an application I have ever seen; the granularity was amazing yet it was easy to understand/administer. Not an easy trick!
- bink 5y agoWith routing it's even worse than that. If they had no out-of-band method to connect to these routers and they botched the routing config then they had no way to route any traffic to them at all. At least with DNS you can still connect to the IPs. I would find it a bit surprising if Facebook didn't have OOB access to their data centers, however.
- withinboredom 5y agoAssuming you don’t need DNS to get authorization to enter the OOB access…
- rifkiamil 5y agoI'm sure they got stuck in a security loop. To get access OOB passwords or IP address, they needed to get into a password vault that is under facebook AN. Next time FB save you passwords in OneDrive and Google Drive as a backup LOL. facebook-oob-password@gmail.com
- ajsnigrutin 5y agoSecurity policies are a pain in cases like this... Laptop + mobile tethering + serial cable to the router + teamvewier for the remote admin to get the access solves problems like this in minutes. Breaking a gajillion security policies by doing that is a different story though.
- skhr0680 5y agoDon’t you think that losing thousands, millions, or billions of dollars is a pain?
- Wolfspirit 5y agoI'm not sure if that is true (and I hope it is not cause that would be fatal) but I read somewhere that with facebook being down also means all internal infrastructure of facebook isn't available at the moment (chats, communication) including remote control tools for the BGP Routers. Therefor they require people to get physical access to the router while many people are working from home cause of the pandemic.
- Hikikomori 5y agoRestoring is just as simple as flipping the switch again, but access to that switch is another matter when your internal network is also down and you cannot even get access to your office or datacenters.
- cyanydeez 5y agoengineers will tell you that not everything is reversible even if theres no specific cqpacity issue
- foobarian 5y agoWhat I don't get is why there is no dead-man switch mechanism in place to roll back the configuration automatically unless someone confirms it positively. Kind of how screen resolution rolls back if you don't ack it. I used to always run a "(sleep 600; iptables -F) &" when messing with remote personal stuff just in case I lock myself out. I suppose with something like BGP it would be very difficult to get such a fallback working given how distributed the system is, and even more difficult to keep it exercised and tested.
- kelp 5y agoThis is a key feature of Junos on Juniper devices. It's called 'commit confirmed' and it will apply the config and then auto-rollback if you don't confirm it within a certain amount of time. https://www.juniper.net/documentation/us/en/software/junos/cli/topics/topic-map/junos-configuration-commit.html https://www.juniper.net/documentation/us/en/software/junos/c... But Facebook's network is certainly much more complex and automated than just doing one commit on one device. But I think they do still uses Juniper devices at the edge.
- deleted 5y ago[deleted]
- dec0dedab0de 5y agoI haven't been following closely, but I think once they moved the prefixes they could no longer access the routers. Coupled with barebones staff at the data center due to the pandemic, and all internal communication being disrupted. Though I really expected it to be up within an hour or two.
- dr_orpheus 5y agoYeah, I think that is true. If you look at the Update near the end of the Cloudflare article there is a huge spike in the BGP activity (I assume re-announcing all of the routes). So that part of it was relatively instantaneous after they got all of their ducks in a row actually getting to the routers and locating the BGP from some earlier version before it went offline this morning that they could use.
- rifkiamil 5y agoWe have had out-of-band management ports & networks design for decades! I know the feeling of driving 8 hours because I lost connection to the device I was configuring. https://en.wikipedia.org/wiki/Out-of-band_management https://en.wikipedia.org/wiki/Out-of-band_management
- ampdepolymerase 5y agoHN engineers believe out of band admin control planes to be surveillance and backdoor firmware so they are disabled for privacy reasons.
- jrochkind1 5y agoWho the hell put HN engineers are in charge of facebook, and why haven't we gotten more out of it than a temporary outage??
- ampdepolymerase 5y agoPreviously they had physical access to the data centers and weren't locked out.
- adamcharnock 5y agoI’m pretty new to BGP, but I’d imagine that cutting off access to an AS is fast because all it takes is for the neighbouring routers update their routes. At which point any traffic that makes it that far is simply dropped. Whereas to make an announcement, the entire internet (or at least all routers between the AS and the user) need to pickup the new announcement. (Note: I still need to read the article)
- bsedlm 5y ago(I'm trying to better understand this) I think it's not so simple because authoritative DNS systems are involved. So it's not just a BGP error. It's a BGP error which disconnected authoritative DNS for all facebook. I'm not quite sure why that makes it so slow to fix. is it just because internal difficulties due to having no DNS at all?
- Hikikomori 5y agoOnce they start adversiting again it should only take a few minutes at most for most ISPs to get it.
- p4bl0 5y agoBut then it is the DNS that had to propagate from the now accessible authoritative servers.
- Hikikomori 5y agoShouldn't take long once they are up and responding again?
- WorldMaker 5y agoI'd assume it is a cache invalidation problem at that point: from my lay understanding BGP probably needs to busts caches on a withdrawal to keep traffic from going to black holes and prevent DDoS attempts, but probably has to wait for TTL timeouts to cache new routes (and those TTLs are going to vary by whatever cache systems the other ASes are running not the timing of Facebook's AS sending the new [old] routes.).
- PeterCorless 5y agoIf an authoritative DNS entry was removed, it can take up to 72 hours for that change to be propagated around the world, though usually just a few hours for some other authoritative DNS systems to get you mostly back: https://ns1.com/resources/dns-propagation#:~:text=DNS%20propagation%20is%20the%20time,typically%20takes%20a%20few%20hours https://ns1.com/resources/dns-propagation#:~:text=DNS%20prop....
- tester756 5y agoWhy it takes this long?
- withinboredom 5y agoCaching
- mnordhoff 5y agoResolvers typically cache successful "does not exist" responses for no more than 1-3 hours. (And authoritative servers often have a lower negative TTL.) (There's a corner case related to DNSSEC that can make it go higher, but that's being worked on, and isn't relevant here.) In this situation, the nameservers were just down. I haven't done exhaustive research, but the resolvers I'm aware of cache that kind of thing for no more than 15 minutes.
- withinboredom 5y agoIf there’s a chain of caches 3 deep, a 15 minute cache on bad responses will take 45 minutes to clear.
- gorgoiler 5y agoAt a guess: reconnecting traffic at the billions-of-people scale has the potential for finding all sorts of weird behaviours. For all we know, it has been connected and disconnected 10x already during the outage, with each reconnect overwhelming some new, deeper level of the system each time. Reconnect and the GLBs fall apart under load as the entire world’s cadre of recursive resolvers hit you. Fix that. Reconnect again. This time your LBs have marked half the servers as offline because their heartbeats have been failing. Fix that. Reconnect again. Now all the memcache data is hours old and so the site business logic fetches straight from databases, knocking them over. Fix the databases. Reconnect again. Ad nauseam.
- Godel_unicode 5y agoI find this kind of uninformed conjecture amusing on a thread full of people complaining about cloudflare doing the same (they didn't). There's no evidence of this kind of flapping behavior in any of the telemetry I've seen posted by network engineer friends, and the blog post explicitly calls out when they saw the BGP updates that brought the site back online.
- atf104 5y agoFor me, the hardest part to believe is there was literally no one on-site at their datacenters. Really? No one? At this scale, literally, there has to be a security guard there who can kick the door open.
- acomjean 5y agoReminds me of when I was working at a startup we DOSed ourselves. We had devices that monitored power minute by minute. A load balancer was mis configured and we went down for a some hours. We came back up and the devices all saw that and flooded us with all the data they’d been storing since we were down.. down again… bring it up and down again…. we needed a better fix. I was up till 2 am with the other developer coding a fix.. the next morning we talked to the CEO (CTO was on vacation)who told us upgrading the was database was going to be too expensive… good times. We did get a firmware fix ( to be honest we were running out of cash..) I can’t imagine trying to restart something as big as facebook…
- nimbius 5y agoUnpopular opinion but I think a talent exodus and turnover are likely at fault here as well. The people who stood up this juggernaut are no longer here, and Facebooks ability to consistently attract talent that is competent enough to maintain such a formidable beast is hampered by repetitive revelations that it operates at the net-loss of humanity as a whole. Facebooks no longer an innovator, just a mining operation with a dwindling population of hateful elderly and bots.
- bsedlm 5y agoI'd say the same thing about Google.
- erhk 5y agoYes well lets have this outage again in that flavour in say... 11 months?
- bhawks 5y agoTaking a bet that complex systems will fail is usually free money :). Having been a part of Google the only thing more awe inspiring than the sheer complexity of production is the fact that it all worked so well. This is not a dig at current Googlers but entropy is cruel and uncaring. Perhaps parts of the stack which have been kept fit & fresh in people's minds due to constant rewrites will last longer but there are tons of places in the depot that are unowned despite serving production query traffic and the number of engineers that have any context to support it grows smaller over time.
- cornel_io 5y agoGoogle (at least the good parts) has a pretty good culture of documentation, which helps a bit. While it drove me batty to need to spend a month writing up 10 page specs and getting sign off from directors for a minor feature that nobody would see and would take 3 days to code up, it was nice to be able to trace through the historical evolution of abandoned features when trying to figure out how they worked. No idea if Facebook is similar.
- termau 5y agoI feel like it just confuses the issue with a bunch of unnecessary babble about DNS, theres better ways to read about how BGP works without confusing a bunch of different things. The only part of the article that was relevant was 'Routes were withdrawn' - the rest being a consequence of that.
- foolfoolz 5y agofast off board slow onboard is a pattern you find all over. it has roots in fraud but there’s many reasons. getting more access is a privilege escalation and requires some trust to achieve
- cortesoft 5y agoYou have to be careful turning something as large as Facebook back on. If you turn on announcements one place first, the entire internet will try to reach you through a single transit and overwhelm it.
- dijit 5y agoThe kind of tail you’re talking about is baked into DNS at least. I don’t know enough about BGP to make an informed decision; but at the point the outage is noticed it’s entirely possible that the system has been unavailable for quite some time already.
- belorn 5y agoIt just a guess, but from experience with BGP and associated redundancy systems that most likely was in place, if everything doesn't return immediately then you have a big fight on your hand to not only stop the non-functional redundancy but also reestablish the peer connections with associted hearthbeat/processes for establishing and maintaining the peer connection. My understanding from what people write about configuring BGP and the system around it seems to imply that the best practice in this circumstance is to kill everything, fix the original error and then turn on things slowly again. Then fix the broken redundancy configuration. Then test the redundancy system regularly in the future.