9 ms·
Route Leak Impacting Cloudflare
- dajonker 7y ago"Cloudflare is observing network performance issues." not just performance issues, our entire website is unavailable because of it. edit: availability has been alternating between available and unavailable
- monkin 7y agoIt's not their fault as accidents happen to everyone. You should be prepared for such a scenario and for many others too.
- dajonker 7y agoYou are right, accidents happen to anyone. I cannot really be prepared for Cloudflare to go down though. What are my alternatives? Turn it off and route traffic to our servers directly? The DNS propagation takes longer than it just took for our website to be available again.
- jasongill 7y agoIt shouldn't - Cloudflare keeps the TTL for their cache-enabled records very low (like 300 seconds). If you just log in to Cloudflare and click the "orange cloud" icon on the DNS tab, which points the domain back directly to your origin, you'll see the site up within a couple minutes.
- yjftsjthsd-h 7y agoIs 300 very low? I've occasionally seen 60 in the wild.
- jasongill 7y agoIt's very low compared to 24 hours, which is what used to be the most common setting and that (among other factors) was a big part of the "DNS propagation takes forever" mentality
- tech5000 7y agoSame here.
- markplindsay 7y agoThis morning I'm finding out just how many of our supporting services rely on Cloudflare as well.
- deleted 7y ago[deleted]
- sudhirj 7y agoIsn't HN on Cloudflare? How are we reading about a CF outage on a site that runs behind CF?
- nodesocket 7y agoSeems to be intermittent, and perhaps dependent on which POP you are hitting. I have been getting PagerDuty alerts that are flapping between triggered and resolved.
- srushtika 7y agoDoesn't seem so http://www.doesitusecloudflare.com/?url=https%3A%2F%2Fnews.ycombinator.com%2F http://www.doesitusecloudflare.com/?url=https%3A%2F%2Fnews.y...
- sudhirj 7y agoI've seen the "checking your browser" on HN message quite a few times, especially when I use a VPN, and I'm pretty certain it's the CF one. Maybe it's turned on selectively for high traffic / bot situations?
- inumedia 7y agoFWIW, this site doesn't seem accurate. I used it against my site that is definitely using cloudflare and it incorrectly reported it as not :p
- has-ams 7y agohttp://www.doesitusecloudflare.com/?url=cloudflare.com http://www.doesitusecloudflare.com/?url=cloudflare.com
- jakejarvis 7y agoI just put in my own Cloudflare-enabled website and it came back negative as well.
- alternize 7y ago
- fluxsauce 7y agoI was going back and forth about whether to post this; we're clearly experiencing some sort of partial outage, New Relic Synthetics shows everything offline (except for Sydney for whatever reason), but the actual webserver logs indicate healthy traffic and the sites load when I visit it. "Cloudflare is observing" is pretty opaque.
- nikisweeting 7y agoWe're seeing the same situation with our alerts.
- jgrahamc 7y agoThis appears to be a routing problem. All our systems are running normally but traffic isn't getting to us for a portion of our domains. 1128 UTC update Looks like we're dealing with a route leak and we're talking directly with the leaker and Level3 at the moment. 1131 UTC update Just to be clear this isn't affecting all our traffic or all our domains or all countries. A portion of traffic isn't hitting Cloudflare. Looks to be about an aggregate 10% drop in traffic to us. 1134 UTC update We are now certain we are dealing with a route leak. @dang etc.: could someone update the title to reflect the status page "Route Leak Impacting Cloudflare" 1147 UTC update Staring at internal graphs looks like global traffic is now at 97% of expected so impact lessening. 1204 UTC update This leak is wider spread that just Cloudflare. 1208 UTC update Amazon Web Services now reporting external networking problem https://status.aws.amazon.com/ https://status.aws.amazon.com/ 1230 UTC update We are working with networks around the world and are observing network routes for Google and AWS being leaked at well. 1239 UTC update Traffic levels are returning to normal.
- holycowzer 7y agoThanks for the updates. I wish I could get this information somewhere other than hacker news though. :(
- sauldcosta 7y agoWe use downdetector.com because status pages tend to take up to an hour or so to update, if they ever do.
- jgrahamc 7y ago1042 UTC First alert of global traffic problem 1057 UTC Internal group chat room up and running 1102 UTC Status page updated So, first alert to status page was 20 minutes.
- jgrahamc 7y agoThe team is updating the status page but not with granular detail because they'd have to spend time discussing what to say. I'm giving you the blow by blow.
- nnx 7y agoSeems to be BGP/routing related, some networks can access CloudFlare networks normally.
- sudhirj 7y agoHow would that kind of disruption happen? Someone else also anycasting the CF IP addresses?
- arghwhat 7y agoBGP route/prefix leaks. BGP is the protocol that deals with routing across the various internet backbones (known in the protocol as "autonomous systems", identified by an AS number). On that protocol, the various systems broadcast what prefixes they can route, which then affects the rest of the networks' routing decisions. By error or malice, a system can report a prefix they cannot or should not route, causing other systems to start routing traffic across it. This will either just cause weird routes (such as ones going through certain suspicious countries), cause poor performance for those routed, or no connection at all for those routed.
- nikisweeting 7y agoAt 3-4 major leaks per year it seems like we should probably fix BGP one of these days...
- purerandomness 7y agoThe way I understand it it's not BGP, it's mostly human error, or malicious intent. The protocol is fine.
- arghwhat 7y agoMalicious intent could likely be mitigated with cryptography. E.g., publishing a prefix requiring a signature from its owner. Such system would also contain human error to a smaller set of possible faults.
- nikisweeting 7y agoAn unauthenticated protocol that allows unsigned routes to be blindly accepted is not a good protocol, that's why Cloudflare has been pushing RPKI for a while https://blog.cloudflare.com/rpki/ https://blog.cloudflare.com/rpki/ https://blog.cloudflare.com/rpki-details/ https://blog.cloudflare.com/rpki-details/
- colinodell 7y agoThis seems like a partial outage, likely region-based. We have a large number of sites routed through Cloudflare and I can access all of them from home, but our HTTP monitoring software reports the sites as down.
- tomcam 7y agoIt’s caused my site monitoring via PagerDuty to go insane, with texts sent every few minutes.
- nikisweeting 7y agoOne of the weirder leaks I've seen, 8.8.8.8 and 1.1.1.1 are both down for me, but everything else is working fine.
- cityzen 7y ago8.8.8.8 is google’s DNS, though.
- nikisweeting 7y agoExactly, that's why it's weird. It would have to be a huge range leaked to get both 8.8.8.8 and 1.1.1.1, surprising that a peer didn't filter it before it worked it's way up the chain.
- ipmb 7y agoSeeing ~60% drop in traffic here.
- steelaz 7y agoDisabling HTTP proxy and leaving "DNS only" option in CloudFlare DNS settings solved the problem for us.
- hugoromano 7y agonot very safe for some users.
- larsthorup 7y agoDoesn't appear to work for us
- Dunning-Kruger 7y agoThe current it stack needs a do-over. These outages already happen on accident often because of human error. Imagine the damage a state actor could inflict by targeting these large data centers. I hope that some of the newer decentralized cloud startups like dfinity or storj takes over.
- jgrahamc 7y agoOne of the reasons we're pushing: https://blog.cloudflare.com/rpki/ https://blog.cloudflare.com/rpki/
- forgottenpass 7y ago>The current it stack needs a do-over. "The network is unreliable" is a rule of thumb that was drilled into my head in network programming class. It always has been, it always will be. Doesn't matter if it's the internet or the link between your computer and a device sitting on your desk. And it doesn't matter what the tech is. Making the internet more resilient only increases the severity of the failure when organizations that don't understand the risk they're taking on experience network outages. The network is unreliable.
- filistar 7y agoDoes anyone know which global sites were unavailable because of Cloudflare crash?
- nikisweeting 7y agoYou're not going to be able to get a solid list, this is a different category of problem than something like CloudBleed, and even then the list wasn't solid. This issue is affecting AWS, Cloudflare, Cloudflare DNS, Google DNS, and the tens of thousands of other services that depend on them, but it's region specific and will break different things for different users as the leak propagates.
- cntlzw 7y agoOne source: https://twitter.com/atoonk/status/1143143943531454464 https://twitter.com/atoonk/status/1143143943531454464 90 AS 13335 Cloudflare, Inc. 18 AS 7018 AT&T Services, Inc. 8 AS 63949 Linode, LLC 8 AS 2828 MCI Communications Services, Inc. d/b/a Verizon Business 6 AS 26769 Bandcon 6 AS 16509 Amazon.com, Inc. 4 AS 6428 CDM 4 AS 2914 NTT America, Inc. 2 AS 9808 Guangdong Mobile Communication Co.Ltd. 2 AS 6939 Hurricane Electric LLC 2 AS 62904 Eonix Corporation 2 AS 55081 24 SHELLS 2 AS 54113 Fastly 2 AS 46606 Unified Layer 2 AS 45899 VNPT Corp 2 AS 4246 New Jersey Institute of Technology 2 AS 3257 GTT Communications Inc. 2 AS 27695 EDATEL S.A. E.S.P 2 AS 22781 Strong Technology, LLC. 2 AS 20473 Choopa, LLC 2 AS 16625 Akamai Technologies, Inc. 2 AS 12129 123.Net, Inc.
- nikisweeting 7y agoI think it's like 2.4k ASNs at this point, each with 10s-1000s of IPs, I guess you can make a list from that but it's going to be as unreliable as the Cloudbleed list was. Also not always easy to do reverse hostname lookups from the IPs to see the site names.
- breakingcups 7y agoIt's not really fair to call it a Cloudflare crash.
- filistar 7y agoDoes anyone know which global sites were or still are unavailable because of Cloudflare crash? Maybe some media sites?
- taf2 7y agoI’m pretty sure 1.1.1.1 for dns is impacted by this. Initially I thought my WiFi was having issues this morning until I realized it must be dns switching 1.1.1.1 out and besides the cloudflare sites everything is normal again
- foobarbazetc 7y agoIt’s definitely all countries, just for a specific range of anycast IPs. Our CloudFlare stuff isn’t even pingable. Sometimes it’ll return an echo from a far away DC. It’s been like this for over an hour now and your status page doesn’t even acknowledge it apart from “Network performance issues”.
- nikisweeting 7y agoIt was updated a few minutes ago confirming that it's a route leak.
- foobarbazetc 7y agoYeah but they just wasted an hour of everyone’s lives trying to figure out WTF was going on at 3:34am. (The average CF user has no idea what a route leak is, tbh.)
- corobo 7y ago"Everyone" speak for yourself, middle of the workday here :P
- nikisweeting 7y agoWhat's weird is that 8.8.8.8 is also intermittently down for me. Are other people having issues with Google DNS too? https://i.imgur.com/3ySmVLW.png https://i.imgur.com/3ySmVLW.png
- RKearney 7y agoGoogle rate limits ICMP to 8.8.8.8. It’s not meant to be used as your personal “is the internet up” test.
- nikisweeting 7y agoI use my own server or 1.1.1.1 for uptime checks, but 8.8.8.8 was my DNS fallback when 1.1.1.1 went down, which then meant I had no DNS working at all, which is why I noticed and tried pinging them.
- tyingq 7y agoHe probably pinged it because DNS wasn't working. It is meant to be his personal DNS resolver.
- mindslight 7y agoSeriously? If true that's an awfully quick bait and switch, even for Google.
- ceejayoz 7y agoHow is it bait and switch? 8.8.8.8 was never marketed as a "ping me to see if the Internet is up" service, as far as I know. Just as a fast, public DNS server.
- mindslight 7y agoAn important use of well known easy to type IP addresses is when you're mucking around to figure out if your upstream network isn't working. I could see if they attempted to set a new standard by just not responding to ICMP at all (although turning around an icmp echo takes less work than a DNS lookup...), but responding intermittently is actively harmful.
- pgt 7y agoWhat is a route leak?
- nikisweeting 7y agoThe Border Gatway Protocol is what network providers use to announce which IP ranges they can route traffic for. The problem is it's almost totally unauthenticated, so rogue ISPs and network operators can suddenly take over parts of the internet by "leaking" routes for ranges they shouldn't be able to control. They do this by announcing something like "send me all traffic for 1.1.1.1 - 1.1.1.255", and if their peers don't verify it, they'll just start routing that traffic to them. Peer by peer, the route then propagates and a larger portion of the internet, and as routers learn the new bad route, more of the traffic to those IPs gets sent to incorrect network.
- andreareina 7y agoHere ya go https://en.wikipedia.org/wiki/BGP_hijacking https://en.wikipedia.org/wiki/BGP_hijacking
- robbiemitchell 7y agoUnless it’s a total coincidence, this looks like this is affecting some Amazon services, including Sagemaker notebooks and Echo devices.
- kristofferR 7y agoI realized that it would be an issue like this when downforeveryoneorjustme.com didn't load either.
- emilstahl 7y ago“AS396531 "Allegheny Technologies Incorporated" is leaking a better-reachable route for AS13335 "Cloudflare, Inc." towards AS701 "Verizon Business/UUnet" explaining the current LSE going on.” https://twitter.com/OhNoItsFusl/status/1143117619106652160 https://twitter.com/OhNoItsFusl/status/1143117619106652160
- jgrahamc 7y agoThey aren't the original leaker. Update: sorry, I may have been wrong. Hard to see clearly in the fog of BGP.
- mbell 7y ago> AS396531 - Allegheny Technologies Incorporated That appears to be a steel/alloys company. Why are they operating BGP equipment?
- bin0 7y agoAny company which operates large factories probably has its own ASN and runs its own networks. Every thing's gotta be internet-enabled these days, and at a certain scale, it becomes cost-effective.
- Hikikomori 7y agoWhy not? Pretty much everyone that needs a redundant internet connection (dual ISP) does it.
- ussrlongbow 7y agoExperiencing around 60% traffic drop on customer's sites.
- lordelph 7y agoWe've been evaluating Cloudflare mainly for doing failovers faster than DNS. This morning I ran some tests to generate graphs to show the typical delay incurred in preparation for a show-and-tell with some key people. I started seeing delays of up to 300 seconds! At best there was a 1 second delay. I wondered if I was going to have present "Why we've decided not to go with Cloudflare!" Any longtime Cloudflare users comment on how rare an event this sort of thing is? It seems rare from eyeballing the recent alert history.
- awinder 7y agoI had to switch off 1.1.1.1 for the first time because of this, can’t speak to their enterprise stuff but if dns resolver is a good test this is the first event I’ve hit since launch
- shdon 7y agoThings like this are not unique to CF and actually originate from outside their network. It does happen every once in a while, but I have far more confidence in CF's ability to resolve it than my own. They have the clout in the industry, the connections and the expertise to deal with this kind of thing. I've been with CF since late 2011 and am quite satisfied with their services.
- lordelph 7y agoThat's a good point, and I must admit I didn't know what a route leak was or that it could inflict this kind of damage. I appreciate now it's not CloudFlare's fault, and my hat is off to the CTO for posting more detail here. On the plus side, I did get to test the "Pause CloudFlare" button in a real-world scenario!
- veswdev 7y agoLongtime Cloudflare user. This isn't specific to Cloudflare, but common sense would be to always have a backup. For example, my sites I've got Cloudflare in front, but in the background I'm caching all my content and pushing to BunnyCDN, so if I need to fallover, I can safely fallover out of the network into a live cache (each request I re-populate cache in background job). It's saved me lots of time and energy.
- swixm 7y agoWhat a great idea it is to have half the internet behind Crimeflare! It shows!
- asymptotically2 7y agoI came here to comment the same thing. Cloudflare is too big.
- joepie91_ 7y agoWhile I agree with the general sentiment - and I've certainly publicly and loudly expressed it in the past - this particular incident can't actually be blamed on that. It's a route leak, which can affect any arbitrary amount of ISPs, because the BGP protocol is totally unauthenticated.
- tomschlick 7y ago> Crimeflare Care to explain that one?
- JakeTheAndroid 7y agothere is a site called crimeflare that can explain it as well as it can be explained.
- shacharz 7y agoHi Shachar from Peer5 here, we're operating a MultiCDN. Cloudflare is actually one of the best performing CDNs. All CDNs encounter issues small to big - that's why using multiple providers and intelligently routing between them is critical for high resilience.
- shacharz 7y agoRight now we're seeing issues in the following ASNs: 9,541 . 59,257 . 38,264 . 132,165 . 23,888 . 55,714 . 45,773 . 45,669 . 9,260 . 58,895 . 17,557 . 38,547 . 38,193 . 135,407 . 23,966 . 7,590 . 136,525 .
- nikisweeting 7y agoAre you seeing ASN 396531 as the original leaker?
- jcalabro 7y agoThis is affecting far more than just Cloudflare.
- deleted 7y ago[deleted]
- jbergstroem 7y agoAs an aside, there is something to be said about the size of Cloudflare when global network problems are reported as being Cloudflare issues.
- nikisweeting 7y agoThere's one thing I don't understand about this all, it looks like Allegheny Technologies Incorporated (AS396531, a suspected original leaker) was originally announcing 192.92.159.0/24. How the heck did their peers not manage to filter a sudden announcement for a range big enough that it managed to snag both 8.8.8.8 and 1.1.1.1. Do upstreams really allow a tiny /24 AS to randomly announce a /4 and get away with it? Or am I misunderstanding something fundamental about how BGP routes are allowed to propagate?
- slenk 7y agoThis is the problem with BGP
- kbirkeland 7y agoLeaking a /4 into BGP would do basically nothing unless the originator was originally advertising a /4. IP forwarding is based on the longest-prefix match. Since allocations are sized from /8 to /24, anybody actually advertising their space would not get hijacked by a /4. The leaker would just get traffic destined toward non-advertised networks.
- nikisweeting 7y agoThen my next question is: If they didn't leak a massive range, then why was it a big problem? I assume if they leaked a bad /24 it surely wouldn't be enough to take down Cloudflare and Google for everyone... no? Did they just leak tons of bad /24s or was it something else?
- kingbirdy 7y agoMy understanding is they had an optimizer that broke the /4 down in to /24s and those got announced
- nikisweeting 7y agoAha! That was the missing piece in my understanding, it all makes sense now! <3 You're the only person out of the ~5 people I asked who explained that bit.
- peterwwillis 7y agoHey CloudFlare: this page is dependent on ajax.googleapis.com, and if js is disabled, googletagmanager.com. (Also, weirdly, they still have a link to Google+ posts?)
- jgrahamc 7y agoFinal update from me. This was a widespread problem that affected 2,400 networks (representing 20,000 prefixes) including us, Amazon, Linode, Google, Facebook and others. https://twitter.com/bgpmon/status/1143149817473847296 https://twitter.com/bgpmon/status/1143149817473847296 https://twitter.com/bgpmon/status/1143149817473847296 https://twitter.com/bgpmon/status/1143149817473847296
- shacharz 7y agoIs the specific IP range of the leak known ?
- throwawayflower 7y agoThe main internet and phone service provider of the Netherlands is down. Even the emergency number (112, our equivalent of 911) is down. Almost everyone is unreachable. The whole telephone network is disrupted. I wonder if it's related to this? It does say this kind of BGP thing can be a deliberate malicious attack. Perhaps this? https://en.wikipedia.org/wiki/BGP_hijacking https://en.wikipedia.org/wiki/BGP_hijacking
- throwawayflower 7y agoOh, and the country's train and public transport infrastructure is experiencing some major problems too due to the phone service outage.
- x86_64Ubuntu 7y agoYou have to wonder if these outages aren't the result of hostile states laying the groundwork and testing the viability of certain attacks.
- nikisweeting 7y agoHeh I think based on BGP's track-record, if a state-level actor wanted to mess up everyone's BGP routes they wouldn't have to try very hard...
- jgrahamc 7y ago@dang etc. Be good if someone changed the title here. 2,400 networks were affected (including parts of Cloudflare, Google, Amazon, Linode, Facebook, ...).
- jiveturkey 7y agoInteresting, isn't it, that when it's a US based steel plant, it's a route leak. When it's China Telecom, it's a route hijack. The description of it as a leak AFAICT seems to be due to CF getting first dibs on the announcement[†] and positioned it as such. However, I firmly believe that had the general tech press gotten ahead of it first, it still would be treated much more generously than we treat China leaks. [†] grin
- jgrahamc 7y agoWe've written this incident up: https://blog.cloudflare.com/how-verizon-and-a-bgp-optimizer-knocked-large-parts-of-the-internet-offline-today/ https://blog.cloudflare.com/how-verizon-and-a-bgp-optimizer-...
- plasma 7y agoThank you!
- btown 7y agoGreat article! A couple missing periods at the ends of paragraphs FYI. I'm curious why so much of this lies on Verizon's shoulders. Couldn't DQE and Allegheny have implemented the exact same best practices that Verizon should have, so it never leaked to Verizon's level? And to the extent non-Verizon subscribers were affected, couldn't their ISPs have implemented the same best practices in distrusting Verizon? Is Verizon directly responsible for routing that much of global traffic?
- foota 7y agoI'm not very knowledgeable on network routing, so be warned. But I think at some point a network peering with verizon trusts it to route things, i.e., if I as an ISP always go through verizon to deliver traffic to cloud flare then it's out of my hands the route they take. As for downstreams adding mitigation, ideally this would happen, but I would think you should place blame proportionally to the resources and criticality. A ten person ISP won't necessarily do everything right, and it shouldn't matter that they do, since there's a small part of the internet.
- deleted 7y ago[deleted]
- dang 7y agohttps://news.ycombinator.com/item?id=20267790 https://news.ycombinator.com/item?id=20267790 is a more recent thread on this.