18 ms·
Cloudflare 1.1.1.1 Incident on July 14, 2025
- Mindless2112 1y agoInteresting that traffic didn't return to completely normal levels after the incident. I recently started using the "luci-app-https-dns-proxy" package on OpenWrt, which is preconfigured to use both Cloudflare and Google DNS, and since DoH was mostly unaffected, I didn't notice an outage. (Though if DoH had been affected, it presumably would have failed over to Google DNS anyway.)
- anon7000 1y agoThey go into that more towards the end, sounds like some smaller % of servers needed more direct intervention
- caconym_ 1y ago> Interesting that traffic didn't return to completely normal levels after the incident. Anecdotally, I figured out their DNS was broken before it hit their status page and switched my upstream DNS over to Google. Haven't gotten around to switching back yet.
- radicaldreamer 1y agoWhat would be a good reason to switch back from Google DNS?
- sammy2255 1y agoDepends who you trust more with your DNS traffic. I know who I trust more.
- nojs 1y agoWho? Honest question
- misiek08 1y agoGoogle is serving you ads, CF isn’t. And it’s not conspiracy theory - it was very suspicious when we did some testing on small, aware group. The traffic didn’t look like being handled anonymously at Google side
- mnordhoff 1y agoUnless the privacy policy changed recently, Google shouldn't be doing anything nefarious with 8.8.8.8 DNS queries.
- DarkCrusader2 1y agoThey weren't supposed to do anything with our gmail data as well. That didn't stop them.
- Tijdreiziger 1y ago[citation needed]
- johnklos 1y agoRead their TOS.
- Tijdreiziger 1y agoIf it’s in the ToS, then it’s not true that “[they] weren't supposed to do anything with our gmail data”.
- daneel_w 1y agoYeah it's not like they have a long track record of being caught red-handed stepping all over privacy regulations and snarfing up user activity data across their entire range of free products...
- Algent 1y agoAfter trying both several time I since stayed with google due to cloudflare always returning really bad IPs for anything involving CDN. Having users complain stuff take age to load because you got matched to an IP on opposite side of planet is a bit problematic especially when it rarely happen on other dns providers. Maybe there is a way to fix this but I admit I went for the easier option of going back to good old 8.8.8.8
- homebrewer 1y agoNo, it's deliberately not implemented: https://developers.cloudflare.com/1.1.1.1/faq/#does-1111-send-edns-client-subnet-header https://developers.cloudflare.com/1.1.1.1/faq/#does-1111-sen... I've also changed to 9.9.9.9 and 8.8.8.8 after using 1.1.1.1 for several years because connectivity here is not very good, and being connected to the wrong data center means RTT in excess of 300 ms. Makes the web very sluggish.
- Aachen 1y agoDoes that setup fall back to 8.8.8.8 if 9.9.9.9 fails to resolve? Quad9 has a very aggressive blocking policy (my site with user-uploaded content was banned without even reporting the malicious content; if you're a big brand name it seems to be fine to have user-uploaded content though) which this would be a possible workaround for, but it may not take an nxdomain response as a resolver failure
- motorest 1y ago> Interesting that traffic didn't return to completely normal levels after the incident. Clients cache DNS resolutions to avoid having to do that request each time they send a request. It's plausible that some clients held on to their cache for a significant period.
- bastawhiz 1y agoIf your Internet doesn't work, you'll get up and do other things for a while. I strongly suspect most folks didn't switch DNS providers in that time.
- CuteDepravity 1y agoIt's crazy that both 1.1.1.1 and 1.0.0.1 where affected by the same change I guess now we should start using a completely different provider as dns backup Maybe 8.8.8.8 or 9.9.9.9
- sammy2255 1y ago1.1.1.1 and 1.0.0.1 are served by the same service. It's not advertised as a redundant fully separate backup or anything like that...
- yjftsjthsd-h 1y agoWait, then why does 1.0.0.1 exist? I'll grant I've never seen it advertised/documented as a backup, but I just assumed it must be because why else would you have two? (Given that 1.1.1.1 already isn't actually a single point, so I wouldn't think you need a second IP for load balancing reasons.)
- ta1243 1y agoFar quicker to type ping 1.1 than ping 1.1.1.1 1.0.0.0/24 is a different network than 1.1.1.0/24 too, so can be hosted elsewhere. Indeed right now 1.1.1.1 from my laptop goes via 141.101.71.63 and 1.0.0.1 via 141.101.71.121, which are both hosts on the same LINX/LON1 peer but presumably from different routers, so there is some resilience there. Given DNS is about the easiest thing to avoid a single point of failure on I'm not sure why you would put all your eggs in a single company, but that seems to be the modern internet - centralisation over resilience because resilience is somehow deemed to be hard.
- yjftsjthsd-h 1y ago> Far quicker to type ping 1.1 than ping 1.1.1.1 I guess. I wouldn't have thought it worthwhile for 4 chars, but yes. > 1.0.0.0/24 is a different network than 1.1.1.0/24 too, so can be hosted elsewhere. I thought anycast gave them that on a single IP, though perhaps this is even more resilient?
- geoffpado 1y agoThis was quite annoying for me, having only switched my DNS server to 1.1.1.1 approximately 3 weeks ago to get around my ISP having a DNS outage. Is reasonably stable DNS really so much to ask for these days?
- bauruine 1y agoWhy not use multiple? You can use 1.1.1.1, your ISPs and google at the same time. Or just run a resolver yourself.
- ripdog 1y ago>Or just run a resolver yourself. I did this for a while, but ~300ms hangs on every DNS resolution sure do get old fast.
- xpe 1y agoOuch. What resolver? What hardware? With something like a N100- or N150-based single board computer (perhaps around $200) running any number of open source DNS resolvers, I would expect you can average around 30 ms for cold lookups and <1 ms for cache hits.
- ripdog 1y agoNot a hardware issue, but a physics problem. I live in NZ. I guess the root servers are all in the US, so that's 130ms per trip minimum.
- johnklos 1y agoThey are not all in the US.
- ripdog 1y agoWell that's the experience I had. Obviously caching was enabled (unbound), but most DNS keepalive times are so short as to be fairly useless for a single user. Even if a root server wasn't in the US, it will still be pretty slow for me. Europe is far worse. Most of Asia has bad paths to me, except for Japan and Singapore which are marginally better than the US. Maybe Aus has one...?
- jallmann 1y agoGood writeup. > It’s worth noting that DoH (DNS-over-HTTPS) traffic remained relatively stable as most DoH users use the domain cloudflare-dns.com, configured manually or through their browser, to access the public DNS resolver, rather than by IP address. Interesting, I was affected by this yesterday. My router (supposedly) had Cloudflare DoH enabled but nothing would resolve. Changing the DNS server to 8.8.8.8 fixed the issues.
- bauruine 1y agoHow does DoH work? Somehow you need to know the IP of cloudflare-dns.com first. Maybe your router uses 1.1.1.1 for this.
- nelox 1y ago[flagged]
- k1t 1y agoSmells like AI and completely fails to answer the question. How is the IP address of the DoH server obtained?
- MayeulC 1y agoFirefox accepts a bootstrap IP, or uses the system resolver: > network.trr.bootstrapAddress > (default: none) by setting this field to the IP address of the host name used in "network.trr.uri", you can bypass using the system native resolver for it. Use this to get the IPs of the cloudflare server: https://dns.google/query?name=mozilla.cloudflare-dns.com https://dns.google/query?name=mozilla.cloudflare-dns.com > Starting with Firefox 74 setting the bootstrap address is no longer required in mode 3. Firefox will attempt to use regular DNS in order to get the IP address of the trusted resolver. However, if DNS resolution of the resolver domain fails, setting the bootstrap address is again necessary. Source: https://wiki.mozilla.org/Trusted_Recursive_Resolver https://wiki.mozilla.org/Trusted_Recursive_Resolver
- 1y ago
- angst 1y agoI wonder how uptime ratio of 1.1.1.1 is against 8.8.8.8 Maybe there is noticeable difference? I have seen more outage incident reports of cloudflare than of google, but this is just personal anecdote.
- ta1243 1y agoI guess it depends on where you are and what you count as an outage. Is a single failed query an outage? For me cloudflare 1.1.1.1 and 1.0.0.1 have a mean response time of 15.5ms over the last 3 months, 8.8.8.8 and 8.8.4.4 are 15.0ms, and 9.9.9.9 is 13.8ms. All of those servers return over 3-nines of uptime when quantised in the "worst result in a given 1 minute bucket" from my monitoring points, which seem fine to have in your mix of upstream providers. Personally I'd never rely on a single provider. Google gets 4 nines, but that's only over 90 days so I wouldn't draw any long term conclusions.
- Pharaoh2 1y agohttps://www.dnsperf.com/#!dns-resolvers https://www.dnsperf.com/#!dns-resolvers Last 30 days, 8.8.8.8 has 99.99% uptime vs 1.1.1.1 has 99.09%
- dawnerd 1y agoOh this explains a lot. I kept having random connection issues and when I disabled AdGuard dns (self hosted) it started working so I just assumed it was something with my vm.
- thunderbong 1y agoHow does Cloudflare compare with OpenDNS?
- blurrybird 1y agoYou’d be better off comparing it to Quad9 based on performance, privacy claims, and response accuracy.
- johnklos 1y agoCloudflare is a for-profit company in the US. Their privacy claims can't be believed. Even if we did believe them, we have no idea if rsolution data isn't taken by US TLA agencies.
- forbiddenlake 1y agoHm, what distinction are you trying to make here? OpenDNS is also an American company, acquired by Cisco (an American company) in 2015.
- johnklos 1y agoI don't know much about OpenDNS, but yes, I wouldn't trust Cisco to do anything that didn't somehow push money in their direction. I was just offering relevant information about Cloudflare.
- johnklos 1y agoIt seems we have a lot of Cloudflare fanbois and apologists here. This is not unexpected. But is anything I'm writing untrue, or just unpopular? Does anyone who's downvoting me care to point out any inaccuracies about what I've written?
- astrange 1y agoIt's illegal for a US company to lie to their investors, so if you believe they're lying to you, you should sue them for securities fraud.
- nodesocket 1y agoI used to configure 1.1.1.1 as primary and 8.8.8.8 as secondary but noticed that Cloudflare on aggregate was quicker to respond to queries and changed everything to use 1.1.1.1 and 1.0.0.1. Perhaps I'll switch back to using 8.8.8.8 as secondary, though my understanding is DNS will round-robin between primary and secondary, it's not primary and then use secondary ONLY if primary is down. Perhaps I am wrong though. EDIT: Appears I was wrong, it is failover not round-robin between the primary and secondary DNS servers. Thus, using 1.1.1.1 and 8.8.8.8 makes sense.
- ta1243 1y agoDepends on how you configure it. In resolv.conf systems for example you can set a timeout of say 1 second and do it as main/reserve, or set it up to round-robin. From memory it's something like "options:rotate" If you have a more advanced local resolver of some sort (systemd for example) you can configure whatever behaviour you want.
- chrismorgan 1y agoI’m surprised at the delay in impact detection: it took their internal health service more than five minutes to notice (or at least alert) that their main protocol’s traffic had abruptly dropped to around 10% of expected and was staying there. Without ever having been involved in monitoring at that kind of scale, I’d have pictured alarms firing for something that extreme within a minute. I’m curious for description of how and why that might be, and whether it’s reasonable or surprising to professionals in that space too.
- TheDong 1y agoI'm not surprised. Let's say you've got a metric aggregation service, and that service crashes. What does that result in? Metrics get delayed until your orchestration system redeploys that service elsewhere, which looks like a 100% drop in metrics. Most orchestration take a sec to redeploy in this case, assuming that it could be a temporary outage of the node (like a network blip of some sort). Sooo, if you alert after just a minute, you end up with people getting woken up at 2am for nothing. What happens if you keep waking up people at 2am for something that auto-resolves in 5 minutes? People quit, or eventually adjust the alert to 5 minutes. I know you often can differentiate no data and real drops, but the overall point, of "if you page people constantly, people will quit" I think is the important one. If people keep getting paged for too tight alarms, the alarms can and should be loosened... and that's one way you end up at 5 minutes.
- mentalgear 1y agoIts not wrong for smaller companies. But there's an argument that a big system critical company/provider like Cloudflare should be able to afford its own always on team with a night shift.
- chrismorgan 1y agoNot even a night shift, just normal working hours in another part of the world.
- 1y ago
- egamirorrim 1y agoWhat's that about a hijack?
- homero 1y agoRelated, non-causal event: BGP origin hijack of 1.1.1.0/24 exposed by withdrawal of routes from Cloudflare. This was not a cause of the service failure, but an unrelated issue that was suddenly visible as that prefix was withdrawn by Cloudflare.
- kylestanfield 1y agoSo someone just started advertising the prefix when it was up for grabs? That’s pretty funny
- woutifier 1y agoNo they were already doing that, the global withdrawal of the legitimate route just exposed it.
- SemioticStandrd 1y agoHow is there absolutely no further comment about that in their RCA? That seems like a pretty major thing...
- JdeBP 1y agoAnd because people highlighted it on social media at the time of the outage, many thought that the bogus route was the cause of the problem.
- ollien 1y agoI'm a bit uneducated here - why was the other 1.1.1.0/24 announcement previously suppressed? Did it just express a high enough cost that no one took it on compared to the CF announcement?
- 1y ago
- 0xbadcafebee 1y ago> A configuration change was made for the same DLS service. The change attached a test location to the non-production service; this location itself was not live, but the change triggered a refresh of network configuration globally. Say what now? A test triggered a global production change? > Due to the earlier configuration error linking the 1.1.1.1 Resolver's IP addresses to our non-production service, those 1.1.1.1 IPs were inadvertently included when we changed how the non-production service was set up. You have a process that allows some other service to just hoover up address routes already in use in production by a different service?
- sneak 1y ago1.1.1.1 does not operate in isolation. It is designed to be used in conjunction with 1.0.0.1. DNS has fault tolerance built in. Did 1.0.0.1 go down too? If so, why were they on the same infrastructure? This makes no sense to me. 8.8.8.8 also has 8.8.4.4. The whole point is that it can go down at any time and everything keeps working. Shouldn’t the fix be to ensure that these are served out of completely independent silos and update all docs to make sure anyone using 1.1.1.1 also has 1.0.0.1 configured as a backup? If I ran a service like this I would regularly do blackouts or brownouts on the primary to make sure that people’s resolvers are configured correctly. Nobody should be using a single IP as a point of failure for their internet access/browsing.
- detaro 1y agoYou don't need to test if peoples resolvers handle this cleanly, because its already known that many don't. DNS fallback behavior across platforms is a mess.
- notpushkin 1y ago> Did 1.0.0.1 go down too? Yes. > Shouldn’t the fix be to ensure that these are served out of completely independent silos [...]? Yes. > If so, why were they on the same infrastructure? Apparently, they weren’t independent enough: something in CF has announced both addresses and that got out. The solution for the end user is, of course, to use 1.1.1.1 and 8.8.8.8 (or any other combination of two different resolvers).
- rswail 1y agoI now run unbound locally as a recursive DNS server, which really should be the default. There's no reason not to in modern routers. Not sure what the "advantage" of stub resolvers is in 2025 for anything.
- i_niks_86 1y agoMany commenters assume fallback behavior exists between DNS providers, but in practice, DNS clients - especially at the OS or router level -rarely implement robust failover for DoH. If you're using cloudflare-dns(.)com and it goes down, unless the stub resolver or router explicitly supports multi-provider failover (and uses a trust-on-first-use or pinned cert model), you’re stuck. The illusion of redundancy with DoH needs serious UX rethinking.
- tankenmate 1y agoI use routedns[0] for this specific reason it handles almost all DNS protocols; UDP, TCP, DoT, DoH, DoQ (including 0-RTT). But more importantly is has a very configurable route steering even down to a record by record basis if you want to put up with all the configuration involved. It's very robust and is very handy, I use 1.1.1.1 on my desktops and servers and when the incident happened I didn't even notice as the failover "just worked". I had to actually go look at the logs because I didn't notice. [0] https://github.com/folbricht/routedns https://github.com/folbricht/routedns
- deleted 1y ago[deleted]
- hkon 1y agoTo say I was surprised when I finally checked the status page of cloudflare is an understatement.
- v5v3 1y ago> For many users, not being able to resolve names using the 1.1.1.1 Resolver meant that basically all Internet services were unavailable. Don't you normally have 2 DnS servers listed on any device. So was the second also down, if not why didn't it go to that.
- rat9988 1y agoNot all users have configured two DNS servers?
- quacksilver 1y agoIt is highly recommended to configure two or more DNS servers incase one is down. I would count not configuring at least two as 'user error'. Many systems require you to enter a primary and alternate server in order to save a configuration.
- tgv 1y agoThe default setting on most computers seems to be: use the (wifi) router. I suppose telcos like that because it keeps the number of DNS requests down. So I wouldn't necessarily see it as user error.
- SketchySeaBeast 1y agoThe funny part with that is that sites like cloudflare say "Oh, yeah, just use 1.0.0.1 as your alternate", when, in reality, it should be an entirely different service.
- daneel_w 1y agoOK. But there's no reason or excuse not to, if they already manually configured a primary.
- rom1v 1y agoOn Android, in Settings, Network & internet, Private DNS, you can only provide one in "Private DNS provider hostname" (AFAIK). Btw, I really don't understand why it does not accept an IP (1.1.1.1), so you have to give an address (one.one.one.one). It would be more sensible to configure a DNS server from an IP rather than from an address to be resolved by a DNS server :/
- udev4096 1y agoThis is why running your own resolver is so important. Clownflare will always break something or backdoor something
- perlgeek 1y agoAn outage of roughly 1 hour is 0.13% of a month or 0.0114% of a year. It would be interesting to see the service level objective (SLO) that cloudflare internally has for this service. I've found https://www.cloudflare.com/r2-service-level-agreement/ https://www.cloudflare.com/r2-service-level-agreement/ but this seems to be for payed services, so this outage would put July in the "< 99.9% but >= 99.0%" bucket, so you'd get a 10% refund for the month if you payed for it.
- philipwhiuk 1y agoProbably 99.9% or better annually just from a 'maintaining reputation for reliability' standpoint.
- stingraycharles 1y agoWhat really matters with these percentages is whether it’s per month or per year. 99.9% per year allows for much longer outages than 99.9% per month.
- kachapopopow 1y agoInteresting to see that they probably lost 20% of 1.1.1.1 usage from a roughly 20 minute incident. Not sure how cloudflare keeps struggling with issues like these, this isn't the first (and probably won't be the last) time they have these 'simple', 'deprecated', 'legacy' issues occuring. 8.8.8.8+8.8.4.4 hasn't had a global(1) second of downtime for almost a decade. 1: localized issues did exist, but that's really the fault of the internet and they did remain running when google itself suffered severe downtime in various different services.
- Tepix 1y agoThere's more to DNS than just availability (granted, it's very important). There's also speed and privacy. European users might prefer one of the alternatives listed at https://european-alternatives.eu/category/public-dns https://european-alternatives.eu/category/public-dns over US corporations subject to the CLOUD act.
- immibis 1y agoHN users might prefer to run their own. It's a low maintenance service. It's not like running a mail server.
- daneel_w 1y agoI think that might be overestimating the technical prowess of HN readers on the whole. Sure, it doesn't require wizardry to set up e.g. Unbound as a catch-all DoT forwarder, but it's not the click'n'play most people require. It should be compared to just changing the system resolvers to dns0, Quad9 etc.
- lossolo 1y agoOne issue here is that you can be tracked easily.
- kachapopopow 1y agoRunning your own and being the sole user is the exact same thing as using a dns server (you need to obtain nameservers for any given domain which you have to contact a dns server for).
- nness 1y agoInteresting side-effect, the Gluetun docker image uses 1.1.1.1 for DNS resolution — as a result of the outage Gluetun's health checks failed and the images stopped. If there were some way to view torrenting traffic, no doubt there'd be a 20 minute slump.
- johnklos 1y agoPersonally, I'd consider any Docker image that does its own DNS resolution outside of the OS a Trojan.
- greggsy 1y agoI’d love to know legacy systems they’re referring to.
- wreckage645 1y agoThis is a good post mortem, but improvements only come with change on processes. It seems every team at CloudFlare is approaching this in isolation, without a central problem management. Every week we see a new CloudFlare global outage. It seems like the change management processes is broken and needs to be looked at..
- sylware 1y agocloudflare is providing a service designed to block noscript/basic (x)html browsers. I know.
- trollbridge 1y agoI got bit by this, so dnsmasq now has 1.1.1.2, Quad9, and Google’s 8.8.8.8 with both primary and secondary. Secondary DNS is supposed to be in an independent network to avoid precisely this.
- neurostimulant 1y agoI never noticed the outage because my isp hijack all outbound udp traffic to port 53 and redirect them to their own dns server so they can apply government-mandated cencorship :)
- nu11ptr 1y agoQuestion: Years ago, back when I used to do networking, Cisco Wireless controllers used 1.1.1.1 internally. They seemed to literally blackhole any comms to that IP in my testing. I assume they changed this when 1.0.0.0/8 started routing on the Internet?
- blurrybird 1y agoYeah part of the reason why APNIC granted Cloudflare access to those very lucrative IPs is to observe the misconfiguration volume. The theory is CF had the capacity to soak up the junk traffic without negatively impacting their network.
- yabones 1y agoThe general guidance for networking has been to only use IPs and domains that you actually control... But even 5-8 years ago, the last time I personally touched a cisco WLC box, it still had 1.1.1.1 hardcoded. Cisco loves to break their own rules...
- homebrewer 1y agoThis is a good time to mention that dnsmasq lets you setup several DNS servers, and can race them. The first responder wins. You won't ever notice one of the services being down: all-servers server=8.8.8.8 server=9.9.9.9 server=1.1.1.1
- mnordhoff 1y agoEven without "all-servers", DNSMasq will race servers frequently (after 20 seconds, unless it's changed), and when retrying. A sudden outage should only affect you for a few seconds, if at all.
- anthonyryan1 1y agoAdditionally, as long as you don't set strict-order, dnsmasq will automatically use all-servers for retries. If you were using systemd-resolved however, it retries all servers in the order they were specified, so it's important to interleave upstreams. Using the servers in the above example, and assuming IPv4 + IPv6: 1.1.1.1 2001:4860:4860::8888 9.9.9.9 2606:4700:4700::1111 8.8.8.8 2620:fe::fe 1.0.0.1 2001:4860:4860::8844 149.112.112.112 2606:4700:4700::1001 8.8.4.4 2620:fe::9 will failover faster and more successfully on systemd-resolved, than if you specify all Cloudflare IPs together, then all Google IPs, etc. Also note that Quad9 is default filtering on this IP while the other two or not, so you could get intermittent differences in resolution behavior. If this is a problem, don't mix filtered and unfiltered resolvers. You definitely shouldn't mix DNSSEC validatng and not DNSSEC validating resolvers if you care about that (all of the above are DNSSEC validating).
- matthewtse 1y agowow good tip I was handling an incident due to this outage. I ended up adding Google DNS resolvers using systemd-resolved, but I didn't think to interleave them!
- karel-3d 1y agodnsdist is AMAZINGLY easy to set up as a secure local resolver that forwards all queries to DoH (and checks SSL) and checks liveliness every second I need to do a write-up one day
- alyandon 1y agoCloudflare's 1.1.1.1 Resolver service became unavailable to the Internet starting at 21:52 UTC and ending at 22:54 UTC Weird. According to my own telemetry from multiple networks they were unavailable for a lot longer than that.
- chrisgeleven 1y agoI been lazy and was using Cloudflare's resolver only recently. In hindsight I probably should just setup two instances of Unbound on my home network that don't rely on upstream resolvers and call it a day. It's unlikely both will go down at the same time and if I'm having an total Internet outage (unlikely as I have Comcast as primary + T-Mobile Home Internet as a backup), it doesn't matter if DNS is or isn't resolving.
- tacitusarc 1y agoPerhaps I am over-saturated, but this write up felt like AI- at least largely edited by a model.
- xyst 1y agoAm not a fan of CF in general due to their role in centralization of the internet around their services. But I do appreciate these types of detailed public incident reports and RCAs.
- zac23or 1y agoIt's no surprise that Cloudflare is having a service issue again. I use Cloudflare at work. Cloudflare has many bugs, and some technical decisions are absurd, such as the worker's cache.delete method, which only clears the cache contents in the data center where the Worker was invoked!!! https://developers.cloudflare.com/workers/runtime-apis/cache/#delete https://developers.cloudflare.com/workers/runtime-apis/cache... In my experience, Cloudflare support is not helpful at all, trying to pass the problem onto the user, like "Just avoid holding it in that way. ". At work, I needed to use Cloudflare. The next job I get, I'll put a limit on my responsibilities: I don't work with Cloudflare. I will never use Cloudflare at home and I don't recommend it to anyone. Next week: A new post about how Cloudflare saved the web from a massive DDOS attack.
- kentonv 1y ago> some technical decisions are absurd, such as the worker's cache.delete method, which only clears the cache contents in the data center where the Worker was invoked!!! The Cache API is a standard taken from browsers. In the browser, cache.delete obviously only deletes that browser's cache, not all other browsers in the world. You could certainly argue that a global purge would be more useful in Workers, but it would be inconsistent with the standard API behavior, and also would be extraordinarily expensive. Code designed to use the standard cache API would end up being much more expensive than expected. With all that said, we (Workers team) do generally feel in retrospect that the Cache API was not a good fit for our platform. We really wanted to follow standards, but this standard in this case is too specific to browsers and as a result does not work well for typical use cases in Cloudflare Workers. We'd like to replace it with something better.
- freedomben 1y agoJust wanted to say, I always appreciate your comments and frankness!
- zac23or 1y ago>cache.delete obviously only deletes that browser's cache, not all other browsers in the world. To me, it only makes sense if the put method creates a cache only in the datacenter where the Worker was invoked. Put and delete need to be related, in my opinion. Now I'm curious: what's the point of clearing the cache contents in the datacenter where the Worker was invoked? I can't think of any use for this method. My criticisms aren't about functionality per see or developers. I don't doubt the developers' competence, but I feel like there's something wrong with the company culture.
- aftbit 1y ago>Even though this release was peer-reviewed by multiple engineers I find it somewhat surprising that none of the multiple engineers who reviewed the original change in June noticed that they had added 1.1.1.0/24 to the list of prefixes that should be rerouted. I wonder what sort of human mistake or malice led to that original error. Perhaps it would be wise to add some hard-coded special-case mitigations to DLS such that it would not allow 1.1.1.1/32 or 1.0.0.1/32 to be reassigned to a single location.
- burnte 1y agoIt's probably much simpler, "I trust Jerry, I'm sure this is fine, approved."
- roughly 1y agoI’m generally more a “blame the tools” than “blame the people” - depending on how the system is set up and how the configs are generated, it’s easy for a change like this to slip by - especially if a bunch of the diff is autogenerated. It’s still humans doing code review, and this kind of failure indicates process problems, regardless of whether or not laziness or stupidity were also present. But, yes, a second mitigation here would be defense in depth - in an ideal world, all your systems use the same ops/deploy/etc stack, in this one, you probably want an extra couple steps in the way of potentially taking a large public service offline.
- deleted 1y ago[deleted]
- b0rbb 1y agoI don't know about you all but I love a well written RCA. Nicely done.
- alexandrutocar 1y ago> It’s worth noting that DoH (DNS-over-HTTPS) traffic remained relatively stable as most DoH users use the domain cloudflare-dns.com, configured manually or through their browser, to access the public DNS resolver, rather than by IP address. I use their DNS over HTTPS and if I hadn't seen the issue being reported here, I wouldn't have caught it at all. However, this—along with a chain of past incidents (including a recent cascading service failure caused by a third-party outage)—led me to reduce my dependencies. I no longer use Cloudflare Tunnels or Cloudflare Access, replacing them with WireGuard and mTLS certificates. I still use their compute and storage, but for personal projects only.
- cadamsdotcom 1y ago> The way that Cloudflare manages service topologies has been refined over time and currently consist of a combination of a legacy and a strategic system that are synced. This writing is just brilliant. Clear to technical and non-technical readers. Makes the in-progress migration sound way more exciting than it probably is! > We are sorry for the disruption this incident caused for our customers. We are actively making these improvements to ensure improved stability moving forward and to prevent this problem from happening again. This is about as good as you can get it from a company as serious and important as Cloudflare. Bravo to the writers and vetters for not watering this down.
- kccqzy 1y agoI can't tell if you are being sarcastic, but "legacy" is a term most often used by technical people whereas "strategic" is a term most often used by marketing and non-technical leadership. Mixing them together annoys both kinds of readers.
- nixpulvis 1y agoFun fact, Verizon cellular blocks 1.1.1.1. I discovered this after trying to use my hotspot from my Linux laptop with it set for my default DNS. Very frustrating.