37 ms·
Akamai Edge DNS was down
- cbono1 5y agoWhy would Google and Amazon be on the downdetector list or experiencing issues? Don't they have their own DNS / nameservers separate from Akamai?
- sathackr 5y agobecause the way downdetector works is it just basically counts how many people are searching/visiting for <site> down and if it's much higher than typical it flags the site as down. So if everyone searched "is google down" and visited the link on downdetector that was returned in the search, that would add to the downdetector count for that site. Downdetector doesn't actually know if the site is up or down.
- brentm 5y agoA more proper name might be PeopleThinkItsDownDetector.com
- cbono1 5y agoNot nearly as SEO friendly
- k1t 5y agoI found this hard to believe, but it's correct. Downdetector only reports an issue if a significant number of users are impacted. To that end, Downdetector calculates a baseline volume of typical problem reports for each service monitored, based on the average number of reports for that given time of day over the last year. Downdetector’s incident detection system compares the current number of problem reports to this baseline and only reports an issue if the current volume significantly exceeds the typical volume of reports. https://www.speedtest.net/insights/blog/how-downdetector-works/ https://www.speedtest.net/insights/blog/how-downdetector-wor...
- topranks 5y agoWhat’s hard to believe? Downdetectors well known for being almost, but not quite, useless. Probably reported Google as “down” because a whole bunch of people use the word “Google” when they mean “internet”.
- mc32 5y agoSo how do they reset status? The number of queries going down signifies return to normal status?
- thunfisch 5y agoYep, all our EdgeDNS zones as well as DSD edgekeys are just returning SERVFAILS. Many big german websites are down right now.
- zhdc1 5y agoSeveral unrelated websites I was trying to visit are down. I figured I would find the answer on HN : )
- mariusseufzer 5y agoSame haha
- realSaddy 5y agoThis is affecting Steam as well
- ssully 5y agoIt is impacting a lot of things: https://downdetector.com/ https://downdetector.com/
- Scoundreller 5y agoAll yuor data are belong to us
- Eikon 5y agoThis is affecting apple as well https://www.apple.com/go/ https://www.apple.com/go/
- iruoy 5y agoFor some reason that url doesn't work for me, but https://www.apple.com/ https://www.apple.com/ and https://www.apple.com/nl/ https://www.apple.com/nl/ do.
- remram 5y agoThat fails with a 404 for me, which is probably not related to DNS at all? archive.org seems to indicate there was never anything there...
- tyingq 5y agoYou can see this on a lot of sites right now. You get the Akamai style error with something like: Reference: #11.453a2f17.1393u44848484.3aee33433 At the bottom of a very bland looking error page.
- lowbloodsugar 5y agoWhat's frustrating is that DNS is returning an address, instead of just failing, and so macos is caching that value (though it might be cloudflare doing that).
- space_ghost 5y agoWildcard DNS should be a prosecutable crime, punishable by no less than 20 years of hard labor. (Edit: Probably should have made it clear that this was a joke)
- breakingcups 5y agoI don't see how wildcard DNS is related to this? Nor how it's bad?
- gokhan 5y agoWildcard DNS helps me to handle multitenancy easily. What's wrong with it?
- dylan604 5y agoWhen did congress members start posting to HN?
- adamdoran 5y agoPresumably you're referring to the practice of answering queries for nonexistent records with an A record belonging to an advertisement page? (instead of doing the right thing answering NXDOMAIN, presuming no records of another type also exist for the queried name.) dnsmasq has a really useful feature for dealing with this: --bogus-nxdomain
- 5y ago
- dbsmith83 5y agohttps://downdetector.com/archive https://downdetector.com/archive So many sites down... and unfortunately not one of them is Twitter
- cpgeier 5y agoAmazing that down detector manages to stay up during these kinds of outages. Noticed it has been a little slow but they really have done a good job keeping it up even though large portions of the internet is down right now.
- mindcrime 5y agoWho detects if Down Detector is down? Is there a isdowndetectordown.com site?
- mcintyre1994 5y agoIt's parked by GoDaddy, but unfortunately their website is fubar by this outage if you try to click through to see how much they want for it :)
- ksec 5y agoI guess the mother of all Network Downtime checker is HN.
- cube00 5y ago"I dunno. Coast Guard?"
- SahAssar 5y agoSounds like when Fuckedcompany put itself on Fuckedcompany.
- mcintyre1994 5y agoIt's interesting that they report an AWS outage but there don't seem to be any issues there. Looks like their methodology is a bit too reliant on those speculative tweets from the first 5 minutes of all these sites going down. https://downdetector.com/status/aws-amazon-web-services/ https://downdetector.com/status/aws-amazon-web-services/ > So many websites are down, are AWS servers down or something? > Amazon web services is down which is affecting a lot of company web sites and services. Not sure what is going on. > Miss us? @aldotcom and a whole bunch of other folks have been knocked off the internet by what appears to be an AWS attack/system failure. We'll be back. ?
- 00deadbeef 5y agoFigured this out almost 30 minutes before they bothered to update their status page.
- rvz 5y agoProbably Akamai needs to use Kubernetes. EDIT: So HN can't even take a joke after this? [0] [0] https://news.ycombinator.com/item?id=27893482 https://news.ycombinator.com/item?id=27893482
- whitepoplar 5y agoProbably caused by Kubernetes
- rvz 5y agoThat's even worse if true; despite HNers creating a storm in a tea cup on DOSing a blog of a service not using K8s when having a blog is not their main service. [0]. Either way, the joke's is now on the HNers in that thread. [0] https://news.ycombinator.com/item?id=27893482 https://news.ycombinator.com/item?id=27893482
- mdtancsa 5y agoSheesh, So yesterday! :)
- unemphysbro 5y agocome on, this is funny. HN needs to lighten-up.
- deleted 5y ago[deleted]
- simonswords82 5y agoI'm sick and tired of these types of services (I'm looking at you too Cloudflare) going down and taking otherwise healthy websites down with them.
- sammy2244 5y agoCloudflare hasnt had an outage in a long time. And when they do they are upfront about it, and post a detailed post-mortem.
- ceejayoz 5y agoMost websites using Akamai aren't gonna be "otherwise healthy" without the CDN handling most of the load.
- tootie 5y agoIt was fastly last time.
- simonswords82 5y agoTrue but cloudflare have been guilty of downtime too.
- ceejayoz 5y agoThere aren't many sites that aren't, including "otherwise healthy websites" hosted without a CDN.
- TheSwordsman 5y agoI think this is a factually true statement if your business uses any computers. ;)
- davidjgraph 5y agoSerious question, has anyone properly solved the issue of DNS as a single point of failure?
- tyingq 5y agoIt's an interesting question, as it's always been solved on the server side. All of the current problem is client side. That is, client resolvers that aren't using diverse providers, and only do things like round-robin with long timeouts.
- kokey 5y agoAnycast for the DNS IPs deals with most of the problems of clients not failing over elegantly when their primary DNS server is broken.
- citrin_ru 5y agoFrom a client (DNS recursor) point of view there is no primary server. There is just multiple NS records which are equal. If one of them is down it can introduce resolving delays, but they are usually small. At least if something like Unbound or Bind is used. Unbound e. g. maintains infra-cache where it tracks RTT and errors for each server and avoid servers which are down.
- sakisv 5y agoDepending on what point you draw the line of "single point of failure" you could use multiple providers for your dns. GOV.UK for example uses both aws and gcp for DNS
- davidjgraph 5y agoSo, NS entries pointing to both? But then take the example your domain was in Route53 and AWS goes down. You can't configure the NS entries to avoid AWS DNS servers. Is the idea that child DNS servers detect the outage and cache the values in the name server(s) that remain up? But then, the cached values from AWS take a while to clear, TTL never seems to be applied properly. It always feels like the worst case in such a scenario is you can point everyone at the right thing within 24 hours.
- tru3_power 5y agoAny idea on cause? Ddos or hardware failure?
- twalichiewicz 5y agoPosted this is the thread about the travel websites being down, but seems Fidelity is entirely impossible to sign in to / trade right now.
- sebyx07 5y agoThe good parts of centralisation
- conqrr 5y agoAffecting Airbnb search
- delgaudm 5y agoLastpass is down, so if you use lastpass the effect is significantly compounded.
- mcintyre1994 5y agoDo they not cache everything locally? I'd have thought a password manager/secure data store would work offline.
- stusmall 5y agoThey do.
- sammy2244 5y agoHaving your passwords only accessible by internet is a stupid idea anyway
- nonfamous 5y agoIt still works in offline mode. You can’t update passwords, but you can retrieve them.
- compscistd 5y agoTo enable offline mode, I had to turn on airplane mode on my phone before logging in.
- memco 5y agoWas just browsing a website where the first page of a query worked, but visiting page 2 of the results was returning a DNS error. Was curious how and why only part of the site was down, but it looks like this was the problem as now the whole site is down.
- katbyte 5y agoaren't short DNS TTLs great?
- sebmellen 5y agoIs this a serious argument for long TTLs? Always wondered why they exist… How interesting.
- slim 5y agoYes it is. The longer the TTL the longer you stay independent from third parties. It's what makes the internet stable.
- remram 5y agoLong TTL makes you independent from DNS third parties, in that your name is still know by clients if DNS is down. Short TTL makes you independent from hosting third parties, in that you can quickly change which hosting provider your domain name points to. You can't win this one by only changing your TTL. The best solution is to use short TTLs and multiple nameservers on different providers.
- gianpaj 5y agohttps://www.interactivebrokers.co.uk/ https://www.interactivebrokers.co.uk/ , a Trading Platform, is also down as well :( How am I going to sell my AMC stock...
- swarnie_ 5y agoYou don't, you hold the dumb, over priced stock as a reminder for future, better informed investing.
- cbeley 5y agoI wonder if this is why LastPass is down. It has completely locked me out of my vault. You'd think it'd continue to work offline in a case like this. :/
- eunai 5y agoI switched to BitWarden and haven't looked back. You can use it on the phone and pc (browser). As well as a desktop client.
- benburleson 5y agoYeah, my path was LastPass -> Bitwarden -> 1Password. Both Bitwarden and 1Password are great.
- JonathanMerklin 5y agoThen what was the impetus to switch off of Bitwarden?
- decrypt 5y agoSame path. It'll be very hard to move away from 1Password. App experience, sync, security features like key in addition to master password, family organizer-based recovery of an account, these are a few things that stand out.
- revscat 5y agoCan you explain what family organizer-based recovery means? It sounds like dad or mom could recover a kids password?
- chewmieser 5y agohttps://support.1password.com/recovery/ https://support.1password.com/recovery/
- 5y ago
- deleted 5y ago[deleted]
- bpye 5y agoThis is apparently why I can't book my COVID vaccine appointment...
- _joel 5y agoYes, was trying to do the same. Getting this 2nd jab has been a nightmare. Places listed as walk-in having Moderna, don't and they ran out of it when I went to get my secheduled jab. Ringing 119 just ends up in a dead line, then this outage. Fun.
- mvanaltvorst 5y agoWhat role does Akamai Edge DNS play in normal internet traffic? DNS responses usually get cached, as far as I understand correctly. And it is usually possible to change your DNS server to e.g. Google's and circumvent the outage. Does Akamai Edge DNS play a role on the server side?
- NeckBeardPrince 5y ago> What role does Akamai Edge DNS play in normal internet traffic? Clearly a big one.
- r1ch 5y agoThe trend these days are DNS TTLs of 60 - 300 seconds, to allow "Cloud agility" or something, so sites are exposed to a much larger risk of authoritative nameservers going down.
- jameshart 5y agoYou say that like it's a bad idea. Services like Akamai use short TTLs for their edge services for a variety of reasons, not least because if one of their edge servers goes offline (for planned or unplanned reasons) it lets them sub in a new one and have it receive traffic immediately, rather than have a bunch of clients continue trying to talk to a dead node. So sure, you can increase those TTLs to trade 'what if the DNS server goes down?' risk with 'what if the edge server goes down?' risk... But keeping the edge servers up and running is probably a lot harder - they need to scale more to handle traffic load, they have to actually handle client data, TLS termination, much more complex configuration.... so if I'm placing bets on which of those things is more likely to die on me, it's the edge node, not the DNS server.
- uncertainrhymes 5y agoIf you use a CDN to front your traffic, you need the CNAME for www (or whatever) to be pointing at their DNS infrastructure, so they can return whichever closest POP is going to serve your traffic. e.g. dig @1.1.1.1 www.nvidia.com +trace ... various things from the root ... www.nvidia.com. 7200 IN CNAME www.nvidia.com.edgekey.net. ;; Received 83 bytes from 208.94.148.13#53(ns5.dnsmadeeasy.com) in 35 ms So the main DNS is fine, but it'll never get an A record because the last link in the chain is toast -- edgekey being Akamai in this case, but all CDNs do this so they can route traffic. Normally, this is a good thing so they can shift traffic within 30 seconds on their side. Unfortunately, it also means it would take nvidia an two hours to point away from Akamai.
- _joel 5y agoSo that's why the NHS website is down
- jamespwilliams 5y agoBack up now by the looks of it
- jdlyga 5y agoOops, someone unplugged the DNS machine
- SjorsVG 5y agoMany bank systems are disrupted by this in the Netherlands
- ricardo81 5y agoMy UK bank (HBOS) seemed to have 'online banking unavailable' though their site was up. No doubt related.
- SjorsVG 5y agoMany banks in the Netherlands are affected by this.
- schemathings 5y agoPossibly related .. Verizon peering issues / ASN701 at Equinix NY2 in Secaucus NJ
- knaik94 5y agoI am surprised financial institutions don't have any regulation for redundancy. The one that stuck out to me is the Navy Federal Credit Union website being down. I have not had any issues logging into mobile though for some of the reported sites.
- christophilus 5y agoI'm not sure how easy it would be to regulate. But yeah. I've got a few short term trades in my brokerage account, and outages really throw a wrench into those.
- xyzzy21 5y agoThe way regulate is like anything else: if they fail to meet QoS uptimes, they get fined in 6-8 figures for every minute of loss.
- brentm 5y agoCapitalOne has a broken login which is pretty surprising to me.
- toomuchtodo 5y agoCommercial banks are held to a different operational resiliency standard than financial infrastructure. (a component of my consulting work is reporting to financial regulators for institutions)
- cryptoz 5y agoAll major Canadian banks were down.
- deckard1 5y agothis is prime shit Hacker News says right here. Wait until you learn banks close on Sunday. Or have maintenance windows for their website, ATM, etc.
- Terretta 5y ago> financial institutions don't have any regulation for redundancy As CTO of a bank, I wasn’t aware of this. So either we wasted a ton of money and time constantly upgrading redundancy and business continuity technologies to satisfy our regulators… or this statement could be mistaken.
- tjpnz 5y agoJust got booted out of Netflix on the PS4 because the console could no longer connect to Sony's license server. Netflix was working just fine by the way.
- hackerbrother 5y agoYup, I learned Hulu on Xbox One relies heavily on some Microsoft authentication during a recent Office 365 or Azure outage (not sure which).
- lxgr 5y agoWas the app installed/running using a secondary PSN account by any chance? This shouldn't be happening on a primary account/console pair.
- tjpnz 5y agoIt should be my primary although I've often seen it revert back after setting it. I did try setting it as my primary again but you know.
- vmception 5y agoAh thats whats going on. Happened to me as well, I just assumed that Sony is neglecting PS4 performance with its new system, while bogging it down with bloatware.
- deleted 5y ago[deleted]
- SandroG 5y agoIs this related to: Multiple websites including DraftKings, Airbnb, FedEx, Delta and others appear to be experiencing issues. https://www.bloomberg.com/news/articles/2021-07-22/multiple-websites-appear-to-be-experiencing-tech-issues?srnd=premium https://www.bloomberg.com/news/articles/2021-07-22/multiple-...
- xyzzy21 5y agoAnd people wonder why I try to avoid depending on online anything...
- soheil 5y agoApp Store on MacOS is down!
- swarnie_ 5y agoI love seeing these issues reverberate around the internet. This time i think /r/sysadmin pegged the issue first, great sub.
- lowbloodsugar 5y agoSo many sites being reported as down, but change your DNS to something else (e.g. Google 8.8.8.8 and 8.8.4.4) and, after flushing your DNS cache, the sites are available. I was unable to get to ups.com or newegg.com (why yes, I am expecting a new toy), but after switching DNS and flushing DNS cache, I was able to get to both. Specifically, 1.1.1.1 provided bad addresses (as opposed to no addresses), and removing 1.1.1.1 fixed my problem. By then it had returned a bunch of bad addresses and I had to flush my DNS cache.
- aix1 5y agoCould you give an example of what you mean by a "bad address" in this context?
- lowbloodsugar 5y agoThis is from the time of incident: Server: 1.1.1.1 Address: 1.1.1.1#53 Non-authoritative answer: Name: newegg.com Address: 23.35.185.6 vs Server: 8.8.8.8 Address: 8.8.8.8#53 Non-authoritative answer: Name: newegg.com Address: 104.80.92.252 104.80.92.252 is newegg.com 23.35.185.6 is a server that provides an error message. So 1.1.1.1 lied. The proper response would be to reply "I don't recognize that domain". Instead it said, "yeah, I know that, its here..." Newegg was not down, and when I got macos to forget what it had cached from 1.1.1.1 I was able to use newegg.com fine.
- blondie9x 5y agoLooks like it is fixed now!
- fredski42 5y agoI thought DNS was supposed to be resilient
- topspin 5y agoDNS is designed to be fault tolerant. Such a design, however, is often not leveraged correctly; the implementation of DNS can be and frequently is subject to SPOFs.
- 00deadbeef 5y agoWell it's been an hour now since I first noticed the effects and their service status still has no useful information or ETA for a fix. It's just an "emerging issue".
- jonnyone 5y agoThe affected sites that I use are now working. Check again.
- mvelie 5y agoAkamai believes they have it fixed. We've seen our traffic return to normal. https://twitter.com/Akamai/status/1418251400660889603 https://twitter.com/Akamai/status/1418251400660889603
- roody15 5y agohmmm does not appear fixed here in the Midwest
- foobarbazetc 5y agoAbsolutely amazing how many billion $+ companies are single homed for DNS. I wonder how much they spend on multi-AZ redundant architectures...
- orblivion 5y agoSo here's a weird question: Supposing companies multi-home for DNS, or whatever other essential service, via multiple service providers. Whatever multi-home means, why can't there just be one service provider that does that? And are we sure that these service providers aren't already doing that as best we might hope for? (For instance, Amazon already has multiple zones, etc.) I suppose the one thing this can't protect against is some sort of political (broadly defined) threat related to the company itself.
- lxgr 5y ago> Whatever multi-home means, why can't there just be one service provider that does that? Many of these outages are due to pushing broken artifacts or configuration to production. A single provider can pretty easily offer geographic or network topological redundancy, but administrative and/or technological independence is pretty hard to achieve in a single company.
- knute 5y agoI believe EasyDNS can automatically push DNS settings to Route53 to host DNS in AWS. Doesn't protect you from fat-fingering a change, but you should be resilient to either EasyDNS or Route53 going down. https://kb.easydns.com/knowledge/easyroute53/ https://kb.easydns.com/knowledge/easyroute53/
- orblivion 5y agoI mean, I guess what I'm saying is that in theory a single provider could purposely keep two different departments that manage their own artifacts independently.
- 5y ago
- geocrasher 5y agoPeople don't believe me when I say how much DNS matters. So I wrote a song about it. https://soundcloud.com/ryan-flowers-916961339/dns-to-the-tune-of-let-it-be https://soundcloud.com/ryan-flowers-916961339/dns-to-the-tun...
- patleeman 5y agoI just teared up
- geocrasher 5y agoLOL! great comment thank you!
- brianjking 5y agolol, thanks for the laugh.
- Frost1x 5y agoThis made my day, thanks!
- geocrasher 5y agoAnd this, mine! Thanks!
- deleted 5y ago[deleted]
- mvanbaak 5y agoAwesome! Thank you.
- ricardo81 5y agoBrilliant. NOERROR for this.
- zyberzero 5y agoThank you!
- nowahe 5y agoI'm in the middle of a migration from Akamai to Cloudfront, time to take a break I guess
- deleted 5y ago[deleted]
- didjathinkmess 5y agoCyberpolygon already? Thought we had at least a month or two
- penultimatebro 5y agoShh, normies are not ready for that. It’s just a completely random DNS outage, nothing more.
- deleted 5y ago[deleted]
- testplzignore 5y agoStrange thing about the duration of this outage... From logs I have, it seems to have lasted exactly one hour, from 15:38 to 16:38. Their Twitter account also said "disruption lasted up to an hour", though they incorrectly said it started at 15:46 (did it take 8 minutes for their monitoring to notice?). That makes me think that whatever the fix was, it had to wait for some one-hour cache to expire before it took effect. I'm very interested to find out what the cache issue was, more so than what the original bug was.
- throwawaysha 5y agoI ran DNS servers, among other things, in the late 90s with better uptime than these "multi-DC/AZ/geo redundant" services everyone uses these days.
- topranks 5y agoWith all due respect, having also run auth DNS servers in the 90s, and seen the inside of Akamai’s CDN/DNS setup more recently, it isn’t remotely at the same level of scale or sophistication.
- throwawaysha 5y ago"Scale and sophistication" scale relatively with time. Those servers we ran were relatively at the same level of scale and sophistication for their time. The only differentiator here is uptime, which has gotten worse as time has gone on. Five 9s used to be the standard. Three 9s seems to be the new standard.
- aliswe 5y agoNot only that their support telephone line (in sweden) was down as well