22 ms·
DNS Outage at DigitalOcean
- pmalynin 11y agoYeah, tried to access our site and it was down. Really was expecting more out of Digital Ocean than to fuck up such an integral part of their infrastructure. In the future we'll be transitioning away from their DNS solution because this is unacceptable.
- tehbeard 11y agoI hope your clients/users are as understanding and civil as you are. In the meantime, I'm going wait for post-mortem before deciding if I should continue using them for dns. Looking back over the status history, 1-2 incidents a year isn't that bad for my needs, but might be too much for you, which is fine (since I'm only hosting a couple of small side projects with them).
- pmalynin 11y agoThe problem is, for an early stage startup incidents like this are deadly. Especially since we just applied to a bunch of accelerators.
- crisopolis 11y agoThe resolution is, for any app/startup/business everything is a risk and if you didn't include the edge-case of "What happens if my primary DNS nameserver goes down for my domain?" into account. Is all you can do is blame DO? If your app goes down do you have failover for that? Or do you blame your devops team?
- pmalynin 11y agoWe have auto failover for server, app, and database failures -- this can be easily managed. DNS Nameserver failover should have been built-in, after all there is a reason we specify 3 DNS nameservers into the domain configuration, since Digital Ocean took on this task (whereby we used Gandi's DNS before that) we expected it to perform as advertised -- so yes blame lies with DO.
- grej 11y agoAny small business has to manage which risks it accepts. OP got the Digital Ocean service to try and mitigate this risk to a degree. Beyond that, it becomes a question of accept risk in other aspects of product or a risk that your primary DNS server provider will fail? The reality is, you have limited development and financial resources so you simple can’t do everything. Sure, in a vacuum, or in a larger enterprise, we’d love to manage every risk. But when we're just starting out that’s not realistic, and we do have a right to be upset at Digital Ocean's service going down while at the same time realizing that yes, ideally we would / should have had redundancy in place already. Your question on blaming the devops team is exactly the mindset of someone who has a lot more resources than a brand new startup. In a brand new startup there IS NO devops team. If you're lucky, there is one person who does the devops work part time, balanced with a bunch of other development work he/she also does.
- tonyarkles 11y agoI get that this is a pain in the ass. I've got a significant chunk of infrastructure on DO, I've got work to do today that depends on those machines. I learned about this simultaneously when a deployment failed and I got a text from an engineer at a company I consult with. Not a great way to start the day, for sure. Know what I'm going to do? I'm going to have a cup of coffee and play with my dogs for a bit. It's inconvenient, it's going to delay things, and I'm a bit choked about it. But it's not worth getting angry over, because there's nothing I can do about it today.
- ju-st 11y agoBe happy that this happend early. Now you know that you should never ever have a single point of failure.
- tomschlick 11y agoIt's still entirely your fault. Something like Route53/Cloudflare is dirt cheap and crazy redundant. Don't risk your business on free/side services.
- paradite 11y agoYou can go to your domain registrar and switch to another DNS provider (GoDaddy has their own DNS service).
- bitJericho 11y agoJust add a second dns provider.
- clentaminator 11y agoIf your site is so critical that it can't suffer any downtime then why is it not provisioned across multiple independent platforms?
- crisopolis 11y agoI also hope the users of your site understand that shit happens. Also as another user said... if DNS is so critical for you then why don't you have proper failover in place?
- yakshaving_jgt 11y agoThis is the second time recently their AMS region has gone down, which is where I host my email. What a pain.
- grej 11y agoThis is causing huge huge pain for us, Digital Ocean.
- karlgrz 11y agoThis is the first DNS outage I've experienced with them in 3+ years, then again I host everything in their NY regions.
- josh_carterPDX 11y agoSame. They've been pretty reliable. Hoping this doesn't last long or we'll be looking to move off.
- crisopolis 11y agoI've never experienced an outage of any kind with DO, so also first time. I also host all my droplets in the NYC regions.
- chronid 11y agoDNS is hard. Very hard. It may seems trivial when it works (hint: it's not), but some of the biggest fuck ups I've seen in my professional life were caused by strange DNS things happening or DNS servers going kaboom. I feel the pain of the DO engineers trying to mitigate this issue. I really do.
- Thaxll 11y agoIt's not hard, the problem is everything relies on DNS so when DNS goes down or has problems you have cascading failure.
- bitJericho 11y agoThat's why you use multiple providers.
- dsr_ 11y agoSuppose you have multiple providers, but one of them screws up and authoritatively denies the existence of all of your hosts?
- bitJericho 11y agoThat's what you keep an extremely low ttl for.
- Karunamon 11y agoWhich doesn't mean much when a nontrivial amount of ISPs out there don't respect the TTL settings. Source: Days-long service degradation caused by customer ISP's caching bad DNS information well beyond the 10 minute TTL we had set.
- jbaptiste 11y agoEven with multi providers, DNS issues are a cluster fuck.
- coreyp_1 11y agoDoes anyone know of a good strategy for DNS failover?
- c17r 11y agoI don't know if DigitalOcean's DNS servers allow AXFR, if they do you can use a secondary DNS service to automatically replicate the DNS. You then list the secondary DNS as a NS for your domain. If they don't allow AXFR -- and after this, they should! -- you can still have a secondary DNS provider but you'd have to duplicate any changes by hand. Not ideal but still doable.
- deleted 11y ago[deleted]
- mattzito 11y agoWell, there's a couple of strategies: - IP-diverse nameservers - TLD-diverse nameservers - BGP anycast IP-diverse nameservers requires that you expect that your DNS servers will go down rather than start returning bad results - I highly recommend having some sort of mechanism to hard-terminate access to those machines. TLD-diverse nameservers is just an extra strategy for reducing the risk that an upstream TLD issue will blow up your spot. And then BGP anycast is the expensive, complicated piece of this - it requires a high level of technical sophistication, lots of moving parts, and the QA/validation piece of it is tricky. When I built an anycast DNS system, we ended up resorting to tricks like having the DNS servers publish routes to the router for redistribution, so that a down or unresponsive server automatically withdrew the routes. Then you do things like TXT records for your zone that respond with which POP you're hitting in some sort of hashed/obfuscated fashion. It's hard and complicated, and unnecessary for most folks. Better to outsource to Route 53 or someone similar.
- blumentopf 11y ago- Implementation-diverse nameservers Use multiple implementations, e.g. NSD/BIND for authoritative servers and Unbound/BIND for resolvers, to mitigate against implementation-specific bugs and vulnerabilities.
- showerst 11y agoFeeling the pain here too. What DNS providers do others use and like? Route53?
- joejoebob 11y agoWhere I work we use Rotue53. For my personal domains I just use my registrar, Namecheap.
- dsp1234 11y agoFor us, Route53 is painful. We host a few thousand zones, and due to rate limiting on APIs, doing something like "Show a list of domain names" or "Give all the domains matching some pattern" were particularly painful. Upwards of 30 seconds to do a simple list of domains meant we were forced to cache locally. A local cache, combined with the fact that zone names are not unique in their system (possible to create multiple abc.com entries, which differ only in an internal id and the list of NS entries) made it hard to ensure that our internal systems matched "reality". Then the administrative nightmare of 3-4 different NS entries for each zone means customized, rather than generic, instructions for validating NS settings at the individual registrars. All in all, it was not a fun experience with such a large volume of zones, but we knew we were an edge case.
- tomschlick 11y agoYou can contact amazon to get rate limits increased if you have the use case
- rbritton 11y agoThe sites I have that are actually up right now are those routed through CloudFlare.
- dboreham 11y agoBind, running on VMs. Not hard.
- 11y ago
- crisopolis 11y agoDigitalOcean uses CloudFlare for DNS - https://www.cloudflare.com/case-studies-digital-ocean/ https://www.cloudflare.com/case-studies-digital-ocean/
- jtokoph 11y agoThis statement can be misleading. If you read the article, they don't use CloudFlare's DNS servers per se. They use CloudFlare's DNS proxy which acts as a DNS firewall between the DigitalOcean DNS servers and the world.
- tonylemesmer 11y agoPeople hating on DO "I'm losing thousands every hour". Well then should have had some failover in place if its that valuable. [1]https://twitter.com/rodrigoespinosa/status/713035637020971009 https://twitter.com/rodrigoespinosa/status/71303563702097100...
- crisopolis 11y agoI've been reading all the comments on Twitter also... like "err mai gawd I'm switching to AWS because of this" and your failure to not have a secondary DNS provider, but I highly doubt you'd switch. Then another... "Today's @digitalocean DNS #outage is a reminder to not trust your entire business to one provider. Spread the love around!" If your company is e-commerce and makes money by being 99.99% available. It's your own fault for no fail-over. another... ".@digitalocean that's two hours without DNS now...my company's websites could be losing thousands of £ in e-commerce! Please, an update!"
- codegeek 11y agoTimes like this makes you realize the difference betweeen good clients and bad clients. Yes, they have a right to be upset but claims like "could be losing thousands of dollars" is mostly exagerrated due to their frustration.
- crisopolis 11y agoheh yeah, I just laugh at all the tweets saying their losing $billions of dollars every minute their site/app is unavailable. All I can think is... if you're the next Amazon.com I'm pretty sure you'd have some type of disaster plan in place should something like this happen.
- colinbartlett 11y agoI can't disagree with what you're saying, but I think we are all guilty of this. We expect more out of big name services than might be reasonable. (100% uptime) How many of us here have failover email services in case Gmail goes down? I think many companies would say they'd lose thousands in productivity if Google Apps suffers an outage yet I'd hazard that very few have failover plans.
- sashk 11y agoMy rule: provider should do single thing: - Hosting provider - host sites - vps/cloud provider - provide VMs - domain registrar - domain related stuff, but not DNS - dns provider - host dns - second dns provider - host dns in case first dns provider fails So many DNS outages recently and all my projects are up.
- copperx 11y agoDoes Amazon's Route 53 count as a DNS provider, or do you treat it a hosting provider?
- sashk 11y agoFor me - neither. But if I'd be tied in into Amazon's cloud infrastructure, I would have to use many of their features going against my rules above.
- ludbb 11y agoHow do you apply your rules considering what's available today? Which services are you using? It sounds like it would be a big headache to orchestrate the automation among all these different providers.
- josh_carterPDX 11y agoI think the most annoying aspect of this outage are their updates. Three updates and they all say the same thing with no meaningful information as to what's causing this. Likely they may not have much information, but you'd think there would be something more than what they've been posting for the past hour. Good times!
- Rezo 11y agoTheir status page at https://status.digitalocean.com https://status.digitalocean.com is also now giving an intermittent "500 Internal Server Error" nginx error, probably from the load. That's why you should use a service like https://www.statuspage.io https://www.statuspage.io for your important stuff, even though creating a status page is a fun side-project for a dev team.
- crisopolis 11y agoSo what you're saying is that instead of running their own Status Page on their own infrastructure that's reachable. They should outsource it to statuspage.io and pay another company to do it?
- Rezo 11y agoYes, that's pretty standard. Availability monitoring and status reporting should be external and separate from your own infrastructure, otherwise neither may be available when you need it the most.
- dsr_ 11y agoAnd don't use statuspage.io if your host is AWS, because theirs is too.
- TheSwordsman 11y agoEh, as I remember it they have stuff in multiple AWS regions. Would take a global AWS failure to bring them down. This was my justification for going with them while working somewhere that used AWS. That said, while it is extremely unlikely it shouldn't be discounted as impossible.
- Rezo 11y agoThey do have geo-region (not just AZ) redundancy and failover, which puts them quite a bit above most home-grown company status pages in my experience. But yes, if the problem was for example in R53 it would indeed be better to have an solution without that AWS dependency.
- traviswingo 11y agoYeah this is pretty unfortunate. We have some big investor meetings today and this unfortunately took our marketing site offline. Hopefully they resolve this soon - it's the first time we've ever experienced an issue with their service. We really need fail-overs in place...small team problems.
- defenestration 11y agoWe feel the pain as well as our platform is unreachable. I'm now using an other DNS server and changed the nameserver in the domain-record. However the DNS propagation is taking some time. What are you doing at the moment as fail-over?
- pbhjpbhj 11y agoSorry if I'm trying to "teach grandma to suck eggs" but can't you just enter the domain in your local hosts file. If it's a network that needs access then presumably you have some sort of proxy/cache that could be seeded with the necessary domain+IP pairing? I suppose these aren't possible if you're trying to demo on someone else's network or in a public space or such.
- traviswingo 11y agoWe just switched it over to Route53 and set up some fail-overs there. Took us 5 mins and we're back online. Looks like DO is still offline so it seems to have been a good call...
- scurvy 11y agoYou're back online from your perspective. What about all the name servers that have your SOA cached still looking at DO? You're still down for them.
- traviswingo 11y agoTough shit, lol. We now have reduced TTL times for future occurrences, but there's nothing we can do for those users who are still experiencing an outage.
- nlivingstone 11y agoHave multiple VMs @ Digital Ocean (TOR1), we use Cloudflare for DNS... All site have remained available and successfully fulfilling requests.
- cleaver 11y agoEvery site where I was using external DNS stayed up.
- NewHatMatt 11y agoFrom @DOStatus a minute ago: "Our engineering team has identified the issue, and are working to resolve connectivity issues to our DNS servers.... http://do.co/status" http://do.co/status" https://twitter.com/DOStatus/status/713043871559655424 https://twitter.com/DOStatus/status/713043871559655424
- jamescun 11y agoI would be interested in the post-mortem from this. While DigitalOcean operate their own DNS, it is only made publicly available though CloudFlares DNS proxying service.
- samgranieri 11y agoA few years ago Slicehost had a DNS outage and the webscrapers I had running were falling over because they couldnt resolve DNS. I had to SSH into 8 boxes and update resolv.conf to add google DNS and openDNS as a backup. (Yes, I should've had centralized config management with chef or puppet or ansible)
- crisopolis 11y agoThat's crazy... I think by default DO droplets use Google DNS for resolving.
- nodesocket 11y agoRecommend AWS Route53 very highly. Route53 also allows you to buy domain names and do lot's of fancy fail-over, geolocation, and CNAME alias at the apex magic.
- nodesocket 11y agoRecommend Route53 very highly. Route53 also allows you to buy domain names as well and automatically adds the base DNS template.
- tyingq 11y agoOne thing hosting providers could do better would be to split the risk a little by not handing the same dns server name to every client that chooses to have the hosting provider supply dns services. The reason this might have some upside is that DDOS attacks against a specific DNS server are often intended to target one specific customer of a hosting provider. The attacker doesn't care about the side effects...just the original target. Say, for example "controversialblog.com" is hosted on DO, and uses DO dns servers. The person attacking "controversialblog.com" looks up the NS records for the domain, and attacks that DNS server. The fact that it's one hostname that serves all of DO is of little interest to the attacker. So, if DO would come up with say, 10 separate hostnames they could hand out, then this sort of thing would take down 10% of their customers instead of 100%.
- r1ch 11y agoI thought their DNS was supposed to be rock solid since they use Cloudflare Virtual DNS. Oh well, lesson learned. Back to running my own DNS servers on each droplet, if the DNS is down the droplet is likely down regardless :).
- colinbartlett 11y agoIf you want to get alerted when it comes back up, or you wish you had been alerted when it went down, check out my project: https://StatusGator.com https://StatusGator.com. StatusGator monitors status pages and sends notifications via email, Slack, and others. You can get alerted to status changes inside Slack and you can ask it the status of a service with a /statuscheck command.
- camikazeg 11y agoA bit of feedback: you should have a link back to your dashboard on every page. That seems like the most important page to me as a user, but if I am changing my notification or account settings, there is no way back to that page.
- colinbartlett 11y agoGreat feedback, thank you! Added that.
- satyajeet23 11y agoThat awkward moment when it shows the status page
- fredophile 11y agoI don't have anything more important than a small personal website but now I'm curious. If you set up a system to handle your main DNS provider failing, how do you test it? Is there a good reference where I can find some best practices on this?
- mrideout 11y agoHere's my testing recommendation: 1. Pick some subset of your DNS records to monitor, or all of them if you want to be extra thorough. If you are picking a subset, then I'd pick whatever records are most critical to your business. 2. Setup monitoring that queries each of your authoritative name servers for each of the records that you identified in the previous step. The monitoring should notify you if any of the name servers are unresponsive, or return a different response than what's expected. If you'd like to dig into the details of DNS, then O'Reilly's "DNS and BIND" is highly recommended, even if you're not using BIND. There are a number of quality hosting providers out there. A rule of thumb that I use is this: If a DNS hosting provider doesn't eat their own dog food, don't trust them to handle your DNS. Digital Ocean doesn't use their own name servers for their main website's domain. Neither does Amazon. Shameless plug: I created a DNS monitoring service that can be used used for monitoring each of your name servers: https://www.dnscheck.co/ https://www.dnscheck.co/
- tonyarkles 11y agoThe flipside to the dog food point: if DigitalOcean did use their own nameservers for their main site, then we wouldn't have been able to see the status page.
- mrideout 11y agoGood point!
- doublerebel 11y agoNo offense to anyone here, but what is DO's SLA? Last time I looked, they did not have one. DO is cheap for a reason. And that's the same reason I don't host with them, I can get SLA-backed infrastructure for a reasonable price and would have no excuse to my customers or cofounders.
- bpicolo 11y agoLooks like they do have one: https://www.digitalocean.com/help/policy/ https://www.digitalocean.com/help/policy/
- xir78 11y agoWe have seamless DNS "failover" by running dnsmasq with the all-hosts option on all our servers. It causes dnsmasq to query all at once so if any go down its transparent to our apps. Works perfectly on our 1500 ec2 instances.