5 ms·
Grrr. So much for global redundancy. What is going to be faster? Updating DNS records with TTL 3600 to point to a single data center or Google fixing their pro
by iowahansen 8y ago
Grrr. So much for global redundancy.
What is going to be faster? Updating DNS records with TTL 3600 to point to a single data center or Google fixing their problem.
We host DNS at AWS, but servers in GCP. Should we use AWS's automatic DNS failover feature to cover for such a case?
- iowahansen 8y agoTraffic is coming back. Looks like Google fixed their load balancer problem within 28 minutes.
- tango24 8y agoHmm, this comment says it’s been happening for hours (below). Maybe their status page isn’t accurate https://news.ycombinator.com/item?id=17552693 https://news.ycombinator.com/item?id=17552693
- partiallypro 8y agoI'm sure it was a cascading event, similar to the one Amazon had yesterday on their own site. Started small until it snowballed and effected everyone.
- jagthebeetle 8y agoI'd wait for further details from the status page, but as a GCP employee (for whatever that claim's worth on the internet), I'm not seeing evidence of an issue earlier than 12:15 PDT.
- colmmacc 8y agoAWS engineer here, I was lead for Route 53. We generally use 60 second TTLs, and as low as 10 seconds is very common. There's a lot of myth out there about upstream DNS resolvers not honoring low TTLs, but we find that it's very reliable. We actually see faster convergence times with DNS failover than using BGP/IP Anycast. That's probably because DNS TTLs decrement concurrently on every resolver with the record, but BGP advertisements have to propagate serially network-by-network. The way DNS failover works is that the health checks are integrated directly with the Route 53 name servers. In fact every name server is checking the latest healthiness status every single time it gets a query. Those statuses are basically a bitset, being updated /all/ of the time. The system doesn't "care" or "know" how many health status change each time, it's not delta-based. That's made it very very reliable over the years. We use it ourselves for everything. Of course the downside of low TTLs is more queries, and we charge by the query unless you ALIAS to an ELB, S3, or CloudFront (then the cost of the queries is on us).
- orf 8y ago> Of course the downside of low TTLs is more queries I was diagnosing a networking issue from one of our service providers last Friday. For whatever indeterminate reason DNS responses from R53 took upwards of 10-15 seconds to return. While I appreciate the non-configurable default TTL of 60 seconds for ELB is not plucked out of thin air and that actual issue seemed to be on the service providers side, the lower limit seems far too low for medium/high latency networks. I wish it was configurable. What's worse is it looks like it's our site that is the issue, so we get the complaints and I have to dig through wireshark logs.
- colmmacc 8y agoIf you have a very high latency network, say a satellite link, make sure that your near-side resolver supports pre-fetching! Unbound is a good choice.
- jniedrauer 8y agoI run unbound on my own workstations. It's so lightweight, you'd never even notice it, but it definitely makes browsing a little more snappy.
- iowahansen 8y agoInteresting, thank you. So a potential mitigation strategy could look like this: - Route 53 failover record * primary record: Google global load balancer IP * secondary record: Route 53 Geolocation set (really need that latency) - Elastic Load balancer record per region * routes to mirror region GCP IP address (ELB's application load balancer seems to able to point to AWS external IPs) * optionally spin up mirror infrastructure in AWS Seems brittle. Does Azure support global load balancing with external IPs? Does anyone have such (or similar) setup actually in production? How did it work today?
- manigandham 8y agoThat would work, and Azure Traffic Manager does support external IPs. CDNs like Cloudflare and Fastly also have built-in load-balancing where they use their internal routing tables for faster propagation.
- ti_ranger 8y ago> We host DNS at AWS, but servers in GCP. Should we use AWS's automatic DNS failover feature to cover for such a case? Well, I would avoid any of GCP's 'Global' features, they are an availability risk. AWS's approach is to rather have inter-region replication, and there are lots of new features that support this.