7 ms·
If you use the rotate option with timeout you'll quickly hop to the next working resolver. http://edwin.io/optimized-resolv-conf http://edwin.io/optimized-resol
by datums 14y ago
If you use the rotate option with timeout you'll quickly hop to the next working resolver.
http://edwin.io/optimized-resolv-conf http://edwin.io/optimized-resolv-conf
Do you have data you can share regarding "Amazon EC2 resolving nameserver (172.16.0.23) is unreachable too often" ?
- dsr_ 14y agoYou don't need to do rotation, either -- suppose the first nameserver, when up, is consistently faster than the others. Just setting a timeout of 1 is appropriate, then.
- kvz 14y agoA timeout of 1 can result in false positives. 3 seconds is recommended according to the manpage. Should the occasional false positive be none of your concern, then you still add at least a second to every request. If you make many, that's many seconds.
- datums 14y agoThis is more of a performance optimization and less about a HA solution. With rotation you also avoid being throttled. If that's not a concern based on the # of request you make then you can simply timeout and go to the next resolver on the list.
- kvz 14y agoFrom the manpage: "[rotate] causes round robin selection of nameservers from among those listed". This means in your configuration, the broken nameserver will be hit once every 3 times, resulting in a 1second timeout before trying the other 2. So in every 9 requests your setup results in 3 seconds delay for as long as the nameserver is down. In case of Amazon, you actually prefer their nameserver as it resolves names to LAN IPs where possible. With your strategy, ~66% of my instance hostnames would be resolved to the less efficient WAN IPs, that's with all 3 nameservers being reachable. As for data, there's papertrail and Premium Support tickets with Amazon admitting to their downtimes. In most cases (e.g. the one on Jan 14, 2013 11:27 PM PST) "there was a problem with the underlying host". Remember it doesn't even need to be the service itself going down, there's weaker links in this chain like misconfiguration of the virtual host. Cables can get tripped over. Anycast is not possible here. I'm not blaming them. This article was about mitigating suckyness in case of outages, not proving that such a thing even exists : )
- datums 14y agoSo there are 2 different things being addressed. There's a HA solution (handle failovers) and then there's operational performance. Introducing cron to manage this solution, means cron needs to be running and monitored at all times. Also worse case scenario if cron dies and you have a now down nameserver configured. Agreed on aws I would use their resolver to request the internal ip. Another option is setting up your environment in VPC. With VPC you'll be able to assign a static internal ip to your instance. btw I did find many reports on the 172.16.0.23 lookup issues.
- kvz 14y agoWe notice the outages because customers tell us to upload x to y.com. If that fails because of 'unable to resolve y.com', we get a support ticket. I guess for your average web 2.0 the effects are less harsh. Also as Amazon said, these outages can be local due to bad iron. It could have been we've been unlucky with that. At any rate, I was pressed to do something about it. Good suggestions otherwise. I'm not sure what will happen with an RDS Multi-AZ failover if you address it by it's static VPC IP? Also about your worst case scenario. It's true that I rely on cron to be working. While other proposed solutions rely on more obscure daemons to be up & running, there still is a risk, and it could be mitigated by just writing all the nameservers to the resolv.conf. I've created an issue for that: https://github.com/kvz/nsfailover/issues/1 https://github.com/kvz/nsfailover/issues/1