4 ms·
AWS uses health checks to solve this problem. When one of your load balanced instances does not respond to a health check often enough for a fixed size window,
by waterside81 12y ago
AWS uses health checks to solve this problem. When one of your load balanced instances does not respond to a health check often enough for a fixed size window, AWS automatically takes your instance offline. It works pretty well.
http://docs.aws.amazon.com/ElasticLoadBalancing/latest/DeveloperGuide/configure-healthcheck.html http://docs.aws.amazon.com/ElasticLoadBalancing/latest/Devel...
- coderzach 12y agoIt doesn't solve the problem where one machine 500's 5% of the time while all the others only 500 0.5%. Health checks also don't work if your health check ping path is healthy but some other part of your app is broken.
- jrochkind1 12y agoWhy doesn't it solve that problem? If the load balancer is monitoring this, it's quite capable of knowing what the average number of 500's returned is (knowing it automatically, via monitoring over time), and noticing that one machine is returning 500's at 10x the rate every other machine is, and then reducing it in the rotation. No? It would not surprise me if this would lead to yet _other_ undesired side effects in certain conditions though -- I'd expect the same of the solution OP proposes though, or anything automated like this.
- nbm 12y agoPretty much all load balancers have health checks - active where they reach out to each server, or passive where they observe the responses of existing requests if they can. One of the issues is making your active health check more like a doctor's physical than "'tis but a scratch" self-reporting. But also ensuring you're not dealing with a whole bunch of hypochondriacs. Passive health checks at least have the property that they fail servers when the servers are unable to serve, even if the active health check does not consider some subsystem in its response. But alone they can easily be fooled by really fast non-error responses. Anyway, saying "name of brand of load balancers" solves this problem is only covering the most basic cases. General solutions are at best only the first step of the full solution. You need to think about the edges - which I suspect is what Rachel is advocating.
- klaruz 12y agoCloudwatch does what you're referring to as well. It's more of a basic server monitoring system that happens to integrate with the load balancer. You get a set of basic VM level metrics, and you can feed it custom metrics from your app, or log files. All of which can be configured to alarm. I don't think it's possible to run advanced statistics on the metrics for alarming (eg, standard deviation from 30 minute exceeds N), but it may be. Usually it's just an event count, like more than N 500 errors over X time. I do agree you need to think deeper than basic health checks though, 'broken server' is always a hard boolean to nail down.
- eranation 12y agoYep, you can configure the number of failed requests to make it unhealthy, and number of successful ones to mark it healthy, also the timeout, and the polling time (e.g. every 30 seconds). But as others also pointed out, you can still get some "captured" servers until the LB realizes the machine is not super fast but simply superbad. Also your health check URL might return 200 and the rest of your app 404 / 500, which makes me think Perhaps the LB should be aware that if it gets back way more 404 and 500 than average for known URLs, then this should be considered as a bad server... I assume advanced LBs support that?
- citrin_ru 12y agoIt is not always practical to create checks for all possible errors. E. g. broken node can return 404 for URL-X and 200 for all other used URLs. And this URL-X can be "hot" at that time. I think it is not bad to monitor response time for backends and flag some alert if one node has response time significantly lower than mean among all nodes (with comparable hardware).