3 ms·
It doesn't solve the problem where one machine 500's 5% of the time while all the others only 500 0.5%. Health checks also don't work if your health check ping
by coderzach 12y ago
It doesn't solve the problem where one machine 500's 5% of the time while all the others only 500 0.5%.
Health checks also don't work if your health check ping path is healthy but some other part of your app is broken.
- jrochkind1 12y agoWhy doesn't it solve that problem? If the load balancer is monitoring this, it's quite capable of knowing what the average number of 500's returned is (knowing it automatically, via monitoring over time), and noticing that one machine is returning 500's at 10x the rate every other machine is, and then reducing it in the rotation. No? It would not surprise me if this would lead to yet _other_ undesired side effects in certain conditions though -- I'd expect the same of the solution OP proposes though, or anything automated like this.