4 ms·
"cattle vs. pets" misses some important nuance though, especially the way it's usually used to emphasize how you need to architect systems in a cloud environmen
by chousuke 6y ago
"cattle vs. pets" misses some important nuance though, especially the way it's usually used to emphasize how you need to architect systems in a cloud environment that way because you have no way of fixing some issues.
What it is actually saying is not that it's a good approach to just throw away servers when they start having problems. What's important is having that ability when it is necessary.
Servers don't just randomly fail. If your instance goes down, sure, your automation will recover your system, but it will still be important to know why it failed, because it may be a symptom of a deeper issue.
It can be due to a hardware issue, but even when that happens, you don't throw away good hardware; you fix it and the server can return to full operation. In the cloud, though, you have no idea what the hardware is doing. Maybe the instance failed because the underlying host failed; maybe it didn't. You should still find out.
- bigiain 6y agoSure, I totally agree about nuance and the "needing to architect it that way" meaning there. But once you have it architected that way, you then gain the ability to mostly ignore single failures, and only look for "deeper issues" if failures persist. Even at not-very-high scale, AWS instances _do_ "just randomly fail", at least for all practical interpretations. I don't run anything like FAANG scale, only hundreds of instances rather than thousands or millions, and I see at least a few "random failures" a year (not including spot instances terminating, which I see in clumps every month or so). I (almost) never try to repair broken a EC2 instance. Wherever I can, they'll be running totally stateless, and I just provision new ones and kill off old ones. I probably won't even bother investigating if it's a rare and singular problem on a known-reliable platform. If one instance wedges and gets replaced, I'll just have a note to investigate if it happens again any time soon. If we get a second failure, we'll go looking in logs and maybe keep and investigate the EBS volume. For platforms running new-ish code, procedures are different. If we see dead instances after deployments we obviously investigate the new code/config there. But a fair chunk of clients where I am only get 6 or 12 (or even 24) month backend update cycles, if I've got dozens of instances running the same code for months on end and _one_ dies, we just bury it and replace it, and keep a closer eye on the rest of the "herd" for a week or two.
- chousuke 6y agoAWS instances usually "randomly fail" because the underlying hardware has issues. I still don't think it's truly random, but the problem is that you don't really get access to any direct indicators that the host is about to fail before it does. You don't even truly know how old the hardware is, so your risk mitigation strategy has to assume that anything can fail at any time with zero indication of issues beforehand. When you manage your own physical servers, you have more knowledge of your risk. The actual time of failure will still be random, but if you've been running a host for 5 years straight, you know the risk is growing. But we're mostly agreeing here. In the scenario where you throw away a "randomly failed" instance, the historical stability is good evidence that it is due to a hardware failure, and you can just replace the instance and move on.