3 ms·
AWS instances usually "randomly fail" because the underlying hardware has issues. I still don't think it's truly random, but the problem is that you don't real
by chousuke 6y ago
AWS instances usually "randomly fail" because the underlying hardware has issues. I still don't think it's truly random, but the problem is that you don't really get access to any direct indicators that the host is about to fail before it does. You don't even truly know how old the hardware is, so your risk mitigation strategy has to assume that anything can fail at any time with zero indication of issues beforehand.
When you manage your own physical servers, you have more knowledge of your risk. The actual time of failure will still be random, but if you've been running a host for 5 years straight, you know the risk is growing.
But we're mostly agreeing here. In the scenario where you throw away a "randomly failed" instance, the historical stability is good evidence that it is due to a hardware failure, and you can just replace the instance and move on.