4 ms·
Even in these cases, however, there are usually architectural methods for mitigating this kind of failure. In the case of EMR, using transient clusters with sta
by bshacklett 7y ago
Even in these cases, however, there are usually architectural methods for mitigating this kind of failure. In the case of EMR, using transient clusters with state stored in durable file stores such as EMRFS and vanilla S3 is one such option.
It is inevitable that systems will fail. The best the industry can do is work to reduce the number of failures and understand failure modes well so that they can be planned for. AWS does a very good job of this in my experience.
Regardless of whether applications are hosted in the cloud, on premises, co-located or in some hybrid configuration, it's important to design for that inevitable failure and keep business decision makers in the loop while doing so. Understanding requirements around RPO and RTO are extremely important in developing an architecture which meets the needs of the business, yet is still cost effective.