3 ms·
Sometimes, you have a component which fails in such a way that your redundancies can't really help. I once had to prepare for a total blackout scenario in a da
by chousuke 5y ago
Sometimes, you have a component which fails in such a way that your redundancies can't really help.
I once had to prepare for a total blackout scenario in a datacenter because there was a fault in the power supply system that required bypassing major systems to fix. Had some mistake or fault happened during those critical moments, all power would've been lost.
Well-designed redundancy makes high-impact incidents less likely, but you're not immune to Murphy's law.
- macintux 5y agoTo my mind, among the more frustrating aspects to implementing protection against failure is that the mechanisms to be added can themselves cause failure. It's turtles all the way down.
- chousuke 5y agoYou need to pick your battles and choose what you want to protect against to mitigate risk and enable day-to-day operations. For example, too often people will set up clustered databases and whatnot because "they need HA" without much thought about all the other potential effects of using a cluster, such as much more complicated recovery scenarios. In the vast majority of cases, an active-passive replicated database with manual failover is likely to have fewer pitfalls and gives you the same operational HA a clustered database would, even though in the case of a (rare) real failure it would not automatically recover like a cluster might.