2 ms·
My personal bet is there's an 80% chance this is caused by some internal bootstrapping problem that they've messed up. AIUI all the main cloud vendors are in tr
by purpleidea 15d ago
My personal bet is there's an 80% chance this is caused by some internal bootstrapping problem that they've messed up. AIUI all the main cloud vendors are in trouble here. The automation project I work on is expressly designed to help folks solve this DR/bootstrapping problem. Soo many people get this wrong. Of course missiles don't help things, but I'd bet AWS is primarily to blame here. I'd love an actual technical report of why they can't recover things.
- jeffrallen 15d agoI work on the same kind of thing, and while we think hard about bootstrap problems, we always find new surprising ones. The problem is you never know until you do it, and creating a faithful test of restarting giant systems is economically impossible. Because if you say to the boss, "look, I need 1 million now to test against a maybe 100 million loss, maybe in 10 years" they don't give you the money (and rightly so).
- AdamN 15d agoEven if you did the $1MM test there is very low likelihood that the $100MM event would be fully mitigated 10 years down the line (after who knows how many changes - physical, logical, and even in the org chart). The only way to approach readiness here is repeated investment - like one team doing the deep dive and another pulling cables and then constantly doing pre- and post-mortems.