3 ms·
There's a particular class of bug that affects all datacenters in a globally distributed service. Almost invariably it comes down to knock-on effects due to mul
by throwaway290232 5y ago
There's a particular class of bug that affects all datacenters in a globally distributed service. Almost invariably it comes down to knock-on effects due to multiple things happening "the wrong way" in tandem. Those individual failures typically happen because somebody said "well yeah we should have a better way to do this, but the other components have redundancy, so shrug". If a single domino isn't tested well and has all potential failure modes well documented and tested for, it can still take down the entire chain of dominos.
- jedberg 5y agoOr just a bad global configuration change. We used to take down Netflix globally with config changes until we changed the system to have config changes scoped to single regions. And even then sometimes the failure wouldn't show up until multiple regions got the config change.
- mrkurt 5y agoAlways bet on a config change. Or DNS.
- pgporada 5y agoNeither of these apply to this situation though. We're working on it.
- mrkurt 5y agoI lose many of my bets.
- terom 5y agoIt sounds like it was a hardware issue [1]. First rule of electronics applies: thou shalt check voltages. > [Update] We're conducting follow-up maintenance to address power supply issues from our earlier service disruption. All production API services may be down for up to 30 minutes. [1] https://letsencrypt.status.io/pages/maintenance/55957a99e800baa4470002da/60f5f2460451310998105e7f https://letsencrypt.status.io/pages/maintenance/55957a99e800...
- dane-pgp 5y agoIt sounds like you're quoting from the classic "How complex systems fail": https://how.complexsystems.fail/ https://how.complexsystems.fail/