3 ms·
They narrowed down the actual problem to some Rust code in the Bot Management system that enforced a hard limit on the number of configuration items by returnin
by HumanOstrich 11mo ago
They narrowed down the actual problem to some Rust code in the Bot Management system that enforced a hard limit on the number of configuration items by returning an error, but the caller was just blindly unwrapping it.
- otterley 11mo agoA dormant bug in the code is usually a condition precedent to incidents like these. Later, when a bad input is given, the bug then surfaces. The bug could have laid dormant for years or decades, if it ever surfaced at all. The point here remains: consider every change to involve risk, and architect defensively.
- tptacek 11mo agoThey made the classic distributed systems mistake and actually did something. Never leap to thing-doing!
- otterley 11mo agoIf they're going to yeet configs into production, they ought to at least have plenty of mitigation mechanisms, including canary deployments and fault isolation boundaries. This was my primary point at the root of this thread. And I hope fly.io has these mechanisms as well :-)
- tptacek 11mo agoWe've written at long, tedious length about how hard this problem is.
- otterley 11mo agoHave a link?
- tptacek 11mo agoMost recently, a few weeks ago (but you'll find more just a page or two into the blog): https://fly.io/blog/corrosion/ https://fly.io/blog/corrosion/
- otterley 11mo agoIt's great that you're working on regionalization. Yes, it is hard, but 100x harder if you don't start with cellular design in mind. And as I said in the root of the thread, this is a sign that CloudFlare needs to invest in it just like you have been.
- tptacek 11mo agoI recoil from that last statement not because I have a rooting interest in Cloudflare but because the last several years of working at Fly.io have drilled Richard Cook's "How Complex Systems Fail"† deep into my brain, and what you said runs aground of Cook #18: Failure free operations require experience with failure. If the exact same thing happens again at Cloudflare, they'll be fair game. But right now I feel people on this thread are doing exactly, precisely, surgically and specifically the thing Richard Cook and the Cook-ites try to get people not to do, which is to see complex system failures as predictable faults with root causes, rather than as part of the process of creating resilient systems. † https://how.complexsystems.fail/ https://how.complexsystems.fail/
- otterley 11mo agoSuppose they did have the cellular architecture today, but every other fact was identical. They'd still have suffered the failure! But it would have been contained, and the damage would have been far less. Fires happen every day. Smoke alarms go off, firefighters get called in, incident response is exercised, and lessons from the situation are learned (with resulting updates to the fire and building codes). Yet even though this happens, entire cities almost never burn down anymore. And we want to keep it that way. As Cook points out, "Safety is a characteristic of systems and not of their components."
- Ekaros 11mo agoSounds like lack of good testing. Too many items in any input should be a boundary case you will get to eventually.