4 ms·
Rachel always writes great stuff. There's one point I want to dig in on. > since they were so low in the stack, when they broke, lots of other stuff broke It
by brodouevencode 3y ago
Rachel always writes great stuff. There's one point I want to dig in on.
> since they were so low in the stack, when they broke, lots of other stuff broke
It is interesting that while engineers are far too happy to complain about dependency hell in terms of libraries and languages they don't complain enough about dependency hell when it comes to systems. Engineers should take a more proactive role in identifying and pushing whatever changes are required to prevent mass outages due to single points of failure. In other words - if more than a significant percentage of your business requires that X system be up and running, then your business has made a grave error. Engineers should be the first to point this out.
"That's the architect's job"
My thoughts on the usefulness of architecture aside, it's also the engineer's job. Not only should an engineer identify and be aware of system dependencies they should also build software that allows, prevents, and gracefully handles system dependency breakage. This is why the engineer should understand the business as much as the technical. (this is especially true for leads)
"But I inherited this system"
Yeah, it sucks. But now you've got work to do to mitigate risk. Risk mitigation outside of the context of cybersecurity is often overlooked. This is not an advocation for deploying to every data center/cloud provider in the world in the name of high availability. You'll need to do the math to make sure it works for those critical pathways.
EDIT: clarity
- michaelt 3y ago> Engineers should take a more proactive role in identifying and pushing whatever changes are required to prevent mass outages due to single points of failure. In my experience even if you've eliminated single points of failure you can still get failures. Sure, my server's got dual NICs and dual power supplies. But if the guy sent to replace server 12 in rack 345 accidentally gets server 12 in rack 346 I'm going to lose both NICs and both power supplies at once. Sure, the network is fully redundant. But we naturally need to keep the settings on the two sets of kit in sync. We've had outages due to them getting out of sync in the past, not going to make that mistake again. So of course if there's a change to the firewall rules it automatically gets rolled out to both of the firewalls. And so on.
- brodouevencode 3y agoYes, this is true. There's no such thing as 100% redundancy or uptime in the real world without significant cost. The likelihood of that happening goes down exponentially with each successive layer of redundancy or fallback in my experience.