5 ms·
You can't add a new thing without adding a new point of failure. Every point is a point of failure. Who deploys the thing? A human? God knows they can screw i
by newobj 10y ago
You can't add a new thing without adding a new point of failure. Every point is a point of failure.
Who deploys the thing?
A human? God knows they can screw it up.
Automated deployment?
Well, that's how you get a simultaneous failure and total outage.
Automated incremental deployment?
Ok, slower road to total outage.
Automated incremental that will halt itself or rollback based on reliability metrics?
Ok, getting there.
Wait, was the local proxy load tested?
Was it load tested when one of your data centers is down and everything is doing 30% more work?
And on and on and on. It's all operational overhead, it's all ways to fail.
Can you tell I used to work in monitoring? Maybe I just have PTSD now. :P
- mattrobenolt 10y ago> You can't add a new thing without adding a new point of failure. Every point is a point of failure. Correct, but it's an existing process. So you're right, we could ship a blatantly bad config. > Who deploys the thing? We do, humans, yes. We can definitely screw up a config. > Automated deployment? We tend to do blue/green deploys on critical pieces of infrastructure just to sanity check it. We might even pull a node out of production, test on a staging server, etc. > Wait, was the local proxy load tested? Yes. The load we need for this case is not even close to significant. > Was it load tested when one of your data centers is down and everything is doing 30% more work? Yes, it's literally just a proxy to S3 doing no additional work. For our traffic, the load is not a concern. Especially since it's running on every machine, it's distributed pretty well. A single box cannot overload our haproxy process compared to the CPU needed to run the Python application itself. > Can you tell I used to work in monitoring? Maybe I just have PTSD now. :P DataDog, it's pretty dope. It gives us lots of super good insight into all of these things and is what alerted us because haproxy reported S3 down in the first place. It'd also tell is the moment a process like this crashes, etc.
- newobj 10y agoDataDog was my customer...
- deleted 10y ago[deleted]