4 ms·
>The postmortems published by Google, Amazon, and Azure (as well as postmortems internal to the company I work for) are nearly always due to some type of change
by superuser2 10y ago
>The postmortems published by Google, Amazon, and Azure (as well as postmortems internal to the company I work for) are nearly always due to some type of change (code or configuration) being rolled out
When you have a really large distributed system that's primarily running on metal, smaller-scale copies of the whole system per developer or even a single companywide staging environment that mirrors production are really hard, and they don't always exist.
Developers work on their components in isolation by mocking out the rest of the system, hopefully there's a thorough code review, and then "integration testing" happens by flipping a feature flag and watching the logs/metrics in production. You might design the feature flagging so it initially only hits test accounts, but some things (like service communication layers, Puppet configs, router configs) don't work that way.
The cost of outages resulting from the lack of a staging environment may well be less than re-architecting production (100% automation, no snowflakes, and probably some kind of IaaS) to allow for disposable dev/test environments which would exhibit the same bugs.
- adrianratnapala 10y agoBut I think what you are saying actually is an argument for @piinbinary's idea of "static analisys". If you can't really test your change, then the other thing you can do is try to rationally analyse its consequences. After all thats why "hopefully there's a throrough code review". Wouldn't it be nice to automate some of that analysis? I have no idea how to go about such a thing, or whether it is even possible -- but the harder testing gets, the more attractive this option is.