3 ms·
The main thing that comes to mind is why they do not deploy these kind of changes to a small slice and smoke test the slice before deploying to all users? This
by roskilli 13y ago
The main thing that comes to mind is why they do not deploy these kind of changes to a small slice and smoke test the slice before deploying to all users? This seems to be a pretty common routine for services at scale nowdays...?
- gfodor 13y agoIt's almost certain that they do in general do this and the fact that this is not what happened is part of the issue. Any blog post describing something like this is going to leave out details such as this.
- pfg 13y agoI'm sure they use Canary Deployments, Gradual Rollouts and what have you to update their services. I suppose this is a hard problem to solve on a configuration change level though. Imagine the configuration change that triggered the bug was something like "hey load balancers, stop sending traffic to the cluster with that new version of service X which seems to cause elevated error rates." You don't really want that kind of change to take too long to propagate.
- dudus 13y agoDepending on what these configuration files are used for it might not be ideal to update only part of the clusters. That might leave the system in an inconsistent way.
- spiderPig 13y agoSeems like the bug was in the config generator/deployer and not the service itself. So it's quite possible that things behaved normally during their dogfooding/smoke testing phase. But you're right, the config should've probably been rolled out region-by-region with a short baking period inbetween instead of a global outage.