3 ms·
I think most people are missing the main failure point: Why does one change propagate automatically to all regions? All this could have been contained if they
by NetStrikeForce 10y ago
I think most people are missing the main failure point: Why does one change propagate automatically to all regions?
All this could have been contained if they deployed changes on different regions at different times. That would also help with screwing less your overseas users by running a maintenance at 10am their local time :-)
- aiiane 10y ago> These safeguards include a canary step where the configuration is deployed at a single site and that site is verified to still be working correctly, and a progressive rollout which makes changes to only a fraction of sites at a time, so that a novel failure can be caught at an early stage before it becomes widespread. In this event, the canary step correctly identified that the new configuration was unsafe. Crucially however, a second software bug in the management software did not propagate the canary step’s conclusion back to the push process, and thus the push system concluded that the new configuration was valid and began its progressive rollout. The system does do progressive rollouts, which are essentially what you are referring to (albeit perhaps at a different pace). The number of changes being rolled out means that it's not really feasible to hand roll out configurations to different regions, so the checks are automated. In this case, the automated checks failed as well.
- NetStrikeForce 10y agoI'm not sure you really understand what I've tried to say, but it's probably my fault because of my poor grasp of the English language. You are just confirming my previous comment. Your rollouts are automated, so pushing a change automatically configures every region, instead of configuring just one and maybe waiting for a prudential time in human scale before the next one because, surprise!, shit happens. I understand your colleagues probably make lots of changes, but if that introduces risks of global outages IMHO you should reconsider your strategy. And I'm not sure why you downvoted my previous comment. It's a perfectly valid observation, based on the published information.
- senderista 10y agoWaiting a longer time between regional rollouts (so monitoring systems would have time to detect serious failures) would sacrifice deployment latency, but not deployment throughput (assuming deployments can be made in parallel). For continuous deployment, throughput really matters more than latency.