3 ms·
Most of the criticisms seem to be around BGP and network management. What I’m seeing here that also is important is that the change was applied to a DC where th
by devonkim 4y ago
Most of the criticisms seem to be around BGP and network management. What I’m seeing here that also is important is that the change was applied to a DC where the route change didn’t trigger the defect. In essence, this is also due to a very classic problem of a test dataset giving a false sense of security due to variation from other configurations. For this reason my team prefers to rollout changes to production using a test region that most customers don’t use yet will have some visible impact if there’s any error in our presumptions so far such as hard-coding regions and relying upon services not present or as capable across all regions. This practice has caught a number of rather serious errors for us that while customer impacting was nowhere near as bad as if we had rolled out simply randomly like many teams do essentially. This is even more important the more difficult it is to perform rollbacks of changes or for rollbacks to take effect such as DNS and CDN caching changes.