5 ms·
It seems to me that the rollout process is flawed. A change should be rolled out to a small fraction of nodes and then monitored for an extended period. 5% of
by eis 4y ago
It seems to me that the rollout process is flawed.
A change should be rolled out to a small fraction of nodes and then monitored for an extended period. 5% of requests returning errors would easily be spotable. Only when a new release has run stable on that portion for some time should it proceed to a bigger subset. You probably want to do this in several steps at the size of CF.
We also learn that it took customer reports to get the investigations rolling but 5xx errors are easily monitored so it points at internal monitoring being lacking even though it's hard for me to believe that they don't have an eye on this already.
It's not the first time that a deploy has brought CloudFlare (partially) down. From the timeline we see there's several hours between incident investigations starting and the rollout being stopped. The rollback should be the first thing considered even before looking at what the actual issue is.
Ideally you have someone sitting next to a red rollback button during a rollout whose only job is it to have an eye on all automatic and customer error reports. :)
- ec109685 4y agoThey did a canary deploy in a data center that didn't have hierarchical caching enabled, so they missed this code path.