4 ms·
Surprised it both took them so long to decide to rollback and that the rollback lasted so long.
by jackblemming 4y ago
Surprised it both took them so long to decide to rollback and that the rollback lasted so long.
- yjftsjthsd-h 4y agoYeah, > Once identified we expedited a rollback which completed in 87 minutes. I appreciate that it's a massive infrastructure spanning the entire globe, but... 87 minutes to revert a change? I wouldn't want the job of fixing it, but that doesn't seem good enough when the impact is that bad.
- systemvoltage 4y agoAs an outsider without any knowledge of the company infra: 1) I wonder if they so many layers of CI/CD and testing, that the whole pipeline, even rollbacks probably re-run their tests which takes time. Are CI/CD runners the problem or it is something like network/DNS thing? 2) What if someone inside Cloudflare glibly yells "Why can't we just rsync our code to all data centers in less than a minute?" What would be the pushback this person would get? Why is it a stupid suggestion? Answering that would reveal some crazy layers of abstractions that get piled up in large companies.
- yjftsjthsd-h 4y agoIn all fairness: If nothing else, a way to quickly deploy or revert is also a way to quickly break everything.
- douglasheriot 4y agoWhat happens if a rollback somehow makes things worse? Even emergency rollbacks need a gradual rollout so you’ve got a chance to catch any new issues you didn’t expect.
- yjftsjthsd-h 4y agoAh, good point - I was reading it as 87 minutes for all of it, but it could easily be 1 minute per datacenter or whatever, and intentionally incremental, which would make a lot more sense.
- hayst4ck 4y agoI don't know how cloudflare is configured, but doing anything fast to caches is generally a really bad idea. Caches often protect upstream data stores from overload. If an upstream data store becomes overloaded with too much traffic, then all requests that use that data store will start to slow down. As requests slow down, they will either timeout and fail resulting in major error spikes or they will start to consume all the resources of the service making requests. So if any caches were cleared in the process of rolling back, a slow rollback makes perfect sense to me. I would expect a company like cloudflare would have previous revisions of software on their machines that could be reverted to near instantly. Another thing to consider is what capacity they run at and how long a service takes to initialize. If you can only take 10% of your machines out at any given moment and each one takes a while to initialize, you don't really have much choice in the matter. Lastly, if you look at the 530 graph, it seems that the drop in 530's was almost instant. > 2022-10-25 17:38: An accelerated rollback continues with large data centers acting as Upper tier for many customers. This leads me to believe they shifted traffic away from the low tiers, which enabled them to roll back as slowly as they like at the cost of a likely minor increase in latency.
- dspillett 4y ago> 87 minutes to revert a change? If their processes require a suite of tests be run on staging systems before the revert hits production, that may take considerable time depending on how much is potentially affected. Also the 87 minutes might be for everything, perhaps it was a rolling update: update some nodes, test, update the next batch, repeat. This might mean that a good chunk of the infrastructure was updated for sooner and traffic could be diverted through those parts (presumably those nodes are over-specified, or just easily scaled, to be able to take a glut of extra load - as they would need to in more normal circumstances in order to contend with waxing & waning traffic patterns). Perhaps it is even the case that the errant update didn't get out to all areas before the problem became apparent, so they were not rolling back over the whole network. Again this may offer the opportunity to divert load to non-affected areas while the others are rolled back. So while the full revert took 87 minutes, things may have been up and running properly (except perhaps a little slower than usual) for most users/processes much quicker than that.
- mritzmann 4y ago> [..] and that the rollback lasted so long That does not surprise me. There are companies working with CI/CD systems whose tests easily take more than an hour. And until they are completed, they cannot start releasing (enforced via system permissions). This is a good thing, but the possibility of an emergency rollback should be considered beforehand and maybe (it always depends) exceptions made for such purposes. And sometimes it makes more sense to revert slowly to make sure it doesn't get worse.