5 ms·
What's the culture like at Cloudflare re: ops/deployment safety? They saw errors related to a deployment, and because it was related to a security issue instea
by flaminHotSpeedo 10mo ago
What's the culture like at Cloudflare re: ops/deployment safety?
They saw errors related to a deployment, and because it was related to a security issue instead of rolling it back they decided to make another deployment with global blast radius instead?
Not only did they fail to apply the deployment safety 101 lesson of "when in doubt, roll back" but they also failed to assess the risk related to the same deployment system that caused their 11/18 outage.
Pure speculation, but to me that sounds like there's more to the story, this sounds like the sort of cowboy decision a team makes when they've either already broken all the rules or weren't following them in the first place
- deadbabe 10mo agoAs usual, Cloudflare is the man in the arena.
- samrus 10mo agoThere are other men in the arena who arent tripping on their own feet
- usrnm 10mo agoLike who? Which large tech company doesn't have outages?
- k8sToGo 10mo agoIt's not about outages. It's about the why. Hardware can fail. Bugs can happen. But to continue a roll out despite warning sings and without understanding the cause and impact is on another level. Especially if it is related to the same problem as last time.
- udev4096 10mo ago[dead]
- deadbabe 10mo agoIt is healthy for tech companies to have outages, as they will build experience in resolving them. Success breeds complacency.
- wizzwizz4 10mo agoYou don't need outages to build experience in resolving them, if you identify conditions that increase the risk of outages. Airlines can develop a lot of experience resolving issues that would lead to plane crashes, without actually crashing any planes.
- k__ 10mo ago"tripping on their own feet" == "not rolling back"
- nish__ 10mo agoGoogle does pretty good.
- hansonkd 10mo agoGoogle docs was just down a couple weeks ago almost the whole day.
- this_user 10mo agoThe question is perhaps what the shape and status of their tech stack is. Obviously, they are running at massive scale, and they have grown extremely aggressively over the years. What's more, especially over the last few years, they have been adding new product after new product. How much tech debt have they accumulated with that "move fast" approach that is now starting to rear its head?
- sandeepkd 10mo agoI think this is probably a bigger root cause and is going to show up in different ways in future. The mere act of adding new products to an existing architecture/system is bound to create knowledge silos around operations and tech debt. There is a good reason why big companies keep smart people on their payroll to just change couple of lines after a week of debate.
- nine_k 10mo ago> more to the story From a more tinfoil-wearing angle, it may not even be a regular deployment, given the idea of Cloudflare being "the largest MitM attack in history". ("Maybe not even by Cloudflare but by NSA", would say some conspiracy theorists, which is, of course, completely bonkers: NSA is supposed to employ engineers who never let such blunders blow their cover.)
- lukeasrodgers 10mo agoRoll back is not always the right answer. I can’t speak to its appropriateness in this particular situation of course, but sometimes “roll forward” is the better solution.
- echelon 10mo agoYou want to build a world where roll back is 95% the right thing to do. So that it almost always works and you don't even have to think about it. During an incident, the incident lead should be able to say to your team's on call: "can you roll back? If so, roll back" and the oncall engineer should know if it's okay. By default it should be if you're writing code mindfully. Certain well-understood migrations are the only cases where roll back might not be acceptable. Always keep your services in "roll back able", "graceful fail", "fail open" state. This requires tremendous engineering consciousness across the entire org. Every team must be a diligent custodian of this. And even then, it will sometimes break down. Never make code changes you can't roll back from without reason and without informing the team. Service calls, data write formats, etc. I've been in the line of billion dollar transaction value services for most of my career. And unfortunately I've been in billion dollar outages.
- drysart 10mo ago"Fail open" state would have been improper here, as the system being impacted was a security-critical system: firewall rules. It is absolutely the wrong approach to "fail open" when you can't run security-critical operations.
- echelon 10mo agoCloudflare is supposed to protect me from occasional ddos, not take my business offline entirely. This can be architected in such a way that if one rules engine crashes, other systems are not impacted and other rules, cached rules, heuristics, global policies, etc. continue to function and provide shielding. You can't ask for Cloudflare to turn on a dime and implement this in this manner. Their infra is probably very sensibly architected by great engineers. But there are always holes, especially when moving fast, migrating systems, etc. And there's probably room for more resiliency.
- rvz 10mo ago> Not only did they fail to apply the deployment safety 101 lesson of "when in doubt, roll back" but they also failed to assess the risk related to the same deployment system that caused their 11/18 outage. Also there seems to be insufficient testing before deployment with very junior level mistakes. > As soon as the change propagated to our network, code execution in our FL1 proxy reached a bug in our rules module which led to the following LUA exception: Where was the testing for this one? If ANY exception happened during the rules checking, the deployment should fail and rollback. Instead, they didn't assess that as a likely risk and pressed on with the deployment "fix". I guess those at Cloudflare are not learning anything from the previous disaster.
- dkyc 10mo agoOne thing to keep in mind when judging what's 'appropriate' is that Cloudflare was effectively responding to an ongoing security incident outside of their control (the React Server RCE vulnerability). Part of Cloudlfare's value proposition is being quick to react to such threats. That changes the equation a bit: any hour you wait longer to deploy, your customers are actively getting hacked through a known high-severity vulnerability. In this case it's not just a matter of 'hold back for another day to make sure it's done right', like when adding a new feature to a normal SaaS application. In Cloudflare's case moving slower also comes with a real cost. That isn't to say it didn't work out badly this time, just that the calculation is a bit different.
- Already__Taken 10mo agothe cve isn't a zero day though how come cloudflare werent at the table for early disclosure?
- flaminHotSpeedo 10mo agoDo you have a public source about an embargo period for this one? I wasn't able to find one
- charcircuit 10mo agoConsidering there were patched libraries at the time of disclosure, those libraries' authors must have been informed ahead of time.
- Pharaoh2 10mo agohttps://react.dev/blog/2025/12/03/critical-security-vulnerability-in-react-server-components https://react.dev/blog/2025/12/03/critical-security-vulnerab... Privately Disclosed: Nov 29 Fix pushed: Dec 1 Publicly disclosed: Dec 3
- drysart 10mo agoThen even in the worst case scenario, they were addressing this issue two days after it was publicly disclosed. So this wasn't a "rush to fix the zero day ASAP" scenario, which makes it harder to justify ignoring errors that started occuring in a small scale rollout.
- liampulles 10mo agoRollback is a reliable strategy when the rollback process is well understood. If a rollback process is not well known and well experienced, then it is a risk in itself. I'm not sure of the nature of the rollback process in this case, but leaning on ill-founded assumptions is a bad practice. I do agree that a global rollout is a problem.
- newsoftheday 10mo agoRollback carries with it the contextual understanding of complete atomicity; otherwise it's slightly better than a yeet. It's similar to backups that are untested.
- marcosdumay 10mo agoComplete atomicity carries with it the idea that the world is frozen, and any data only needs to change when you allow it to. That's to say, it's an incredibly good idea when you can physically implement it. It's not something that everybody can do.
- newsoftheday 10mo agoNo, complete atomicity doesn't require a frozen state, it requires common sense and fail-proof, fool-proof guarantees derived from assurances gained from testing. There is another name for rolling forward, it's called tripping up.
- programd 10mo agoGlobal rollout of security code on a timeframe of seconds is part of Cloudflare's value proposition. In this case they got unlucky with an incident before they finished work on planned changes from the last incident.
- flaminHotSpeedo 10mo agoThat's entirely incorrect. For starters, they didn't get unlucky. They made a choice to use the same system they knew was sketchy (which they almost certainly knew was sketchy even before 11/18) And on top of that, Cloudflare's value proposition is "we're smart enough to know that instantaneous global deployments are a bad idea, so trust us to manage services for you so you don't have to rely on in house folks who might not know better"
- otterley 10mo agoFrom the post: “We have spoken directly with hundreds of customers following that incident and shared our plans to make changes to prevent single updates from causing widespread impact like this. We believe these changes would have helped prevent the impact of today’s incident but, unfortunately, we have not finished deploying them yet. “We know it is disappointing that this work has not been completed yet. It remains our first priority across the organization.”
- NoSalt 10mo agoOoh ... I want to be on a cowboy decision making team!!!
- deleted 10mo ago[deleted]
- ignoramous 10mo ago> this sounds like the sort of cowboy decision Ouch. Harsh given that Cloudflare's being over-honest (to disabling the internal tool) and the outage's relatively limited impact (time wise & no. of customers wise). It was just an unfortunate latent bug: Nov 18 was Rust's Unwrap, Dec 5 its Lua's turn with its dynamic typing. Now, the real cowboy decision I want to see is Cloudflare [0] running a company-wide Rust/Lua code-review with Codex / Claude... cf TFA: if rule_result.action == "execute" then rule_result.execute.results = ruleset_results[tonumber(rule_result.execute.results_index)] end This code expects that, if the ruleset has action="execute", the "rule_result.execute" object will exist ... error in the [Lua] code, which had existed undetected for many years ... prevented by languages with strong type systems. In our replacement [FL2 proxy] ... code written in Rust ... the error did not occur. [0] https://news.ycombinator.com/item?id=44159166 https://news.ycombinator.com/item?id=44159166
- deleted 10mo ago[deleted]
- NicoJuicy 10mo agoWhere I work, all teams were notified about the React CVE. Cloudflare made it less of an expedite.
- crote 10mo ago> They saw errors related to a deployment, and because it was related to a security issue instead of rolling it back they decided to make another deployment with global blast radius instead? Note that the two deployments were of different components. Basically, imagine the following scenario: A patch for a critical vulnerability gets released, during rollout you get a few reports of it causing the screensaver to show a corrupt video buffer instead, you roll out a GPO to use a blank screensaver instead of the intended corporate branding, a crash in a script parsing the GPOs on this new value prevents users from logging in. There's no direct technical link between the two issues. A mitigation of the first one merely exposed a latent bug in the second one. In hindsight it is easy to say that the right approach is obviously to roll back, but in practice a roll forward is often the better choice - both from an ops perspective and from a safety perspective. Given the above scenario, how many people are genuinely willing to do a full rollback, file a ticket with Microsoft, and hope they'll get around to fixing it some time soon? I think in practice the vast majority of us will just look for a suitable temporary workaround instead.