11 ms·
Cloudflare outage should not have happened
- etchalon 11mo ago"This massive, accomplished engineering team whose software operates at a scale nearly no one else operates at missed this basic thing" is a hell of a take.
- zahlman 11mo agoHonestly it's a quite lukewarm take. See for example https://danluu.com/algorithms-interviews/ https://danluu.com/algorithms-interviews/. This sort of thing happens constantly.
- mikece 11mo agoYes, pretty basic looking mistakes that, from the outside, make many wonder how this got through. Though analyzing the post-mortem makes me think of the MV Dali crashing into the Francis Scott Key bridge in Baltimore: the whole thing started with a single loose wire which set off a cascading failure. CF's situation was similar in a few ways though finding a bad query (and .unwrap() in production code rather than test code) should have been a lot easier to spot. Have any of the post-mortems addressed if any of the code that led to CloudFlare's outage was generated by AI?
- bell-cot 11mo ago> ...makes me think of the MV Dali crashing... Yes. Though compared to Cloudflare's infrastructure, the Dali is a wooden rowboat. And CF doesn't have the "...or people will die" safety criticality.
- jacquesm 11mo ago> And CF doesn't have the "...or people will die" safety criticality. I disagree with that. Just because you can't point to people falling off a bridge into the water doesn't mean that outages of the web at this scale will not lead to fatalities.
- bell-cot 11mo agoTechnically true. OTOH...whether you describe it as regulations, an SLA, or otherwise - "150,000 ton freighter destroys a major bridge and kills people" is a far worse violation of expected behavior than "lots of web sites went down".
- jacquesm 11mo agoI see where people use CF and I actually think that 'lots of websites went down' has the potential these days to in aggregate kill far more people than were killed by the Dali losing control over their helm. The Dali accident could also have been avoided by simply requiring ships with the gross tonnage to do damage to the bridge to have mandatory tugs, and I'm not so sure there is a clean and effective solution for the kind of issues that CF can create. They're more like 'the shipping industry' than they are like 'a single out of control vessel'. Keep in mind that half of the health care industry or more uses CF to protect their assets.
- cmckn 11mo agoI agree it should not have happened, but I don’t agree that the database schema is the core problem. The “logical single point of failure” here was created by the rapid, global deployment process. If you don’t want to take down all of prod, you can’t update all of prod at the same time. Gradual deployments are a more reliable defense against bugs than careful programming.
- yodon 11mo ago>Gradual deployments are a more reliable defense against bugs than careful programming The challenge, as I understand it, is that the feature in question had an explicit requirement of fast, wide deployment because of the need to react in real time to changing external attacker behaviors.
- cmckn 11mo agoYeah, I don’t know how fast “fast” needs to be in this system; but my understanding is this particular failure would have been seen immediately on the first replica. The progression could still be aggressive after verifying the first wave.
- packetslave 11mo agoyep, and it was this exact requirement that also caused the exact same outage back in 2013 or so. DDoS rules were pushed to the GFE (edge proxy) every 15 seconds, and a bad release got out. Every single GFE worldwide crashed within 15 seconds. That outage is in the SRE book.
- locknitpicker 11mo agoThis sort of Monday morning quarterbacking is pointless and only serves as a way for random bloggers to try to grab credit without actually doing or creating any value.
- nmoura 11mo agoI disagree. I learnt good stuff from this article and it’s enough.
- locknitpicker 11mo ago> I disagree. I learnt good stuff from this article and it’s enough. That's perfectly fine. It's also besides the point though. You can learn without reading random people online cynically shit talking others as a self promotion strategy. This is junior dev energy manifesting junior level understanding of the whole problem domain. There's not a lot to learn from claims that boil down to "don't have bugs".
- rvnx 11mo agoIt's very similar to LinkedIn posts, where everybody seems to know better than the people actually running the platforms.
- alpinisme 11mo agoNot commenting on the quality of this post but occasional writing that responds to an event provides a good opportunity to share thoughts that wouldn’t otherwise reach an audience. If you post advice without a concrete scenario you’re responding to, it’s both less tangible for your audience and less likely to find an audience when it’s easier to shrug off (or put off).
- galleywest200 11mo ago> You can learn without reading random people online Somebody has to write something in the first place for one to learn from it, even if the writing is disagreeable.
- vessenes 11mo ago"If they had a perfectly normalized database, no NULLing and formally verified code, this bug would not have happened." That may be. What's not specified there is the immense, immense cost of driving a dev org on those terms. It limits, radically, the percent of engineers you can hire (to those who understand this and are willing to work this way), and it slows deployment radically. Cloudflare may well need to transition to this sort of engineering culture, but there is no doubt that they would not be in the position they are in if they started with this culture -- they would have been too slow to capture the market. I think critiques that have actionable plans for real dev teams are likely to be more useful than what, to me, reads as a sort of complaint from an ivory tower. Culture matters, shipping speed matters, quality matters, team DNA matters. That's what makes this stuff hard (and interesting!)
- SoftTalker 11mo agoWhy is being able to "capture the market" something we want to encourage? This leads to monopolies or oligopolies and makes possible various types of abuse that a free competitive market would normally correct. If you're going to step into the role of managing a large percentage of public internet traffic, maybe you need to be held to a different standard and set of rules than a startup trying to get a foothold among dozens or hundreds of other competitors. Something more like a public utility than a private enterprise.
- immibis 11mo agoIt doesn't matter what "we" "encourage". This is a natural selection process: all sorts of teams exist, and then the market decides to be captured by certain ones. We do not prescribe which attributes capture the market; we discover them.
- xeromal 11mo agoI assume wanting a company to succeed is fundamental to hacker news. The world is better of with CF being around for sure
- dzikimarian 11mo ago
- hvb2 11mo ago> A central database query didn’t have the right constraints to express business rules. Not only it missed the database name, but it clearly needs a distinct and a limit, since these seem to be crucial business rules. In a database, you wouldn't solve this with a distinct or a limit? You would make the schema guarantee uniqueness? And yes, that wouldn't deal with cross database queries. But the solution here is just the filter by db name, the rest is table design.
- nine_k 11mo ago* The unwrap() in production code should have never passed code review. Damn, it should have been flagged by a linter. * The deployment should have followed the blue/green pattern, limiting the blast radius of a bad change to a subset of nodes. * In general, a company so much at the foundational level of internet connectivity should not follow the "move fast, break things" pattern. They did not have an overwhelming reason to hurry and take risks. This has burned a lot of trust, no matter the nature of the actual bug.
- whazor 11mo agoThe scale of the outage was so big and global, that the biggest failure was indeed the blast radius.
- echelon 11mo agounwrap() and the family of methods like it are a Rust anti-pattern from the early days of Rust. It dates back to before many of the modern error-handling and safety-conscious features of the language and type system. Rust is being pulled in so many different directions from new users that the language perhaps never originally intended. Some engineers will be fine with panicky behavior, but a lot of others want to be able to statically guarantee most panics (outside of perhaps memory allocation failures) cannot occur. We need more than just a linter on this. A new language feature that poisons, marks, or annotates methods that can potentially panic (for reasons other than allocation) would be amazing. If you then call a method that can panic, you'll have to mark your own method as potentially panicky. The ideal future would be that in time, as more standard library and 3rd party library code adopts this, we can then statically assert our code cannot possibly panic. As it stands, I'm pretty mortified that some transitive dependency might use unwrap() deep in its internals.
- burntsushi 11mo ago> As it stands, I'm pretty mortified that some transitive dependency might use unwrap() deep in its internals. You'll have to go without std and even the `core` library then.
- 11mo ago
- tptacek 11mo agoCloudflare doesn't seem to have called it a "Root Cause Analysis" and, in fact, the term "root cause" doesn't appear to occur in Prince's report. I bring this up because there's a school of thought that says "root cause analysis" is counterproductive: complex systems are always balanced on the precipice of multicausal failure.
- Analemma_ 11mo agoWhen I was at AWS, when we did postmortems on incidents we called it "root cause analysis", but it was understood by everyone that most incidents are multicausal and the actual analyses always ended up being fishbone diagrams. Probably there are some teams which don't do this and really do treat RCA as trying to find a sole root cause, but I think a lot of "getting mad at RCA" is bikeshedding the terminology, and nothing to do with the actual practice.
- tptacek 11mo agoRight, I'm not a semantic zealot on this point, but the post we're commenting on really does suggest that the Cloudflare incident had a root cause in basic database management failures, which is the substantive issue the root-cause-haters have with the term.
- otterley 11mo agoThe layered-swiss-cheese model of understanding incidents tends to map to the real world better than the alternatives.
- cyberax 11mo ago> to find a sole root cause "Six billion years ago the dust around the young Sun coalesced into planets"
- luhn 11mo ago"Workaround: If we wait long enough, the earth will eventually be consumed by the sun." https://xkcd.com/1822/ https://xkcd.com/1822/
- kjuulh 11mo agoIt did happen, and cloudflare should learn from it, but not just the technical reasons. Instead of focusing on the technical reasons why, they should answer how such a change bubbled out to cause such a massive impact instead. Why: Proxy fails requests Why: Handlers crashed because of OOM Why: Clickhouse returns too much data Why: A change was introduced causing double the amount of data Why: A central change was rolled out immediately to all cluster (single point of failure) Why: There are exemptions or standard operating procedure (gate) for releasing changes to the hot path for cloudflares network infra. While the Clickhouse change is important, I personally think it is crucial that Cloudflare tackles the processes, and possibly gates / controls rollout for hot path system, no matter what kind of change they are when they're at that scale it should be possible. But that is probably enough co-driving. It to me seems like a process issue more than a technical one.
- lysace 11mo agoVery quick rollout is crucial for this kind of service. On top of what you wrote, institutionalizing rollback by default if something catastrophically breaks should be the norm. Been there in those calls, begging to people in charge who perhaps shouldn't have been, "eh, maybe we should attempt a rollback to the last known good state? cause, it, you know.... worked". But investigating further before making any change always seems to be the preferred action to these people. Can't be faulted for being cautious and doing things properly, right? I kid you not - this is their instinct. If I recall correctly it took CF 2 hours to roll back the broken changes. So if I were in charge of Cloudflare (4-5k employees) I'd both look at the processes and the people in charge.
- vablings 11mo agoIt does seem insane to me that there isnt a process to catch the panic, unwind back to a reasonable place in the call stack, load the last known good configuration and continue execution as normal. You would go from having a global 2 hour outage to a warning on a dashboard that can be investigated in a timely manner rather than blowing up half the internet
- juujian 11mo agoThey are not going as far as to blame PostgreSQL, but their switch to ClickHouse seems to suggest that they see PostgreSQL as part of the equation. Would ClickHouse really prevent this type of error from occurring? PostgreSQL already has so many options for setting up solid constrains for data entry. Or do they not have anyone on the team anymore (or never had) who could set up a robust PostgreSQL database? Or are they just piggybacking on the latest trend?
- jrm4 11mo agoNothing in this thread about "this should not have happened because Cloudflare is too centralized?" We have far better ideas and working prototypes in terms of how to prevent this from happening again to be up here trying to "fix Cloudflare." Think bigger, y'all.
- nyrikki 11mo agoHindsight bias is always easier but: > FAANG-style companies are unlikely to adopt formal methods or relational rigor wholesale. But for their most critical systems, they should. It’s the only way to make failures like this impossible by design, rather than just less likely. That relational rigor imposes what one chooses to be true, it isn’t a universal truth. The frame problem and the qualification problem apply here. The open domain frame problem == HALT. When you can for a problem into the relational model things are nice but not everything can be reduced to a trivial property. That is why Codd had to as nulls etc.. You can choose to decide that the queen is rich OR pigs can fly; but a poor queen doesn’t result in flying pigs. Choice over finite sets == finite indexes over sets == PEM If you can restrict your problems to where the Entscheidungsproblem is solvable you can gain many benefits But it is horses for courses and sub TC.
- nyrikki 11mo agoI expect the downvotes here but it is important. It doesn’t matter if you get there through Trakhtenbrot or Rice. Codd’s normal form is a projection, it will turn your fancy model logic into classic logic. IMHO it is always something to look for to use as a default, but fails if it is a hard requirement. One classic way to describe the problem is the White king and Alice. > ‘I see nobody on the road,’ said Alice. > ‘I only wish I had such eyes,’ the King remarked in a fretful tone. ‘To be able to see Nobody! And at that distance, too! Why, it’s as much as I can do to see real people, by this light!’ Codd added nulls to handle unknowns or missing data. The proper use of them is a complex subject. But they are required if you care about semantic correctness and not just logical validity in many cases. Diaconescu-Goodman-Myhill theorem[0] will show the equivalence between PEM, finite indexes, and choice [0] https://ncatlab.org/nlab/show/Diaconescu-Goodman-Myhill+theorem https://ncatlab.org/nlab/show/Diaconescu-Goodman-Myhill+theo...
- PunchyHamster 11mo ago> but it clearly needs a distinct and a limit, since these seem to be crucial business rules. Isn't that just... wrong ? Throwing arbitrary limit (vs maybe having some alert when the table is too long) would just silently truncate the list Anybody can be backseat engineer by throwing out industry's best practices like they were gospel but you have to look at entire system, not just the database part
- RenThraysk 11mo agoWould be interesting to see the DDL of the table, to see if it had unique constraints. The query not utilising an unique constraint/index should have raised a red flag.
- ruuda 11mo agoSure, a different database schema may have helped, but there are going to be bugs either way. In my view a more productive approach is to think about how to limit the blast radius when things inevitably do go wrong.
- aforwardslash 11mo agorolls eyes No, their error was that they shouldn't be querying system tables to perform field discovery; the same method in postgresql (pg_class or whatever its called) would have had the same result. The simple alternative is to use "describe table <table_name>". On top of that, they shouldn't be writing ad-hoc code to query system tables, but having a separate library instead to perform those kind of task mixed with business logic (crappy application design). Also, this should never have passed code review in the first place, but lets assume it did because errors happen, and this kind of atrocious code and flaky design is not uncommon. As an example, they could be reading this data from CSV files *and* have made the same mistake. Conflating this with "database design errors" is just stupid - this is not a schema design error, this is a programmer error.
- this_user 11mo agoOf course it shouldn't have happened. But if you run infrastructure as complex as this on the scale that they do, and with the agility that they need, then it was bound to happen eventually. No matter how good you are, there is always some extremely unlikely chain of events that will lead to a catastrophic out. Given enough time, that chain will eventually happen.
- pizlonator 11mo ago> No nullable fiels. If you take away nullability, you eventually get something like a special state that denotes absence and either: - Assertions that the absence never happens. - Untested half-baked code paths that try (and fail) to handle absence. > formally verified Yeah, this does prevent most bugs. But it's horrendously expensive. Probably more expensive than the occasional Cloudflare incident
- wat10000 11mo agoAre there outages that should have happened?
- renewiltord 11mo agoOne of the things I recommend most engineers do when they write a bug is to first take a look and see if the bug is required. Very often, I see that the codebase doesn't need the bug added. Then I can just rewrite that code without the bug.
- k3vinw 11mo agoI was expecting a critique on the centralized nature of the infrastructure and the fragility that comes with it.
- foresto 11mo agoDo you mean Cloudflare's design, or the widespread reliance on Cloudflare? I was hoping for a critique of the latter.
- jmull 11mo agoI think the author is trying to apply a preconceived cause on to the cloudflare outage, but there’s not a fit. E.g., they should try to work through how their own suggested fix would actually ensure the problem couldn’t happen. I don’t believe it would… lack of nullable fields and normalization typically simplify relational logic, but hardly prevent logical errors. Formal verification can prove your code satisfies a certain formal specification, but doesn’t prove your specification solves your business problem (or makes sense at all, in fact).
- block_dagger 11mo agoI initially read the title as "Cloudflare outrage.." and I was thinking how nice someone is thinking of the poor engineers who crashed the Internet.
- yakovsi 11mo agoAdding distinct or group by to a query is not some advanced technic comments are suggesting. It does not slow down development one bit, if you expect distinct result you put explicit distinct in the query, it's not a "safety measure for insulin pumps". Scratching my head what I've missed here, please enlighten me.
- anonymars 11mo agoDISTINCT would just be masking the query bug Random DISTINCT is usually a code smell that indicates an incorrect join / filter
- hodgesrm 11mo ago> I base my paragraph on their choice of abandoning PostgreSQL and adopting ClickHouse(Bocharov 2018). The whole post is a great overview on trying to process data fast, without a single line on how to garantee its logical correctness/consistency in the face of changes. I'm completely mystified how the author concludes that the switch from PostgreSQL to ClickHouse shows the root of this problem. 1. If the point is that PostgreSQL is somehow more less prone to error, it's not in this case. You can make the same mistake if you leave off the table_schema in information_schema.columns queries. 2. If the point is that Cloudflare should have somehow discovered this error through normalization and/or formal methods, perhaps he could demonstrate exactly how this would have (a) worked, (b) been less costly than finding and fixing the query through a better review process or testing, and (c) avoided generating other errors as a side effect. I'm particularly mystified how lack of normalization is at fault. ClickHouse system.columns is normalized. And if you normalized the query result to remove duplicates that would just result in other kinds of bugs as in 2c above. Edit: fix typo
- linsomniac 11mo agoI'd be wanting to have some sort of a "dry run" on the produced artifact by the rust code consuming it, or a deploy to some sort of a test environment before letting it roll out to production. I've been surprised that no mention of that sort of thing in the Cloudflare after-action or here.
- ramon156 11mo agoWhile this blog post is pretty useless, it's a hell of a lot better than the LinkedIn posts about the outage... my god, I wish the "Not interested" button worked.
- necovek 11mo agoI have to disagree on the tests not potentially helping here. Finding the right abstraction layer is hard, but there was obviously no integration test that tested wherever the original query was being constructed and where the output was being used. A single smoke test would have failed the same way their actual infra failed when the change was introduced. Obviously, that's not to say that writing normalized database schemas and formal specification won't reduce the number of problems you will introduce. But people make mistakes anywhere, which could have been the case here with the query even if the DB was in a NF (and it still could have been in their case), or in the formal spec as well. There is no magic bullet for correctness, unfortunately.
- 2d8a875f-39a2-4 11mo agoTFA has a point that it should never have happened, and that CF software engineering practices are likely to blame. But a BCNF (or 5NF or whatever) database without nullable columns wouldn't have prevented it. Formally verified code might have but that remains a pipe dream for any significant code base. The proposed cure is worse than the disease.
- devy 11mo agoAuthor's real cause prevention notes. > 1. No nullable fiels. Is that a typo there? fiels should be fields?
- b-man 11mo agofixed
- 9cb14c1ec0 11mo agoAs an aside, I find it really interesting how Cloudflare has morphed from CDN/DDOS protection into a services conglomerate that many startups could use for every compute need they have.
- 1a527dd5 11mo agoUnless you work at Cloudflare or have worked at Cloudflare I'm not sure opinions like this help. You don't know the context, you don't know _anything_ except for what Cloudflare chooses to share. There are very few companies who deal with the kind of load that Clouldflare does, I dread to think what weird edges cases they've run into because of their sheer scale.
- IshKebab 11mo agoCasually suggesting formally verifying the software too.
- knorker 11mo agoNo, this is nonsense and look like university student naivety. What caused it was rolling out a change and moving on to the next recipient without checking if the previous task instantly died. You can't prevent all crash bugs, but you can check if you are lasering your whole prod.
- notepad0x90 11mo agoThe real RCA (IMHO) is not simulating outages in production as part of reliability engineering. Whatever process was stuck in a loop, crashed, or whatever service (db, dns,etc..) was unavailable, that outage scenario can be simulated. Changes can have an automated rollback requirement. My take away is that CF has single points of failure they're aware of, and for business reasons, they've decided to not have a redundancy/failover. > ...and formally verified code, this bug would not have happened. That's what I mean, "we should have caught the bug" , yeah, but that isn't reliability engineering. You assume there will be bugs/outages and prepare for them instead. What happens if the entire DB entered a weird state and was spitting out valid results with incorrect values? What happens if it accepts connections and just stalls? You prepare for bugs that don't yet exist, you fix bugs that do exist.
- avereveard 11mo ago"If only the world was perfect the world would be perfect" Author fails to mention how to actually formally verify this asynchronous globally replicated product. He may have solved the delivery theorem and if that's so I encourage him sharing the results. > No nullable fiels. Author appears to have not formally verified his post's grammar.
- testemailfordg2 11mo agohttps://blog.cloudflare.com/18-november-2025-outage/ https://blog.cloudflare.com/18-november-2025-outage/ "Customers deployed on the new FL2 proxy engine, observed HTTP 5xx errors. Customers on our old proxy engine, known as FL, did not see errors, but bot scores were not generated correctly, resulting in all traffic receiving a bot score of zero." This simply means, the exception handling quality of your new FL2 is non-existent and is not at par / code logic wise similar to FL. I hope it was not because of AI driven efficiency gains.
- mvkel 11mo agoThis piece feels a lot like someone criticizing an umpire's call after watching the slo-mo fifteen times and concluding the ball was actually a strike. Way different from the umpire's pov
- mediumsmart 11mo agoCloudflare is actually an internet outage waiting to happen.
- ku1ik 11mo agoAlso please appreciate how fast this site is. The average website bloat is imperceptible until you open a page like this.