7 ms·
This is architectural problem, the LUA bug, the longer global outage last week, a long list of earlier such outages only uncover the problem with architecture u
by mixedbit 10mo ago
This is architectural problem, the LUA bug, the longer global outage last week, a long list of earlier such outages only uncover the problem with architecture underneath. The original, distributed, decentralized web architecture with heterogeneous endpoints managed by myriad of organisations is much more resistant to this kind of global outages. Homogeneous systems like Cloudflare will continue to cause global outages. Rust won't help, people will always make mistakes, also in Rust. Robust architecture addresses this by not allowing a single mistake to bring down myriad of unrelated services at once.
- WD-42 10mo agoIn other words, the consolidation on Cloudflare and AWS makes the web less stable. I agree.
- amazingman 10mo agoUsually I am allergic to pithy, vaguely dogmatic summaries like this but you're right. We have traded "some sites are down some of the time" for "most sites are down some of the time". Sure the "some" is eliding an order of magnitude or two, but this framing remains directionally correct.
- PullJosh 10mo agoDoes relying on larger players result in better overall uptime for smaller players? AWS is providing me better uptime than if I assembled something myself because I am less resourced and less talented than that massive team. If so, is it a good or bad trade to have more overall uptime but when things go down it all goes down together?
- VorpalWay 10mo agoFrom a societal view it is worse when everything is down at once. Leads to a less resilient society: It is not great if I can't buy essentials from one store because their payment system is down (this happened to one super market chain in Sweden due to a hacker attack some years ago, took weeks to fully fix everything, and then there was that whole Crowdstrike debacle globally more recently). It is far worse if all of the competitors are down at once. To some extent you can and should have a little bit of stock at home (water, food, medicine, ways to stay warm, etc) but not everything is practical to do so with (gasoline for example, which could have knock on effects on delivery of other goods).
- pas 10mo agoit's not that simple, no? users want to do things, if their goal depends on a complex chain of functions (provided by various semi-independent services) then the ideal setup would be to have redundant providers and users could simply "load balance" between them and that separate high-level providers' uptime state is clustered (meaning that when Google is unavailable Bing is up, and when Random Site A, goes down their payment provider goes down too, etc..) So ideally sites would somehow sort themselves nearly to separate availability groups. Otherwise simply having a lot of uncorrelated downtimes doesn't help (if we count the sum of downtime experienced by people). Though again it gets complicated by the downtime percentage, because likely there's a phase shift between the states when user can mostly complete their goals and when they cannot because too many cascading failures.
- Lamprey 10mo agoWhen only one thing goes down, it's easier to compensate with something else, even for people who are doing critical work but who can't fix IT problems themselves. It means there are ways the non-technical workforce can figure out to keep working, even if the organization doesn't have on-site IT. Also, if you need to switchover to backup systems for everything at once, then either the backup has to be the same for everything and very easily implementable remotely - which to me seems unlikely for specialty systems, like hospital systems, or for the old tech that so many organizations still rely on (and remember the CrowdStrike BSODs that had to be fixed individually and in person and so took forever to fix?) - or you're gonna need a LOT of well-trained IT people, paid to be on standby constantly, if you want to fix the problems quickly, on account of they can't be everywhere at once. If the problems are more spread out over time, then you don't need to have quite so many IT people constantly on standby. Saves a lot of $$$, I'd think. And if problems are smaller and more spread out over time, then an organization can learn how to deal with them regularly, as opposed to potentially beginning to feel and behave as though the problem will never actually happen. And if they DO fuck up their preparedness/response, the consequences are likely less severe.
- Aeolun 10mo ago> AWS is providing me better uptime than if I assembled something myself because I am less resourced and less talented than that massive team. Is it? I can’t say that my personal server has been (unplanned) down at any time in the past 10 years, and these global outages have just flown right past it.
- Aperocky 10mo agoHave your ISP never went down? Or did it went down in some night and you just never realized.
- UltraSane 10mo agoAWS and Cloudflare can recover from outages faster because they can bring dozens (hundreds?) of people to help, often the ones who wrote the software and designed the architecture. Outages at smaller companies I've worked for have often lasted multiple days, up to an exchange server outage that lasted 2 weeks.
- ivanjermakov 10mo agoRobust architecture that is serving 80M requests/second worldwide? My answer would be that no one product should get this big.
- chickensong 10mo agoYou're not wrong, but where's the robust architecture you're referring to? The reality of providing reliable services on the internet is far beyond the capabilities of most organizations.
- coderjames 10mo agoI think it might be a organizational architecture that needs to change. > However, we have never before applied a killswitch to a rule with an action of “execute”. > This is a straightforward error in the code, which had existed undetected for many years So they shipped an untested configuration change that triggered untested code straight to production. This is "tell me you have no tests without telling me you have no tests" level of facepalm. I work on safety-critical software where if we had this type of quality escape both internal auditors and external regulators would be breathing down our necks wondering how our engineering process failed and let this through. They need to rearchitect their org to put greater emphasis on verification and software quality assurance.
- cyanydeez 10mo agoBro, but how do we make shareholder value if we don't monopolize and enshittify everything
- tobyjsullivan 10mo agoI’m not sure I share this sentiment. First, let’s set aside the separate question of whether monopolies are bad. They are not good but that’s not the issue here. As to architecture: Cloudflare has had some outages recently. However, what’s their uptime over the longer term? If an individual site took on the infra challenges themselves, would they achieve better? I don’t think so. But there’s a more interesting argument in favour of the status quo. Assuming cloudflare’s uptime is above average, outages affecting everything at once is actually better for the average internet user. It might not be intuitive but think about it. How many Internet services does someone depend on to accomplish something such as their work over a given hour? Maybe 10 directly, and another 100 indirectly? (Make up your own answer, but it’s probably quite a few). If everything goes offline for one hour per year at the same time, then a person is blocked and unproductive for an hour per year. On the other hand, if each service experiences the same hour per year of downtime but at different times, then the person is likely to be blocked for closer to 100 hours per year. It’s not really bad end user experience that every service uses cloudflare. It’s more-so a question of why is cloudflare’s stability seeming to go downhill? And that’s a fair question. Because if their reliability is below average, then the value prop evaporates.
- gerdesj 10mo agoAll of my company's hosted web sites have way better uptimes and availability than CF but we are utterly tiny in comparison. With only some mild blushing, you could describe us as "artisanal" compared to the industrial monstrosities, such as Cloudflare. Time and time again we get these sorts of issues with the massive cloudy chonks and they are largely due to the sort of tribalism that used to be enshrined in the phrase: "no one ever got fired for buying IBM". We see the dash to the cloud and the shoddy state of in house corporate IT as a result. "We don't need in-house knowledge, we have "MS copilot 365 office thing" that looks after itself and now its intelligent - yay \o/ Until I can't, I'm keeping it as artisanal as I can for me and my customers.
- foobarkey 10mo agoSorry for the downvotes but this is true many times with some basic HA you get better uptime than the big cloud boys, yes their stack and tech is fancier but we also need to factor in how much CF messes with it vs self hosted, anyway the self hosted wisdom is RIP these days and I mostly just run cf pages / kv :)
- 3rodents 10mo agoWould you rather be attacked by 1,000 wasps or 1 dog? A thousand paper cuts or one light stabbing? Global outages are bad but the choice isn’t global pain vs local pleasure. Local and global both bring pain, with different, complicated tradeoffs. Cloudflare is down and hundreds of well paid engineers spring into action to resolve the issue. Your server goes down and you can’t get ahold of your Server Person because they’re at a cabin deep in the woods.
- jchw 10mo agoIn most cases we actually get both local and global pain, since most people are running servers behind Cloudflare.
- gblargg 10mo agoWhy would there be a centralized outage of decentralized services? The proper comparison seems to be attacked by a dog or a single wasp.
- psunavy03 10mo agoIf you've allowed your Server Person to be a single point of failure out innawoods, that's an organizational problem, not a technological one. Two is one and one is none.
- Lamprey 10mo agoIt's not "1,000 wasps or 1 dog", it's "1,000 dogs at once, or "1 dog at once, 1,000 different times". Rare but huge and coordinated siege, or a steady and predictable background radiation of small issues. The latter is easier to handle, easier to fix, and much more suvivable if you do fuck it up a bit. It gives you some leeway to learn from mistakes. If you make a mistake during the 1000 dog siege, or if you don't have enough guards on standby and ready to go just in case of this rare event, you're just cooked.
- philipallstar 10mo agoI don't quite see how this maps onto the situation. The "1000 dog seige" also was resolved very quickly and transparently, so I would say it's actually better than even one of the "1 dog at once"s.
- UltraSane 10mo agoThey badly need smaller blast radius and to use more chaos engineering tools.
- NicoJuicy 10mo agoYou should really check Cloudflare. There is not a single company that makes their infrastructure as globally available like Cloudflare. Additionally, the downtime of Cloudflare seems to be objectively less than the others. Now, it took 25 minutes for 28% of the network. While being the only ones to fix a global vulnerability. There is a reason other clouds wouldn't touch the responsiveness and innovation that Cloudflare brings.
- rossjudson 10mo agoYou have a heterogeneous, fault-free architecture for the Cloudflare problem set? Interesting! Tell us more.
- JumpCrisscross 10mo ago> Homogeneous systems like Cloudflare will continue to cause global outages But the distributed system is vulnerable to DDOS. Is there an architecture that maintains the advantages of both systems? (Distributed resilience with a high-volume failsafe.)
- terminalshort 10mo agoIt's not as simple as that. What will result in more downtime, dependency on a single centralized service or not being behind Cloudflare? Clearly it's the latter or companies wouldn't be behind Cloudflare. Sure, the outages are more widespread now than they used to be, but for any given service the total downtime is typically much lower than before centralization towards major cloud providers and CDNs.
- rekrsiv 10mo agoOn the other hand, as long as the entire internet goes down when Cloudflare goes down, I'll be able to host everything there without ever getting flack from anyone.
- Klonoar 10mo ago> Rust won't help, people will always make mistakes, also in Rust. They don't just use Rust for "protection", they use it first and foremost for performance. They have ballpark-to-matching C++ performance with a realistic ability to avoid a myriad of default bugs. This isn't new. You're playing armchair quarterback with nothing to really offer.
- m00dy 10mo agoObviously Rust is the answer to these kind of problems. But if you are cloudflare and have an important company at a global scale, you need to set high standarts for your rust code. Developers should dance and celebrate end of the day if their code compiles in rust.
- cbsmith 10mo agoI find this sentiment amusing when I consider the vast outages of the "good ol' days". What's changed is a) our second-by-second dependency on the Internet and b) news/coverage.
- johncolanduoni 10mo agoActually, maybe 1 hour downtime for ~ the whole internet every month is a public good provided by Cloudflare. For everyone that doesn’t get paged, that is.
- jonhess 10mo agoYeah, redundancy and efficiency are opposites. As engineers, we always chase efficiency, but resilience and redundancy are related.
- delusional 10mo agoWhat you've identified here is a core part of what the banking sector calls the "risk based approach". Risk in that case is defined as the product of the chance of something happening and the impact of it happening. With this understanding we can make the same argument you're making, a little more clearly. Cloudflare is really good at what they do, they employ good engineering talent, and they understand the problem. That lowers the chance of anything bad happening. On the other hand, they achieve that by unifying the infrastructure for a large part of the internet, raising the impact. The website operator herself might be worse at implementing and maintaining the system, which would raise the chance of an outage. Conversely, it would also only affect her website, lowering the impact. I don't think there's anything to dispute in that description. The discussion then is if cloudflares good engineering lowers the chance of an outage happening more than it raises the impact. In other words, the things we can disagree about is the scaling factors, the core of the argument seems reasonable to me.
- psychoslave 10mo agoThat's a reflect of social organisation. Pushing for hierarchical organisation with a few key centralising nodes will also impact business and technological decisions. See also https://en.wikipedia.org/wiki/Conway%27s_law https://en.wikipedia.org/wiki/Conway%27s_law
- lxgr 10mo agoNot too long ago, critical avionics were programmed by different software developers and the software was run on different hardware architectures, produced by different manufacturers. These heterogeneous systems produced combined control outputs via a quorum architecture – all in a single airplane. Now half of the global economy seems to run on same service provider, it seems…
- theoldgreybeard 10mo agoNotwithstanding that most people using Cloudflare aren't even benefiting from what it actually provides. They just use it...because reasons.
- steelblueskies 10mo agoReductionist, but it's a backup problem. Data matters? Have multiple copies, not all in the same place. This is really no different, yet we don't have those redundancies in play. Host, and paths. Every other take is ultimately just shuffling justification around the least bad for everyone lack of backups for cost saving.