4 ms·
It is definitely a neat feature. And just scratches the surface of what workers and durable objects can do. To what extent though, would such a feature be nec
by network2592 5y ago
It is definitely a neat feature. And just scratches the surface of what workers and durable objects can do.
To what extent though, would such a feature be necessary if the application had been designed properly from the ground up to scale? That could still mean using workers and durable objects.
Having the logic handled on workers themselves rather than some inherently unscalable legacy stack may be a better solution. Bottom line, if your application needs a waiting room, you have more fundamental problems.
This is a poor use of the engineering talents at Cloudflare. Probably something pushed for by the sales team seeking to satisfy potential clients using legacy stacks. Going forward, hope this does not skew the objectives of the engineering teams.
What would a better use of engineering talents at Cloudflare be while addressing the same root problem of overload? Exploring the ways in which the needless complexity of kubernetes can be supplanted - most notably with shared multiple tunnel instances.
- kentonv 5y agoFor the use cases that Waiting Room is designed for, "just design for scale" is really not so simple. In a "typical" web app, you have a stateless web server backed by a distributed database. Each user mostly interacts with their own data, or collaborates with a few other users. In this case, scaling is "easy": just add instances, to both the web server and the database. But now imagine you're trying to distribute a limited number of vaccine appointment slots (or concert tickets, or whatever) to a huge number of people. This is a very different problem, because everyone is accessing the same data. "Just add instances" may work for your web server layer, but it won't necessarily help at the database layer. The most scalable distributed databases are eventually consistent ones. But you can't use those here. You can't afford to accidentally give two people the same appointment slot. So, you absolutely need strong consistency. But strongly-consistent databases will typically have an inherent limit on how many clients can be accessing the same data at the same time. For any one piece of data, there will typically be 3-to-5 designated nodes which form the consensus group for that piece, and at least a quorum of those must each track every single change. That means the throughput of transactions on a single key in the database is limited to what a single machine can process. With really careful database tuning, you might be able to make it work. Perhaps you could carefully ensure that appointment slots are well-distributed across chunks, and make sure to utilize replicas for read operations that don't necessarily need up-to-the-moment consistency. It might work... but needless to say, to most developers a database is a black box, and they have no idea how to think about this kind of tuning. Durable Objects don't magically solve this. They do make "chunks" much more explicit, which might make it easier to think about how to build a scalable solution. But each object has that same inherent scalability limit as a database chunk. Durable Objects are a building block; you still need distributed systems skills to utilize them properly for these kinds of problems. Meanwhile, the kinds of people that have these problems to solve -- local governments, event venues, etc. -- aren't exactly the types of people who employ distributed systems experts. (And even if we did magically have infinitely scalable technology, you still have a problem: humans. If everyone is fighting over the same appointment slots, how do you make sure you aren't just rewarding the person who clicks fastest?) The genius of Waiting Room is that you can hire any random contractor to build your web app, and then you make it scale by slapping Waiting Room on top. There's no need for a distributed systems expert to think about your use case. You don't even have to do much of a load test. (And you don't have to worry that your event will be overrun by professional StarCraft players.) I actually do think this is a great use of engineering talent at Cloudflare, to have one group of distributed systems experts solve the problem in a general way that everyone can then easily apply to their own sites. (Disclosure: I am the tech lead for Cloudflare Workers. Waiting Room was built by a separate team.)
- deleted 5y ago[deleted]
- deleted 5y ago[deleted]
- deleted 5y ago[deleted]
- deleted 5y ago[deleted]
- deleted 5y ago[deleted]
- network2592 5y agoTLDR at the bottom. First off, thank you for taking the time for a thorough reply. Clearly, I am a fan of you and your team's work on workers. There are some issues that are beyond the scope of this reply. Overall though, a well thought out and executed product. On the specific point of the waiting room, however, I have to respectfully disagree. Let's just begin very generally by the mission of Cloudflare - a better internet. Does giving a waiting room for the internet sound like a better internet? Everything Cloudflare does is predicated on reducing waiting times. This is what has informed the expansion of PoPs, the edge cdn, human captcha, and even your own work with workers and v8 isolates to reduce cold start times. At the very least, the waiting room is inconsistent with Cloudflare's mission. Now, I understand Cloudflare has onboarded a bunch of new folks lately. But it remains important that new members are clear and consistent about the mission. Attention to this detail can fall by the wayside when you just need to increase the body count quickly. Next point is to qualify the problem at the root of the waiting room. How many websites can afford to have a waiting room? How many users of those websites will wait patiently in a waiting room for their turn? I am sure you are well aware of the impact of even minuscule delays on user retention. But, let's just point them out for others just in case. Amazon found that 0.1 seconds results in 1 percent less sales. Google found that 0.5 seconds results in 20 percent drop in traffic [1]. The idea that users will wait patiently in an internet waiting room is ludicrous. Dedicating engineering talent to this is just sad. Now let's look at the specific issue of a waiting room for vaccines. The situation of vaccines is of course unique due to supply being heavily controlled. But having an internet waiting room hardly seems to be the solution. For those that want to get vaccinated, the demand is very inelastic. Whether the appointment is at 10:30 or 11:00, they will accommodate their schedule for it. In these pandemic times, if you genuinely want to get vaccinated, what is more important than getting vaccinated? Sure, some folks are reluctant to get vaccinated. But that has more to do with misinformation than not getting their ideal appointment time. Besides, these days, most vaccine appointment slots are empty. Finding solutions to deal with misinformation would be more effective than an internet waiting room. An internet waiting room for vaccination is just the wrong product solution. Now let's look at the more technical side of things. Thank you for the refresher on the distinction between strong and eventual consistency as well highlighting the limits of strong consistency. It may be useful to remind ourselves of the fallacies of distributed computing [2]. Despite what the marketing material of Cloudflare occasionally implies, there are inherent limits with distributed computing. It does not operate on some sort of network laws defying magic. That being said, the answer is in your answer. The Cloudflare group of distributed systems experts can still solve the root problem in a general way for those that do not have distributed systems experts. But, does that way inevitably have to be an internet waiting room? What the technical problem warrants is a strong globally consistent kv data source. It could be in contrast to the current kv store which has weaker consistency. It may mean sacrificing some latency. In other words, a copy of the data will not reside on every edge. But at least, the data will be strongly consistent. If a user undertakes some action that changes the data which potentially affects other users, how could we address that? What other network standard is really good at real time? How about looking at a strongly consistent kv store with built in websocket communication? I could go on about what such a solution would entail. But it is beyond the scope of this reply. I can jump on a call if you think that would help. In short, directing engineering talents towards such a solution would not only be a better use of time as well as more consistent with Cloudflare's mission but also yield more impactful results. TLDR Better internet does not mean an internet waiting room. Better vaccination does not mean an internet waiting room. Better internet means better solution. Better potential solution could be a strongly consistent websocket based kv store. Disclosure: I do not work at Cloudflare. All these statements are based on incomplete information. I am open to completely changing these statements if information is presented which warrants it. These statements should not be construed as personal statements. You can assume I am a person with very little or no knowledge. I will not take any personal offence as a result. [1] https://www.gigaspaces.com/blog/amazon-found-every-100ms-of-latency-cost-them-1-in-sales/ https://www.gigaspaces.com/blog/amazon-found-every-100ms-of-... [2] https://en.wikipedia.org/wiki/Fallacies_of_distributed_computing https://en.wikipedia.org/wiki/Fallacies_of_distributed_compu...