3 ms·
TLDR at the bottom. First off, thank you for taking the time for a thorough reply. Clearly, I am a fan of you and your team's work on workers. There are some i
by network2592 5y ago
TLDR at the bottom.
First off, thank you for taking the time for a thorough reply. Clearly, I am a fan of you and your team's work on workers. There are some issues that are beyond the scope of this reply. Overall though, a well thought out and executed product.
On the specific point of the waiting room, however, I have to respectfully disagree. Let's just begin very generally by the mission of Cloudflare - a better internet. Does giving a waiting room for the internet sound like a better internet?
Everything Cloudflare does is predicated on reducing waiting times. This is what has informed the expansion of PoPs, the edge cdn, human captcha, and even your own work with workers and v8 isolates to reduce cold start times. At the very least, the waiting room is inconsistent with Cloudflare's mission.
Now, I understand Cloudflare has onboarded a bunch of new folks lately. But it remains important that new members are clear and consistent about the mission. Attention to this detail can fall by the wayside when you just need to increase the body count quickly.
Next point is to qualify the problem at the root of the waiting room. How many websites can afford to have a waiting room? How many users of those websites will wait patiently in a waiting room for their turn? I am sure you are well aware of the impact of even minuscule delays on user retention.
But, let's just point them out for others just in case. Amazon found that 0.1 seconds results in 1 percent less sales. Google found that 0.5 seconds results in 20 percent drop in traffic [1]. The idea that users will wait patiently in an internet waiting room is ludicrous. Dedicating engineering talent to this is just sad.
Now let's look at the specific issue of a waiting room for vaccines. The situation of vaccines is of course unique due to supply being heavily controlled. But having an internet waiting room hardly seems to be the solution.
For those that want to get vaccinated, the demand is very inelastic. Whether the appointment is at 10:30 or 11:00, they will accommodate their schedule for it. In these pandemic times, if you genuinely want to get vaccinated, what is more important than getting vaccinated?
Sure, some folks are reluctant to get vaccinated. But that has more to do with misinformation than not getting their ideal appointment time. Besides, these days, most vaccine appointment slots are empty. Finding solutions to deal with misinformation would be more effective than an internet waiting room. An internet waiting room for vaccination is just the wrong product solution.
Now let's look at the more technical side of things. Thank you for the refresher on the distinction between strong and eventual consistency as well highlighting the limits of strong consistency. It may be useful to remind ourselves of the fallacies of distributed computing [2]. Despite what the marketing material of Cloudflare occasionally implies, there are inherent limits with distributed computing. It does not operate on some sort of network laws defying magic.
That being said, the answer is in your answer. The Cloudflare group of distributed systems experts can still solve the root problem in a general way for those that do not have distributed systems experts. But, does that way inevitably have to be an internet waiting room?
What the technical problem warrants is a strong globally consistent kv data source. It could be in contrast to the current kv store which has weaker consistency. It may mean sacrificing some latency. In other words, a copy of the data will not reside on every edge. But at least, the data will be strongly consistent. If a user undertakes some action that changes the data which potentially affects other users, how could we address that? What other network standard is really good at real time? How about looking at a strongly consistent kv store with built in websocket communication?
I could go on about what such a solution would entail. But it is beyond the scope of this reply. I can jump on a call if you think that would help. In short, directing engineering talents towards such a solution would not only be a better use of time as well as more consistent with Cloudflare's mission but also yield more impactful results.
TLDR
Better internet does not mean an internet waiting room. Better vaccination does not mean an internet waiting room. Better internet means better solution.
Better potential solution could be a strongly consistent websocket based kv store.
Disclosure: I do not work at Cloudflare. All these statements are based on incomplete information. I am open to completely changing these statements if information is presented which warrants it. These statements should not be construed as personal statements. You can assume I am a person with very little or no knowledge. I will not take any personal offence as a result.
[1]
https://www.gigaspaces.com/blog/amazon-found-every-100ms-of-latency-cost-them-1-in-sales/ https://www.gigaspaces.com/blog/amazon-found-every-100ms-of-...
[2]
https://en.wikipedia.org/wiki/Fallacies_of_distributed_computing https://en.wikipedia.org/wiki/Fallacies_of_distributed_compu...
- kentonv 5y agoWaiting Room is intended for a very specific use case, where there is an enormous amount of demand for a limited supply. We don't expect most web sites to use it. > these days, most vaccine appointment slots are empty. This is only true in the US and a couple other countries. Most of the world is still waiting. > What the technical problem warrants is a strong globally consistent kv data source. As I tried to explain in my previous reply, there is no general solution that would allow infinite scalability of overlapping transactions. There may be solutions in specific use cases. For example, if you want to store a counter that can be atomically incremented, there are ways to scale that by using a tree topology with branches aggregating batches of requests to submit to the root node. But that's a specific solution to a specific problem, and it involves a use-case-specific API (not a general transaction API). We probably will start offering such specific primitives for specific problems, in due time. But this is beside the point. Even if we offered a magical database that is infinitely scalable, everyone would have to rewrite their applications to use it to get the benefit. If we had offered such a thing at the start of this year, I guarantee you that zero vaccination distribution sites would have had time to build on top of it. You have to consider not just the technical challenge here, but the human organizational challenge of software development and speed of deployment. Waiting Room can be slapped on top of any existing HTTP application built on any stack, without requiring a rewrite. That's valuable, and it has actually helped a lot of people get vaccines in the real world, where a magical new database would not have. Anyway, we're obviously still working on the magical database solutions, too. Building Waiting Room did not distract us from that; the team that built Waiting Room is entirely separate from the Workers team.
- network2592 5y ago> We don't expect most web sites to use it. A fortiori, this indicates it was perhaps not a great use of engineering resources. > Most of the world is still waiting. Sadly true. If anyone wants a sobering overview of vaccination around the world, I suggest checking out this site (https://timetoherd.com/ https://timetoherd.com/). But this has to do with the fact that vaccines are in short supply rather than websites crashing due to http requests. A Cloudflare internet waiting room does nothing to alleviate that global supply problem. > As I tried to explain in my previous reply, there is no general solution that would allow infinite scalability of overlapping transactions. True as well. There is no disagreement here. There will probably never be absolute perfect global consistency in a network of networks with distributed computing and storage. That does not mean you cannot attempt to increase the consistency. Similarly, you will probably never reduce http request response times to 0ms or reduce worker processing times to 0ms. But you can (and are trying to) reduce the latency as much as possible. > For example, if you want to store a counter that can be atomically incremented, there are ways to scale that by using a tree topology with branches aggregating batches of requests to submit to the root node. That seems like an interesting avenue to pursue. It definitely would be a better use of engineering resources. > You have to consider not just the technical challenge here, but the human organizational challenge of software development and speed of deployment. Sure. You also have to consider that this is not just a network engineering problem. It is a product design problem as well. And that there is a risk of over-engineering an unwarranted solution. When I said design to scale from the ground up, I did not mean exclusively from a network standpoint. Let's quickly walk through a potential product design for vaccination. As has been done in many cases, folks can be pre-assigned an appointment time and day and location. There are indications that this actually increases overall attendance. This also has the added benefit of reducing the operational software challenges you allude to. No need to worry much about devops. This would simplify implementation greatly. The appointments can even be batch processed offline one by one asynchronously. I can elaborate on this design and could anticipate your likely objections. But that is beyond the scope of this reply. The takeaway is that there is a product design solution that largely does away with the network engineering problem of having too many synchronous realtime requests for potentially inconsistent data. > That's valuable, and it has actually helped a lot of people get vaccines in the real world Let's assume the synchronous solution for the entire population is the only solution. And that the server crashes returning a 503. It may not be good from Cloudflare's perspective. But the question is did it help or hinder vaccination. You seem to accept as a foregone conclusion that it would hinder. Let's entertain the possibility that it helped. Those that are keen to vaccinate will try again in an hour or whenever the server is back up. If anything, it will more evenly distribute the requests over time. Those folks would vaccinate regardless of whether there is a Cloudflare internet waiting room. How about those that were reluctant to vaccinate? Let's assume they think getting the vaccine is unreasonable. Stories of the website crash may be on the local evening news. These previously reluctant folks may feel like they might miss out. They might begin viewing getting the vaccine as a more reasonable option. You might think this is absurd. Is it as absurd as otherwise reasonable folks hording toilet paper during the early lockdowns? Are you certain none of your Cloudflare colleagues were among those folks? The internet system is complex. The vaccination system is complex. I do not presume to understand all the ins and outs of either let alone both. I am not criticising the team members. I am sure they are competent and well intentioned individuals who just wanted to do their part to help out with the vaccination. What I posit is that the internet waiting room is inconsistent with the mission of Cloudflare for a better internet. Any suggestions that an internet waiting room for vaccination was a net positive for vaccination and better than all other practical alternatives are dubious at best. The information presented thus far has not led me to reconsider this position. But I am still open to reconsidering. You are more than welcome to invite the actual team behind the internet waiting room. Perhaps they can share some insights into their rationale at the time. Hopefully they can approach the exercise with a reciprocal willingness to reconsider their own positions.