3 ms·
The article considers efficiency from the client side, but it doesn't really consider scaling effects on the server side. Simple polling requests coming in are
by fooblitzky 7y ago
The article considers efficiency from the client side, but it doesn't really consider scaling effects on the server side. Simple polling requests coming in are easy to load-balance, either randomly or by a round-robin scheme. It's resilient, because if one request fails, the next will likely succeed and the user won't even notice.
On the other hand, the complexity of trying to manage millions of open web socket connections filled with state would make things difficult. You would have an uneven distribution of load across servers, the challenge of knowing which connections to close in the lb, and probably the best way to keep the user experience smooth would be to have the client assume failure rather than delay, and aggressively disconnect and establish a new web socket anyway.
- xt00 7y agoDefinitely, 90s tech sometimes is the right answer.. HTTP request compared to HTTPS is far less back and forth between the server and the client, and websockets (I've used them) don't scale well into the super high concurrent connection counts (like more than 1 million) -- would love to know if there is an implementation that does? Or if its just implemented by lots of parallel servers each running like 100k websockets each?
- rauhl 7y agoAgreed that websockets would be overkill and probably less reliable, but why not Server-Sent Events[0]? They are quite simple and should get the job done. For unidirectional information I find that they’re a pretty decent technology, and one which is surprisingly-often overlooked. 0: https://www.w3.org/TR/eventsource/ https://www.w3.org/TR/eventsource/
- matt_oriordan 7y agoYup, the article covers SSE, long polling and XHR streaming as all vastly superior solutions that would have improved things. So yes :+1:
- q3k 7y agoI would guess it's because they likely weren't supported yet by Google infrastructure at the time of implementation. Or that it was must simpler to just add a pollable endpoint.
- paulddraper 7y agoBrowser support may be holding SSE back. [1] It will be nice once Edge (as branded Chromium) supports it. [1] https://caniuse.com/#search=server%20side%20events https://caniuse.com/#search=server%20side%20events
- Matthias247 7y agoThose still tie up request handlers and connections on the backend side, whereas short request/response requests do not. Those things can in some cases operationally introduce a lot more pain (e.g. in terms of resource exhaustion) than the additional short lived requests.
- matt_oriordan 7y agofooblitzky it's interesting you think that handling millions of open websocket connections with state is needed. Firstly, all of these browsers already have connections to Google's servers. HTTP keeps a connection pool open, so keeping connections open is still needed regardless of the transport. Given that, why not use long polling? No state is needed in that situation. And if you're doing that, what's wrong with one additional step, and stream updates over an XHR connection, which too is effectively stateless given any request can fail, reconnect, and subscribe to the stream of updates. I appreciate streaming is harder than polling. But IMHO, if that's Google's thinking, that feels a little defeatist that has designed and built complex systems like Millwheel (https://ai.google/research/pubs/pub41378 https://ai.google/research/pubs/pub41378) to solve these types of problems.
- fooblitzky 7y agoYou are still describing a situation where a client connects directly to a server. At Google, those client connections are probably kept open at a load balancer. The load balancer has its own connections to servers on the other side, and those will be as short-lived as possible, so that the server is available to service the next request the lb sends its way. If you start talking about server-side push technologies, you are keeping a long-lived connection open from the server, through the load balancer, to the client. That will affect the dynamics of the load balancing - it's not as simple as round-robin or random distribution that short-lived requests allow. Consider, for example, a server that becomes overloaded. With short-lived requests you can just add more hosts to the LB pool. With long-lived requests you might need to start thinking about migrating connections, or forcefully terminating some connections if the server hits some kind of load threshold, as well as adding hosts. It's not impossible, it just makes it more complicated and less reliable, which is not what you want for something that just updates a score. As others have pointed out, it's probably just fetching a static file that gets updated periodically.
- nostrademons 7y agoIt's actually worse than that: IIRC (this is 10 years old), traffic to Google passes through a custom DNS resolver that geolocates the lowest-latency active datacenter; a load balancer; 2 levels of reverse proxies; DDoS protection; a bastion server that lets SRE quickly kill non-critical traffic if it threatens the stability of websearch; the webserver; numerous application servers (for core websearch this could be tens of thousands per request, but for a feature like this it's probably just one); and the storage layer. With various push technologies, all of these need to be made aware of the need for a server push, and modified to handle potentially long-running connections that receive events at the data-source's command rather than at the user's request. With polling, you write the new data into the data source once and then the infrastructure just picks it up periodically. Websockets etc. are great technologies if you can handle all traffic on a single box; I use them extensively for my startup. They are a terrible technology at scale. If you want to run websockets at scale, I would actually recommend pulling all the real-time features into a side-channel that's updated via RPC when a relevant event happens and listens directly on the Internet, and then making that as lightweight as possible so that it can scale vertically and as non-critical as possible so that users don't care if it goes down momentarily. If your users don't really care about up-to-the-second information (as with sports scores), just don't do it, and use polling instead.