4 ms·
When this happens to our websites (when all servers go down), we need to rate limit and/or shut down traffic at the load balancer level as we bring things back
by terryjsmith 16y ago
When this happens to our websites (when all servers go down), we need to rate limit and/or shut down traffic at the load balancer level as we bring things back online, otherwise everything just continues to get swamped and goes right back down. This would be nearly impossible in a P2P network and coordinating it between locations would be an even bigger nightmare. I imagine this is why turning on an entirely new network is a more viable option for them.
- dennisgorelik 16y agoWhy would overloaded server go down? Shouldn't it simply stop serving incoming requests if it's overloaded?
- barrkel 16y agoIs there a meaningful distinction between a server that doesn't serve requests, and a server that is down?
- dennisgorelik 16y agoThere is a meaningful distinction between a server that serves 1% of requests and server that is down. I assume that overloaded server still serves some requests. But my assumption could be wrong, so if you have experience with that -- please share your knowledge. Another thing that baffles me: if I introduce new supernode that is sitting on new IP address -- why would all the traffic suddenly hit that node? Shouldn't it be just gradual increase in requests while more and more Skype clients discover that new supernode?
- terryjsmith 16y agoIt handling 1% of connections would assume that the network was the bottleneck. You are much more likely to use up the CPU, RAM, or other resources before hitting the maximum number of available sockets. In which case things are swapping or waiting for available CPU time and each individual requests becomes seconds or tens of seconds to get handled. For all intents and purposes, that machine is dead. As to your second point, I think you are right. I assume that is what the mega-supernode is: a network of machines who's resources are as high as can be to handle all of the connections and try to beat the bottlenecks.
- moe 16y agoI assume that overloaded server still serves some requests. See http://en.wikipedia.org/wiki/Thundering_herd_problem http://en.wikipedia.org/wiki/Thundering_herd_problem Many types of systems need a warm up period before they can realize their full performance. In web applications a controlled warm up is often needed to prime the caches. In P2P applications - which you can't easily "reboot" as a whole - the restoration of a steady state can be much more complex. Shouldn't it be just gradual increase in requests while more and more Skype clients discover that new supernode? In theory, yes. In practice this seems to be a case of http://en.wikipedia.org/wiki/Cascading_failure http://en.wikipedia.org/wiki/Cascading_failure The remaining supernodes either can't handle the aggregate load alone. Or they are being overwhelmed because the re-connection attempts from clients are not evenly distributed. Shouldn't it be just gradual increase in requests while more and more Skype clients discover that new supernode? In theory, yes. In practice there's probably a lot of http://en.wikipedia.org/wiki/Positive_feedback http://en.wikipedia.org/wiki/Positive_feedback and perhaps even http://en.wikipedia.org/wiki/Monster_wave http://en.wikipedia.org/wiki/Monster_wave going on in the skype network right now.
- jemfinch 16y agoThis has absolutely nothing to do with the thundering herd problem. Why are you linking to seemingly random Wikipedia articles? The GP's question is valid: there's no reason why server software should crash when overloaded instead of simply degrading service.
- lusis 16y agoTrue, there's no reason but that doesn't mean it doesn't happen. Many people (I would argue, rightly) equate degraded service to being out of service. Made-up Scenario: My cluster can handle XXXk users with an SLA of YYms response time. In degraded mode, I'm only handling XXk users with YYYYms response times. I'm not meeting my SLA for the remaining number of users so I am, in essence, offline. As to your specific point of "crashing", look at what happened with 37s. Should the server have crashed? No but there was a bug. The reason you add more capacity in the FIRST place is because the existing number of nodes cannot handle the volume. Depending on any number of bugs, issues or configuration your degraded capacity is for all intents and purposes "crashed". Made-up scenario #2: A single server in your apache configuration can handle 200k concurrent connections reliably with fast response times. Double that load and response times are so long that various devices on the path are timing out the connections as stale. Apache hasn't crashed but it's not really doing anything. Fast failure is an accepted best practice. Shit, it's baked into Erlang. Kill the process, start a new one and move on. Depending on the nature of the crash, you're doing nothing but churning processes and not actually servicing requests. The bigger problem is that people don't design for this type of scenario. Static landing pages. Decoupled services instead of monolithic all-in-one containers. Look at github. That's an awesome example of how to degrade service during an outage. Only certain components are "crashed" because everything is fairly decoupled. Meanwhile there's a guy over here running 4 apps in the same tomcat container that communicate over memory transport with each other or even if he had the common sense to decouple each app, didn't bother to fail fast and was busy spinning up threads trying to communicate with the rest of the services that he can't actually respond to anything externally.
- barrkel 16y agoIt usually doesn't work that way. Usually, when your server is overloaded, instead of serving some fraction of the requests in reasonable time, it tries to queue the requests and serve them in order, and as a result response time goes through the roof. Often, the user at the other end of the causal chain starts getting frustrated and refreshes, adding more requests to the queue. When servers are getting overloaded, you need to throttle the incoming requests aggressively and early, to keep response time reasonable.
- JoachimSchipper 16y agoIn network engineering, systems quite often use RED (http://en.wikipedia.org/wiki/Random_early_detection http://en.wikipedia.org/wiki/Random_early_detection) - proactively drop a percentage of connections/packets/... when loads begins to climb, scaling from 0% (any easily handled load) to 100% (at capacity). In practice, this results in pretty robust systems. (But note that TCP et al. have mechanisms that help this, e.g. slowing down when packets are lost and retrying a connection with a backoff.)
- dalore 16y agoUsually too many requests mean the machine runs out of memory. How it handles that is different, but it usually means the service goes down and might need a restart.