4 ms·
This seems like a nightmare situation for P2P company. Your super nodes get knocked offline and there's no way to force update them until they come back online
by terryjsmith 16y ago
This seems like a nightmare situation for P2P company. Your super nodes get knocked offline and there's no way to force update them until they come back online, which is sporadic at best (based on my own experience today), which means either manually updating them or completely re-seeding your network (which seems to be the route they're going -- though so far it doesn't appear to be going all that well).
As other people have stated on other articles posted here, no discredit to them for having some downtime (I'm sure they're working round the clock on it), but I have to wonder if this could have been avoided with a phased roll out of new version to the super nodes.
- dennisgorelik 16y agoWhy is it a problem to bring supernode back online?
- barrkel 16y agoI can imagine it would self-DDoS from lack of peer supernodes to share the load.
- dennisgorelik 16y agoBut at least it would successfully reply to at least some of requests, right? And would still keep trying to serve the requests, right? So if sufficient number of supernodes is brought back online then the problems should disappear.
- terryjsmith 16y agoWhen this happens to our websites (when all servers go down), we need to rate limit and/or shut down traffic at the load balancer level as we bring things back online, otherwise everything just continues to get swamped and goes right back down. This would be nearly impossible in a P2P network and coordinating it between locations would be an even bigger nightmare. I imagine this is why turning on an entirely new network is a more viable option for them.
- dennisgorelik 16y agoWhy would overloaded server go down? Shouldn't it simply stop serving incoming requests if it's overloaded?
- barrkel 16y agoIs there a meaningful distinction between a server that doesn't serve requests, and a server that is down?
- dennisgorelik 16y agoThere is a meaningful distinction between a server that serves 1% of requests and server that is down. I assume that overloaded server still serves some requests. But my assumption could be wrong, so if you have experience with that -- please share your knowledge. Another thing that baffles me: if I introduce new supernode that is sitting on new IP address -- why would all the traffic suddenly hit that node? Shouldn't it be just gradual increase in requests while more and more Skype clients discover that new supernode?
- terryjsmith 16y agoIt handling 1% of connections would assume that the network was the bottleneck. You are much more likely to use up the CPU, RAM, or other resources before hitting the maximum number of available sockets. In which case things are swapping or waiting for available CPU time and each individual requests becomes seconds or tens of seconds to get handled. For all intents and purposes, that machine is dead. As to your second point, I think you are right. I assume that is what the mega-supernode is: a network of machines who's resources are as high as can be to handle all of the connections and try to beat the bottlenecks.
- moe 16y agoI assume that overloaded server still serves some requests. See http://en.wikipedia.org/wiki/Thundering_herd_problem http://en.wikipedia.org/wiki/Thundering_herd_problem Many types of systems need a warm up period before they can realize their full performance. In web applications a controlled warm up is often needed to prime the caches. In P2P applications - which you can't easily "reboot" as a whole - the restoration of a steady state can be much more complex. Shouldn't it be just gradual increase in requests while more and more Skype clients discover that new supernode? In theory, yes. In practice this seems to be a case of http://en.wikipedia.org/wiki/Cascading_failure http://en.wikipedia.org/wiki/Cascading_failure The remaining supernodes either can't handle the aggregate load alone. Or they are being overwhelmed because the re-connection attempts from clients are not evenly distributed. Shouldn't it be just gradual increase in requests while more and more Skype clients discover that new supernode? In theory, yes. In practice there's probably a lot of http://en.wikipedia.org/wiki/Positive_feedback http://en.wikipedia.org/wiki/Positive_feedback and perhaps even http://en.wikipedia.org/wiki/Monster_wave http://en.wikipedia.org/wiki/Monster_wave going on in the skype network right now.
- tommi 16y agoThat can be avoided by altering the logic by which supernodes are queried. Not all nodes must try all supernodes.
- terryjsmith 16y agoSince it seems a good chunk of the super nodes went down, I imagine every running unconnected Skype instance is checking all of the super nodes it can find (or actively searching for them). Changing the logic at this point isn't really an option for them.
- tommi 16y agoWell, even if we were to talk about this point, it actually might be. I have no knowledge of Skype architecture, so I'm totally guessing here. But, those unconnected Skype instances have to have some kind of directory of supernodes: either dynamic or static. If it's dynamic, that is from Skype's servers, it can be affected.
- terryjsmith 16y agoThat's fair, but even if it is dynamic, I doubt it's checked on a regular basis; I doubt it's a scenario they've even considered before now. This would be on par with the DNS root servers changing and somehow letting everyone know about it on the fly.
- tommi 16y agoI tend to disagree. Error handling exactly in such cases where you have to bring the network online from a complete halt or other catastrophe is at the very heart of the architecture of these systems. These scenarios should and are at the minds of the architects and coders.
- terryjsmith 16y agoThat's a lot easier when it's your network. Skype is obviously designed to be used in "hostile" environments; super nodes are likely expected to go away and come back on a regular basis. Likewise, I don't think it's an unreasonable expectation that there will be a given number of nodes online at any given time. If they all disappear, I doubt there's a contingency plan for that.
- dedward 16y agoSupernodes are just guys like you and me with the right amount of bandwidth and the right network connection - if bad code or bad data exposed an existing bug, causing the supernodes in the meshed p2p network to go down - you have a chicken and egg problem. Skype doens't own the supernodes.... skype works because it uses the user's own resources to help route calls for other users.
- dennisgorelik 16y agoSo the bug is in Skype client code that is responsible for handling supernode behavior. The fix should probably be about developing deploying new version of Skype client. Such fix can take a week. Especially considering that engineering might be not in a good shape in Skype few years after acquisition from original founders.
- ddlatham 16y agoThere's an important lesson here for P2P system design. As in all systems, things will go wrong, whether it's your own bug or someone else's. Your system needs to not only be stable in the steady state, but to be able to return to the steady state when something interrupts it.
- terryjsmith 16y agoAbsolutely. We designed a P2P file system a few years ago and actually gleaned a good number of tricks from Skype for dealing with NATs and constructing your network in general. Dropbox (and many others I'm sure) have all said the same thing: you need to design your system to function in the most hostile conditions you can think of; for Skype this seems especially devastating because so many components are beyond their control.
- amackera 16y agoDo you know of any resources where I could learn about Skype's network structure? I'm very interested in distributed systems, especially ones with the scope of Skype's network.
- terryjsmith 16y agoSure, so there were a lot of them that we compiled, here are the ones and topics I remember. Protocol: we used a custom psuedo-TCP protocol built on UDP based on libjingle from Google (used in Google Talk): http://code.google.com/apis/talk/libjingle/file_share.html http://code.google.com/apis/talk/libjingle/file_share.html. Libjingle's filesharing itself was a decent resource for learning some P2P stuff as well. NATing and routing: we used Skype's UDP hole punching: http://www.h-online.com/security/features/How-Skype-Co-get-round-firewalls-747314.html http://www.h-online.com/security/features/How-Skype-Co-get-r.... Skype configures each install to be a supernode by default and then disables it based on certain criteria (low speeds, behind a NAT/firewall, etc.). Discovery of other network nodes is an entire subject unto itself and a lot of other network discovery protocols are well documented; our system used a central server to track up-time statistically and map clients to one another based on finding a good fit. Happy to answer any questions here or at my e-mail address in profile.