9 ms·
Signal apps DDoS'ed their own server
- dboreham 6y agoThermal runaway. Easily done and I suppose hard to see as a potential problem until you do it to yourself.
- nullify88 6y agoExplains why DNS was resolving to 127.0.0.1 during the outage. Just didn't think it could be caused by their own app.
- alexandrerond 6y agoTo keep clients from DDoSing the server. Probably they set a % of requests to resolve to localhost and work with less load. It's a pretty good kill switch.
- hyuen 6y agoFirst rule of network programming: conservative sending, liberal receiving
- ncmncm 6y agoThat has turned out to be a really terrible rule. It makes it very hard, sometimes impossible to upgrade a protocol, because some peer somewhere will try to interpret your new thing if it were a broken old thing; so security holes get enshrined. It is much better to fail fast, and make it easy to see why. The only people who benefit from enabling crap implementations, in the end, are makers of crap implementations. Their stuff should fail as soon as possible, so they have to fix it. In some cases, it is acceptable have a "permit crap" mode that can be turned on in cases where the other end is not negotiable. But ideally the crap fails to work the first time it is plugged in, and returned to the manufacturer. Stuff that causes no end of trouble gets replaced eventually, but somebody has to expose the trouble, and it is best to have lots of company so the culprit is blamed.
- samoa42 6y ago> That has turned out to be a really terrible rule. maybe in the ivory tower of web shit, but it is essential for internet/society. like you said > in cases where the other end is not negotiable. aiming for robustness is key if you want to enable a diverse ecosystem.
- tpirx 6y agoIs it common for linear growth to cause these sorts of issues? Any postmortems to share from similar app failures?
- hayst4ck 6y agoWhat are you assuming is linear growth?
- Exuma 6y agoWhat's up with Elon's tweet about Signal... wtf?
- approxim8ion 6y agoHe's just trying to be part of the conversation in a "how do you do, fellow kids" way.
- lrossi 6y agoIt’s easy to judge, but distributed systems are hard to program correctly, and hard to test. One would need to run failure injection tests at scale to detect this. This is made worse by mobile device virtualization being difficult to achieve as well. Is there a system I could use to spin up 10k emulated iPhones, to run a test in a CI job, for example?
- jackcosgrove 6y agoIt is easy to judge. At an old job of mine we had devices we controlled that could be rebooted. The initialization sequence after a reboot is much more network heavy than normal operation. We tried to push a critical security patch to all devices at once, meaning they all rebooted at once. Whoops self-DDOS.
- ffggvv 6y agothat’s not a ddos. it’s just called not being able to handle a spike in traffic.
- uncledave 6y agoUSDOS - Unexpected Success Denial Of Service. I woke up one Monday morning to crashed servers and angry phone calls once in that situation. Unfortunately it wasn't success on my part but what I had done the Friday before was so crap it couldn't handle a hundred users. But we blamed it loudly on unexpected success, quickly patched the mistake and everyone was happy. YMMV :)
- nitrogen 6y agoLesson learned: never ship on Friday
- Daniel_sk 6y agoThe application has a retry mechanism that will keep trying until a connection succeeds (with an exponential backoff), but it doesn't handle a case where the server is basically rejecting the messages on purpose (due to capacity issues). Suddenly all applications will start retrying the connection at once and there is no way to turn if off. So it will make the problem even much worse than it already is. They have added a server feature flag yesterday to change the maximum backoff time and they also added the handling of HTTP 508 error response.
- williamdclt 6y agoI'm wondering if another way to handle that would be the server accepting the request, but not actually servicing it. Basically let it timeout so that the client doesn't retry (at least, does not retry before timeout)
- Daniel_sk 6y agoCommits from yesterday: Feature flag for maximum retry backoff time: https://github.com/signalapp/Signal-Android/commit/93e9dd6425b22db13baaf7e7f780a93bda0a457e https://github.com/signalapp/Signal-Android/commit/93e9dd642... Add jitter to backoff time: https://github.com/signalapp/Signal-Android/commit/8f7fe5c3eeb693e132b3c7d8bc692546bd70d27d https://github.com/signalapp/Signal-Android/commit/8f7fe5c3e... Handle ServerRejectedException (HTTP 508): https://github.com/signalapp/Signal-Android/commit/c95f0fce6ee3b78ff82fde865b2ee49288e1303f https://github.com/signalapp/Signal-Android/commit/c95f0fce6... Feature flag for automatic session reset: https://github.com/signalapp/Signal-Android/commit/a3c7e7e552f35751c43e79481357980bb36a8404 https://github.com/signalapp/Signal-Android/commit/a3c7e7e55...
- LurkersWillLurk 6y agoUnfortunately, I believe this is correct. I received over 1,000 messages from one contact of mine that had the message of "secure session reset". It seems his phone tried to reset the encrypted connection with me over and over. Considering that I have 3 devices in total, that's thousands of messages just from one user alone. I'm sure millions of devices doing the same thing probably bogged them down.
- spurgu 6y agoI had this with one contact as well, the message on Android was Bad encrypted message and on desktop it was Error handling incoming message.
- rvz 6y agoThe general case to use Erlang for your reliability and scalability issues. Plenty of suggestions and time to switch. [0] Now after the fact of its popularity from the source of it all [1]. My response: Use Erlang. Downvoters: So for a chat app that is still down for several hours and needs to be reliable to handle the scale of new users, what should they (Signal) have used instead? Java? [0] https://mobile.twitter.com/ejpcmac/status/977192665489035265 https://mobile.twitter.com/ejpcmac/status/977192665489035265 [1] https://mobile.twitter.com/anildigital/status/1350274900737585152 https://mobile.twitter.com/anildigital/status/13502749007375...
- vips7L 6y agoHow does erlang fix the problem of not having enough physical servers?
- rvz 6y agoFirst of all, they’re on AWS. They don’t ‘have’ physical servers. Secondly, the point of Erlang is that you do not need that many physical/virtual servers to scale to these millions of users. Perhaps just a few or the same number of servers they have, but running Erlang instead. The BEAM VM and languages based on this is quite frankly designed for this scale. Still working well for WhatsApp and Discord.
- lsllc 6y agoI think what GP is referring to is that WhatsApp's Erlang based system was able to handle 2M connections per server allowing you to do more with less [0] (although I think I read since the FB acquisition, WhatsApp has moved off from Erlang). [0] http://highscalability.com/blog/2014/2/26/the-whatsapp-architecture-facebook-bought-for-19-billion.html http://highscalability.com/blog/2014/2/26/the-whatsapp-archi...
- darksaints 6y agoThere are lots of languages that can do this, and it requires deliberate engineering effort in all of them, including erlang. I mean, you have to have all of your ports open, with multiple IP addresses allocated to a single server, and the process has to have efficient routing algorithms to do so. Erlang doesn't make this job more efficient than other languages. Erlang has two killer use cases: fault tolerance and horizontal scalability. Computational resource efficiency has never been a goal and it shows. The grandparent comment is nothing more than a cargo cult. "If we use the same language as WhatsApp, we'll magically be just as scalable as WhatsApp". It's horrendously naive.
- Vinnl 6y agoThey're asking people to share their debug logs to help diagnose this potential issue: https://community.signalusers.org/t/help-needed-please-send-me-your-android-debug-logs/23266 https://community.signalusers.org/t/help-needed-please-send-...
- Bucephalus355 6y agoFYI the debug logs are a very curious part of all this. Signal streams them usually to a domain they keep their ownership of somewhat hidden. The domain is: debuglogs.org and the endpoint is just api.debuglogs.org. It appears to ultimately just front for AWS S3 backend so a very common architectural pattern.
- zamadatix 6y agoI'm not sure I understood what the curious part about the debug logs was.
- faitswulff 6y agoCame here to post this. Extra important since they don't collect analytics!
- maxpert 6y agoIDK why mobile dev folks do blind retries thinking server is some kind of mage. I’ve had debates on how much retry makes sense for a login service I built. People have the tendency of hey this endpoint used to work with these parameters didn’t work? Fine I will keep retrying! The only difference here I think is in company paid job you step up immediately put in a server side fix to fix outage and then make these changes to roll them out. All of this is achieved in matter of hour or so rather than 12+ hours. Circuit breaking, exponential backoff, bulkheading should be key pillars and no dev should be allowed to write a client without these key skills. Edit: Please add Jitter to the list of items. Any good library like https://resilience4j.readme.io/ https://resilience4j.readme.io/ will give you all of this for free out of box!
- tyingq 6y agoHas the "batteries included" aspect of things like exponential backoff improved? Last I was in the details, most of the popular client libraries offered little, and you had to implement it yourself.
- maxpert 6y agoResilience4J has improved quite a lot. I have been using for over two years now, never missed out on anything.
- abhi_kr 6y agoAnother scenario to keep in mind is the Thundering Herd Problem[0]. Exponential backoff without added jitter could still DDoS the servers. [0]: https://en.wikipedia.org/wiki/Thundering_herd_problem#Mitigation https://en.wikipedia.org/wiki/Thundering_herd_problem#Mitiga...
- VWWHFSfQ 6y agoexponential backoff will still kill the servers because the first thing people do is kill/restart the app or reload the webpage and it will just restart the backoff again
- mwcampbell 6y agoHaving gotten client-side retry logic wrong before, while also being the sysadmin responsible for keeping the server side up (back in the days of using a single dedicated server), I'm sympathetic to the Signal staff right now. We should go easy on them.
- deleted 6y ago[deleted]
- Daniel_sk 6y agoI have done the same mistake with my own app a few years ago. Fortunately I had a spend limit set on Google App Engine and it was set quite low, it could have bankrupted me :-). I believe that not many mobile developers actually think about this "retry hug" problem until it bites them. It's easy to just write few lines of automatic retries and think that the server will be fine.
- cloudking 6y agoThis is why you need to implement exponential backoff on retries for anything at scale https://en.wikipedia.org/wiki/Exponential_backoff https://en.wikipedia.org/wiki/Exponential_backoff
- hnarn 6y agoIs this what's commonly seen as "Retrying connection in:" with a number of seconds increasing like 5 -> 10 -> 20 -> 40 -> 80 etc?
- whycombagator 6y agoCould be. If implementing exponential backoff it’s always a good idea to add some jitter (randomness) to the backoff, otherwise you can end up with multiple processes retrying all at the same time, all backing off again, and then all retrying at the same time etc etc
- manojlds 6y agoCalled Thundering Herd problem.
- fitblipper 6y agoThis is why you need some form of network analytics to figure out if the new version you are rolling out is breaking anything or even seeing of something _is_ breaking. I get the "we don't want to compromise security even an inch", but how secure is a messaging app that cannot send messages? Even more critical since anyone can write clients that function with the server and thus a malicious actor could attach them without them knowing how to identify it or stop it.
- zamadatix 6y agoI'm not sure it had anything to do with "a new version rolling out", the issue has always been there it just never had a trigger (massive enough load) up until this point. I suppose the app that can't send messages is the most secure of all! All joking aside if their focus is security over anything else then this is acceptable. Unfortunately I don't think most of the users have the same order of priorities so it's bound to create some tension over time in multiple ways such as testing analytics, features that don't get added, or just general friction due to always requiring "most secure". The last part I agree is probably the most concerning though, it seems like Signal's centralized services aren't ready to be battle tested from attacks bound to come due to it's popularity growth. I'm not sure matrix is the silver bullet to that problem either... email isn't resilient to DoS because it's distributed it's resilient because the key centralized players like Gmail can withstand constant attack without affecting service due to their scale. Neither Signal nor Matrix are ready for the attention that comes with serving billions from that perspective. It's something I think will come with growing pains though.
- tyingq 6y agoI remember a problem like this I ran into with a website where a short outage occured, then an avalanche of demand as end users reloaded or retried the page. We settled on throttling demand based on the incoming IP at the firewall...dropping incoming http packets for 3/4 of the addr space, then later 1/2, then 1/4, etc. That was long ago, though, long enough that it was an NetScape webserver.
- bilal4hmed 6y agoSo does this mean every message I sent till the update isn't rolled out to the client is still DDoSing them ?
- nimbius 6y ago>I think the @signalapp apps DDoS'ed the server. and so the chickens come home to roost. Moxies vehement rejection of a distributed design seems less and less tenable each outage. Last time it was what...verification numbers that werent getting sent? and FWIW the divinations from the community are exceedingly helpful in a time when not even the signal website seems to confirm or deny any sort of outage. Signups are still being taken and the outage page at signal.org is still as static as ever. Even the twitter hasnt seen an update in nearly a day. Is anyone at Signal foundation at the controls? As a signal user myself I know this is going to sound rude but at this point other than the endorsements from musk and dorsey, why would any new user consider this service at all if its been down for nearly two days? theres never any postmortem, and communication is generally evangelical or solicitous in nature for either installs or donations.
- rglullis 6y agoIn some ways, it seems like we are watching a warp-speed demonstration of how evolutionary processes work. Environmental pressure, lots of contenders coming with slightly mutations and all of them passing through some fitness filter. It's just too bad that those arguing for decentralized systems fail to spread the meme more effectively. We federalists are like cockroaches.
- LurkersWillLurk 6y agoRespectfully, how exactly would federation have helped with this outage? Signal's own official clients failed to properly back off from spamming the server with incessant requests. I'm not entirely sure how more third party clients would have helped with this issue.
- bilal4hmed 6y agoMy question exactly, if majority of the users have an account on the @matrix instance and if that fails, wont the same issue happen ?
- rglullis 6y ago
- fareesh 6y agoI've noticed that on Android, WhatsApp and Gmail sync their data on a periodic basis (probably using WorkManager or something like that). What is an efficient architecture for something like this? Are millions of devices hitting some gateway->load balancer everytime the Workmanager wakes up and polling for whether new messages have arrived? For example, if my phone does a "sync" 3 minutes from now, as do millions of others around the world, do we all get routed to our respective box that's keeping a queue of our pending messages in memory? If we happen to receive these messages in between the poll intervals via persistent socket or push notification, then the system wipes the in-memory queue. Am I imagining this correctly or is there a more optimal architecture for something like this?