23 ms·
Slack’s Outage on January 4th 2021
- fullstop 6y agoI was kind of surprised to see that they are using Apache's threaded workers and not nginx.
- dijit 6y agoApache mod_php is still much faster than php-fpm, and since slack uses a lot of PHP on the backend it makes a lot of sense for them.
- saurik 6y agoI loved this article where someone started with the goal of writing an article about how much faster nginx was, but then discovered the opposite (and for some reason didn't change the title of the article, which is hilarious)... both because it showed the author "cared", but also because it showed that people just assume what amounts to marketing myths (such as that Apache and mod_php are ancient tech vs. the more modern php-fpm stack) before bothering to verify anything. https://www.eschrade.com/page/why-is-fastcgi-w-nginx-so-much-faster-than-apache-w-mod_php/ https://www.eschrade.com/page/why-is-fastcgi-w-nginx-so-much...
- fullstop 6y agoWhen I moved things from Apache -> nginx years ago, I did it not because it was faster but because the resource requirements of nginx were so much more predictable under load.
- Denvercoder9 6y agoThat's not the reason, as mod_php only supports the prefork MPM, and not the threaded MPMs.
- johannes1234321 6y agoThreaded MPMs and PHP work quite well, I fixed a few bugs in that space some time ago. There however might be issues in some of the millions libraries PHP potentially links in an called from it's extensions, and those sometimes at emit thread safe, but finding that and finding bypasses isn't easy ... The other issue is that it's often slower (while there recently were changes from custom thread local storage to more modern one) and if there is a crash (i.e. Recursion stack overflow ...) it affects all requests in that process, not only the one.
- cholmon 6y agoAre you saying that Slack uses mod_php? According to a Slack Engineering blog post[0] from 9 months ago, Slack has been using Hack/HHVM since 2016 in place of PHP. My understanding is that HHVM can only be run over FastCGI, unless there's a mod_hhvm that I'm unaware of. [0]: https://slack.engineering/hacklang-at-slack-a-better-php/ https://slack.engineering/hacklang-at-slack-a-better-php/
- kmavm 6y agoThis is accurate: Slack is exclusively using Hack/HHVM for its application servers. HHVM has an embedded web server (the folly project's Proxygen), and can directly terminate HTTP/HTTPS itself. Facebook uses it in this way. If you want to bring your own webserver, though, FastCGI is the most practical way to do so with HHVM.
- user5994461 6y agomod_xyz plugins were deprecated long ago because they are unstable. If you do a comparison you should compare fastcgi versus fastcgi, that is the standard way to run web applications. Running with mod should be faster because it's running the interpreter directly into the apache process but it's also making apache unstable. mod_python was abandoned around a decade ago. It's crashing on python 2.7. mod_perl was dropped in 2012 with the release of apache 2.4. It was kicked out of the project but continues to exist as a separate project (not sure if it works at all).
- nickthemagicman 6y agoWow. The trail leads back to AWS. Wasn't there a number of other companies that were down around that same time or was that a different time?
- jjtheblunt 6y agoIs this the recent event you refer to? https://aws.amazon.com/message/11201/ https://aws.amazon.com/message/11201/
- nickthemagicman 6y agoAh yes, that was the one I was referring to. Looks like a different event.
- conradfr 6y agoDoes AWS compensate you in cases like this?
- dewey 6y agoIt depends: https://aws.amazon.com/legal/service-level-agreements/ https://aws.amazon.com/legal/service-level-agreements/
- dandigangi 6y agoThey do. I can't remember where the documentation is but there's a clause where they payout as long as they aren't meeting their SLAs. Who knows the different ways they may be able to get out of that. I assume this wasn't one of those times.
- sargun 6y agoI wonder why Slack uses TGW instead of VPC peering.
- tikkabhuna 6y agoThe "I've just done my Solutions Architect exam" answer would be that TGW simplifies the topology by having a central hub, rather than each VPC having to peer with all the other VPCs. I wonder how many VPCs people have before transitioning over to TGW.
- saurik 6y agoI miss EC2 Classic :/. It always feels like the entire world of VPCs must have come from the armies of network engineers who felt like if the world didn't support all of the complexity they had designed to fix a problem EC2 no longer had--the tyranny of cables and hubs and devices acting as routers--that maybe they would be out of a job or something, and so rather than design hierarchical security groups Amazon just brought back in every feature of network administration I had been happily prepared to never have to think about every again :(.
- BillinghamJ 6y agoGenerally inclined to agree, but to be fair you can operate a VPC in exactly the same way as EC2 Classic - give everything public IPs, public subnets and ignore the internal IPs. Pretty sure those are the defaults too
- whatisthiseven 6y agoAgreed, I always thought VPC and all that complexity was a big step backwards. My org is moving from a largely managed network into AWS, and now we have to configure the whole network and external gateways ourselves? What engineer wants to do this? VPCs are virtual, but I don't need VPCs, I need the entire network layer virtualized and abstracted. As you suggested,just grouping devices in a single network and saying "let them all talk to each other, let this one talk to that one over this port/IP" should be all I describe. Let AWS figure out CIDR, routing, gateways, etc.
- dplgk 6y ago> On January 4th, one of our Transit Gateways became overloaded. The TGWs are managed by AWS and are intended to scale transparently to us. However, Slack’s annual traffic pattern is a little unusual: Traffic is lower over the holidays, as everyone disconnects from work (good job on the work-life balance, Slack users!). On the first Monday back, client caches are cold and clients pull down more data than usual on their first connection to Slack. We go from our quietest time of the whole year to one of our biggest days quite literally overnight. What's interesting is that when this happened, some HN comments suggested it was the return from holiday traffic that caused it. Others said, "nah, don't you think they know how to handle that by now?" Turns out occam's razor applied here. The simplest answer was the correct one. Return-from-holiday traffic.
- cett 6y agoThough the nuance is Slack did know how to handle it, AWS didn't.
- tempest_ 6y agoWell Slack depended on the Cloud(tm). It is a interesting though because a lot of the blog posts like "How we handled a 3000% traffic increase overnight!" boil down to "We turned up the AWS knob". What happens when the AWS knob doesn't work?
- buildawesome 6y agoYou do what Slack did and call the maker of the AWS Knob:tm:.
- coldcode 6y agoI actually thought something in AWS was a cause but did not know about these internal systems.
- coldcode 6y agoI actually thought something in AWS was a cause but did not know anything about how TGW works internally.
- bovermyer 6y agoI really enjoy reading these write-ups, even if the causal incident is not something I enjoy.
- ignoramous 6y agoNot at all fun when you're actively involved in mitigating these, though. Pretty rough and sometimes scars you for life.
- cle 6y agoI work on critical services, similar to this. While nobody likes the overall business impact of events like this, I love being in the trenches when things like this happen. I enjoy the pressure of having to use anything I can to mitigate as quickly as possible. I used to work in professional kitchens before software, and it feels a lot like the pressure of a really busy night as a line cook. Some people love it.
- deleted 6y ago[deleted]
- bovermyer 6y agoI've been in the middle of events like these and had to write my share of postmortem documents. Hearing about others' similar experiences makes me feel a connection to them, and often teaches me something.
- gscho 6y ago> our dashboarding and alerting service became unavailable. Sounds like the monitoring system needs a monitoring system.
- TonyTrapp 6y ago"Who monitors the monitors?"
- jrockway 6y agoIt is quite awkward that the output of "working" and "completely broken" alerting systems have the same visible effect -- no alerts. For Prometheus users, I wrote alertmanager-status to let a third-party "website up?" monitoring server check your alertmanager: https://github.com/jrockway/alertmanager-status https://github.com/jrockway/alertmanager-status (I also wrote one of the main Google Fiber monitoring systems back when I was at Google. We spent quite a bit of time on monitoring monitoring, because whenever there was an actual incident people would ask us "is this real, or just the monitoring system being down?" Previous monitoring systems were flaky so people were kind of conditioned to ignore the improved system -- so we had to have a lot of dashboards to show them that there was really an ongoing issue.)
- ianrw 6y agoI wonder if you can pre-warm TGWs like you can ELB? It would be annoying to have to have AWS prewarm a bunch of you stuff, but it's better than it going down.
- Thaxll 6y agoSurprised that there is just a few metrics available for TGW: https://docs.aws.amazon.com/vpc/latest/tgw/transit-gateway-cloudwatch-metrics.html https://docs.aws.amazon.com/vpc/latest/tgw/transit-gateway-c... Probably the only way to see a problem is if you have a flat line for bandwidth, but as the article suggested they had packet drop wich does not appear on the cloudwatch metrics, aws should add those metrics imo
- miyuru 6y agosome TGW limits are documented here: https://aws.amazon.com/transit-gateway/faqs/#Performance_and_limits https://aws.amazon.com/transit-gateway/faqs/#Performance_and...
- kparaju 6y agoSome lessons I took from this retro: - Disable autoscaling if appropriate during outage. For example if the web server is degraded, it's probably best to make sure that the backends don't autoscale down. - Panic mode in Envoy is amazing! - Ability to quickly scale your services is important, but that metric should also take into account how quickly the underlying infrastructure can scale. Your pods could spin up in 15 seconds but k8s nodes will not!
- tobobo 6y agoThe big takeaway for me here is that this “provisioning service” had enough internal dependencies that they couldn’t bring up new nodes. Seems like the worst thing possible during a big traffic spike.
- throwdbaaway 6y ago> We run a service, aptly named ‘provision-service’, which does exactly what it says on the tin. It is responsible for configuring and testing new instances, and performing various infrastructural housekeeping tasks. Provision-service needs to talk to other internal Slack systems and to some AWS APIs. The "configuring and testing new instances" part also sounds very fishy to me. Configuration should be done when creating the image and launch template, while testing should be the job of the load balancing layer. Why do we need a separate "provision-service" to piece everything together?
- jeffrallen 6y agotldr: "we added complexity into our system to make it safer and that complexity blew us up. Also: we cannot scale up our system without everything working well, so we couldn't fix our stuff. Also: we were flying blind, probably because of the failing complexity that was supposed to protect us." I am really not impressed... with the state of IT. I could not have done better, but isn't it too bad that we've built these towers of sand that keep knocking each other over?
- bobthebuilders 6y agoThe thing is anyone can build a system that can scale out to Slack's level given enough machines and money. What's harder is scaling out to that level and not burning gobs of cash. It's similar to the whole buildings a long time ago last much longer than those today. Its true in the literal sense, but it ignores the fact that we've gotten at reducing the cost of stuff like skyscrapers and bridges. In our pursuit of efficiency, we do things like JIT delivery, dropshipping, scaling, building to the minimum spec. Sometimes, we get it wrong and it comes tumbling down (covid, HN hug of death, earthquakes).
- danw1979 6y agoSo many fails due to in-band control and monitoring are laid bare, followed by this absolute chestnut - > We’ve also set ourselves a reminder (a Slack reminder, of course) to request a preemptive upscaling of our TGWs at the end of the next holiday season.
- jrockway 6y agoThe thing that always worries me about cloud systems are the hidden dependencies in your cloud provider that work until they don't. They typically don't output logs and metrics, so you have no choice to pray that someone looks at your support ticket and clicks their internal system's "fix it for this customer" button. I'll also say that I'm interested in ubiquitous mTLS so that you don't have to isolate teams with VPCs and opaque proxies. I don't think we have widely-available technology around yet that eliminates the need for what Slack seems to have here, but trusting the network has always seemed like a bad idea to me, and this shows how a workaround can go wrong. (Of course, to avoid issues like the confused deputy problem, which Slack suffered from, you need some service to issue certs to applications as they scale up that will be accepted by services that it is allowed to talk to and rejected by all other services. In that case, this postmortem would have said "we scaled up our web frontends, but the service that issues them certificates to talk to the backend exploded in a big ball of fire, so we were down." Ya just can't win ;)
- bengale 6y agoExperienced something similar with mongo atlas today. Our primary node went down and the cluster didn’t failover to either of the secondaries. We got to sit with our production environment completely offline while staring at two completely functional nodes that we had no ability to use. Even when we managed to get hold of support they also seemed unable to trigger a failover and basically told us to wait for the primary node to come back up. It took 90 minutes in the end and has definitely made us rethink about the future and the control we’ve given over.
- deleted 6y ago[deleted]
- gen220 6y agoSome customers have a hard requirement that their slack instances be behind a unique VPC. Other customers are easier to sell to if you sprinkle some “you’ll get your own closed network” on top of the offer, if security is something they’ve been burned by in the past. I agree with you the mTLS is the future. It exists within many companies internally (as a VPC alternative!) and works great. There’s some problems around the certificate issuer being a central point of failure, but these are known problems with well-understood solutions. I think there’s mostly a non-technical barrier to be overcome here, where the non-technical executives need to understand that closed network != better security. mTLS’s time in the sun will only come when the aforementioned sales pitch is less effective (or even counterproductive!) for Enterprise Inc., I think.
- johnnymonster 6y agoTL;DR outage caused by traffic spike from people returning to work after the holiday.
- jeffbee 6y agoWhy don't they mention what seems like a clear lesson: control traffic has to be prioritized using IP DSCP bits, or else your control systems can't recover from widespread frame drop events. Does AWS TGW not support DSCP?
- nhoughto 6y agoNever seen Transit Gateway, I assume they wouldn't have this problem if it was just single VPC or done via VPC peering?
- keyle 6y agoDidn't we just read a story about the exact same issue? Traffic picked up heavily on some website or app, AWS didn't auto-scale fast enough or at all and the very systems that are designed to be elastic just tumbled down to a grinding halt?
- keyle 6y agoUpdate: it was Advent Of Code 2020, where he reported the exact same issue. The AWS auto scaling framework rocked a pooper when the site exploded on release day.
- plaidfuji 6y agoMaybe I’m being naive, but what I’m curious about is: how did their whole team communicate through this triage? I assume not Slack?
- allannienhuis 6y agoour emergency backup for slack is zoom. Horrible UX for group chats, but everyone already has it installed and it's quick and simple to set up a new room for each team. For temporary use you can put up with a fair bit of annoying behaviour or lack of features.
- grumple 6y agoAutomated scaling has been a persistent problem for me, especially if I try to scale on simple metrics, or even worse (in Slack's case) on metrics that could potentially compete. The situations in which multiple metrics could compete are sometimes difficult to conceive, but it will always happen if you aren't performing something more sophisticated than "up if metric > value, down if < value" for multiple metrics. I think you've got to combine these somehow into a custom metric and scale just on one metric. I'm totally unsurprised to see that autoscaling failed for both Slack and AWS in this case. I think you really have to look at metric-based autoscaling and say: is it worth the X% savings per month? Or would I rather avoid the occasional severe headaches caused by autoscaling messing up my day? Obviously this depends on company scale and how much your load varies. I'd rather have an excess of capacity than any impact on users.