9 ms·
When I was at Uber, we noticed that most incidents are directly caused by human actions that modify the state of the system. Therefore, a large "backlog" of hum
by exhaze 6y ago
When I was at Uber, we noticed that most incidents are directly caused by human actions that modify the state of the system. Therefore, a large "backlog" of human actions that modify the system state have a much higher chance of causing an incident.
My bet is that this incident is caused by a big release after a post-holiday "code freeze".
- nkassis 6y agoI would bet it's just the influx of traffic post holiday with systems that haven't been updated in so long maybe some annoying memory leaks have crept up and gone unnoticed or some other bad state that was exacerbated by return to work day for most NA folks. Code freezes were good at identifying bugs that only show up after long periods. Doubt anyone releasing big changes Monday morning.
- hnlmorg 6y agoThat might be true but when you take the global usage of Slack and their respective time zones, more than half the world would have signed into Slack this morning before SV had and I certainly didn't notice any downtime this morning in my time zone.
- radicalbyte 6y agoIt was ropey before SV woke up, I thought it was just my (normally rock solid thanks to using Ubiquity) network having issues. Guess it was Slack being Slack.
- exhaze 6y agoI haven't worked at Slack, so I can't speak with high confidence. A traffic spike is a possible reason, but I'm willing to bet that it's not the reason: > Doubt anyone releasing big changes Monday morning. This is definitely an engineering best practice, and by best practice, I mean something that Uber's, I mean Slack's SRE team strongly pushed for, and got politely overruled on. After a code freeze is lifted, it's quite common for lots of promotion-eager engineers to release big changes.
- agrippanux 6y agoIn my experience it's not promotion-eager engineers that want to push after a code freeze, it's antsy product managers. YMMV tho.
- glouwbug 6y agoWhat's there to change in Slack, though? It's arguably a messaging system, and that feature is tried and tested. That, and giphys, to be honest. EDIT: Guys it was a joke, chill
- VectorLock 6y agoHN's tolerance for jokes and sarcasm is extremely low.
- deleted 6y ago[deleted]
- brlewis 6y agoI'm not sure about that. I feel like I get more upvotes from sarcasm and jokes than from insight. In this instance, I think it's because when people hear something dumb said seriously in real life, they're not going to readily recognize online that it's a joke.
- Cederfjard 6y agoYeah, Poe’s law applies here. That’s definitely something someone less informed might say in earnest.
- mewpmewp2 6y agoYeah, there was other thread about Uber, where similar sentiment was seriously debated there, so I didn't recognise this as sarcasm either.
- bobthepanda 6y agoWhat would make that strange? Where I work it is frowned upon to do releases on weekends and so bad changes due to buildups happen on Monday. Although, we also don’t close the pipeline for just any holiday break. In fact low holiday traffic is a good time to keep pipelines open, since changes will impact less people.
- deleted 6y ago[deleted]
- alfalfasprout 6y agoThis is very likely a broken release. The timing lines up with pacific time too well.
- kevinmchugh 6y agoThey declared the issue at 7:14AM PST. How long is their deploy process? That sounds pretty early to think somebody on the west coast did something, other than maybe acknowledge the pages and declare the incident.
- NewEntryHN 6y agoSlack does progressive roll-outs. The broken release hypothesis seems very unlikely.
- deleted 6y ago[deleted]
- exhaze 6y agoTo elaborate a bit more on this point, you have to think about it like any complex system failure - it's almost never one thing, but rather a combination of many different factors. The factors around post NYE releases: - high risk changes that weren't released pre-holidays get released. Depending on the company, this could mean a 1-week to 1-month delay between implementation and release. The greater that interval, the higher the divergence world of production and the world of the new feature - lots of new hires (new year = new hiring budget). New hires are missing some tribal knowledge about the system and make a production-breaking release. I tried to think of other reasons, but these two overwhelmingly stand out as the two biggest reasons. Would love to hear from others.
- brundolf 6y agoSudden surge of traffic as all their users returns to work?
- ciceryadam 6y agoCould be, it's the perfect time overlap between US-West, US-East, and Europe.
- johnmaguire2013 6y agoYes - I wondered if they took some servers down prior to the break as a cost saving measure, and forgot to reinstate them.
- fragmede 6y agoDoubtful. It's not impossible a company the size of Slack would be reliant on a specific engineer logging on in the morning before a traffic spike so the service can handle the spike in load, but that's a misuse of modern distributed cloud-based computing. Hate on the cloud all you want, but AWS has (several flavors of) load balancers and various ways to automatically scale up and down resources (and if you're conservative, you can disable the 'down' part). If you're operating a major SaaS company like Slack and not taking advantage of them, something's gone wrong.
- rwc 6y agoSeems to be more than that. Even slack.com in an incognito browser fails.
- zwily 6y agoWhat does an incognito browser have to do with anything?
- SQueeeeeL 6y agoThat means it's not a user auth error
- johannes1234321 6y agoIt means that one is not sending a session cookie of any kind, thus should be sent to a 100% cached version. No "Are you XYZ and what to log into ABC's Slack again?" box.
- sbilstein 6y agonon-logged in user may not go through all the same codepaths as a user with cookies present.
- derin 6y agoAn incognito browser would ignore all client-side cookies, so the Slack web client would not try to - say - resume a previous user's session or re-use any previously saved data. Likewise, incognito mode will also ignore most cached web content, meaning all assets on the Slack web app will get loaded again from scratch. This "clean state" start could, theoretically, get around issues with old - potentially incorrect/outdated - assets being loaded, even though that really shouldn't happen under most circumstances.
- zwily 6y agoSure, but why does that indicate the issue probably wasn’t related to a code push, like the person I responded to said?
- cratermoon 6y agoI have definitely worked in places where the times right before and right after a change freeze were the most unstable, so that could be it. However, as others have mentioned, it's pretty early on the west coast of the US. Unless some engineer was up extra early (perhaps at the behest of an anxious project manager) it seems unlikely to be a release. What it could be is some engineer somewhere coming in after the holiday, noticing a slightly flaky thing, and thinking, "I'll reboot/redeploy/refresh this thing so the flakiness doesn't get worse". Only it turns out the flaky thing was a signal of something else falling over. Or maybe the redeploy was the wrong version because of bad CI/CD, or maybe the person just fat-fingered it.
- savo92 6y agoOr unless that engineer was not in the US
- cratermoon 6y agoVery possible. I don't know what Slack's workforce distribution is. In places I've worked there have definitely been some incidents in US off-hours triggered by someone on the other side of the world.
- ikiris 6y agoMost releases are automated with time lockouts.
- cratermoon 6y agoIn what companies?
- ikiris 6y agoCompetent ones like those you'd hear about being down on HN. At least that how it worked at one FAANG
- 6y ago
- ThePadawan 6y agoThis is one of the original concepts why to go capital-A Agile. Make smaller releases more often, so at least if something breaks, it's (hopefully) something small, and least it's easier to trace. (I'm not making a statement if that's good or bad or if it works or whatever. Please don't read an opinion into it.)
- erik_seaberg 6y agoThis. If you roll many changes into a single deployment, you don’t know which change broke what. But if you have two or three weeks of commits waiting, it’s hard to do otherwise.
- Cthulhu_ 6y agoThat's why good regression tests and CI are so important; in an ideal world (which we were close to in one of my projects), every change is pending in a pull request; the CI rebases the change on top of its upstream (e.g. master/main), simulating the state the codebase will be in once merged, and runs the full suite of tests. The build is invalidated and has to be re-run if either the branch or upstream is changed. Now, caveats etc, this was a collection of single applications in a big microservices architecture, and as the project grows it becomes more and more difficult to manage something like this, especially if you get more pull requests in the time it takes to do a build. But it is the way to go, I think. Anyway, since tests and CI are not definitive, you also need a gradual rollout - 1%, 5%, etc - AND you need a similar process for any infrastructure change, which gets more and more tricky as you go down to the hardware level.
- NewEntryHN 6y agoAnother common cause is resource exhaustion as a result of poorly monitored resources (or bugged monitoring). For example Google's authentication was down because their system reported (wrongly) available quota of 0. The last two incidents at my company were also related to resource exhaustion.
- alfiedotwtf 6y agoThis is why Change Management is the main tenant of principles like ITIL
- reportgunner 6y agoIf you think about it, modifications to state of the system caused by human actions are the sole purpose of computers.