6 ms·
To elaborate a bit more on this point, you have to think about it like any complex system failure - it's almost never one thing, but rather a combination of man
by exhaze 6y ago
To elaborate a bit more on this point, you have to think about it like any complex system failure - it's almost never one thing, but rather a combination of many different factors. The factors around post NYE releases:
- high risk changes that weren't released pre-holidays get released. Depending on the company, this could mean a 1-week to 1-month delay between implementation and release. The greater that interval, the higher the divergence world of production and the world of the new feature
- lots of new hires (new year = new hiring budget). New hires are missing some tribal knowledge about the system and make a production-breaking release.
I tried to think of other reasons, but these two overwhelmingly stand out as the two biggest reasons. Would love to hear from others.
- brundolf 6y agoSudden surge of traffic as all their users returns to work?
- ciceryadam 6y agoCould be, it's the perfect time overlap between US-West, US-East, and Europe.
- johnmaguire2013 6y agoYes - I wondered if they took some servers down prior to the break as a cost saving measure, and forgot to reinstate them.
- fragmede 6y agoDoubtful. It's not impossible a company the size of Slack would be reliant on a specific engineer logging on in the morning before a traffic spike so the service can handle the spike in load, but that's a misuse of modern distributed cloud-based computing. Hate on the cloud all you want, but AWS has (several flavors of) load balancers and various ways to automatically scale up and down resources (and if you're conservative, you can disable the 'down' part). If you're operating a major SaaS company like Slack and not taking advantage of them, something's gone wrong.
- onefuncman 6y agoIt's easy to fall behind on bumping up the high watermark for your max autoscaling or for new traffic patterns to cause emergent instability. New code paths are taking unprecedented amounts of traffic all the time. In 2021, how does one keep track of resource starvation at the process, container, os, service, pod, cluster, availability zone and region levels?
- kevinmchugh 6y agoIf new hires tends to break production, it's not in the first business day of the calendar year. December gets really quiet for recruiting, typically, as candidates get busy with their social lives, and scheduling interviews gets harder. January is busy for recruiting, but given a week or two of interviewing and negotiating, two weeks notice, it's probably February before new employees are starting, and they're not making big, production-damaging deploys for a week or two after that.
- likpok 6y agoYou will also get a pause in new hires in late December for the same reason. I've certainly accepted an offer late in the year and then didn't start until the new year. Probably not as big of a rush as the end of school year rush in summer though. I also doubt that new people will be breaking production on day one. Even at a fast moving startup I'd expect it to take a bit to go through the onboarding paperwork, get a laptop and actually try pushing a change to production.
- sudhirj 6y agoI think some big company (maybe Facebook) has this rule that you had to deploy something to production on your first day. They seemed pretty confident in their processes and devops teams. A company trying to imitate that policy without doing the work necessary to make it possible would probably have outages on days when lots of new people joined :-P
- spiralx 6y agoCould be Facebook as I think production releases are always rolled out in phases e.g. first to 10 users, then 100, then 1000 and so on. That means there's much less chance of even the worst mistake having a serious effect.
- laci37 6y agoWow, onboarding new hires here is going good, if they can access slack, O365, LDAP, VPN and clone the repo by the end of the first day. Tho we have the initiation ritual of installing the OS to your laptop.
- adrianpike 6y agoI think you're right on the first bullet, but not the second. If it was mid-Feb, then maybe, but the next FY hasn't even started yet for a ton of companies, let alone onboarding newbies to production.
- lwedel 6y agoI would add here the potential scaling issue - holidays were a dry season - less meeting. So if they have some automation for scaling down to reduce cost, it may have bitten them in their arses now. People came back to work, and most of them start around the same time (US wise at least). Hence kids - a vital lesson for all of us - don't start the call at a full hour, give it 3-7 min to make your coworkers confused and give some time for the systems to auto-scale ;)
- binaryblitz 6y agoI hope this isn't the case. It's not like this is the first holiday season for slack.
- darkerside 6y agoPeople returning to work and downloading a huge backlog of messages from the past two weeks.
- pgAdmin4 6y agoAs an Microsoft Teams Ex-Dev, I can vouch that message retrieval after vacation puts a lot of stress on storage systems before it stabilizes :)
- darkerside 6y agoYeah, makes sense. A system typically optimized for performance and real time delivery is suddenly asked to perform multiple batch retrievals in large chunks. Ouch!