3 ms·
I’m not intelligent in the huge scale distributed systems stuff but if there are outages every couple of weeks, why would you push code to production several ti
by blntechie 6y ago
I’m not intelligent in the huge scale distributed systems stuff but if there are outages every couple of weeks, why would you push code to production several times each day?
I honestly don’t know what’s in Slack which requires pushing code to production several times every day. I completely understand having that capability is required for any serious engineering org but what kind of churn is that? Maybe someone knowledgeable can help me understand.
- AznHisoka 6y agoOutages are not always caused by bugs you deploy to production. They can be due to capacity issues with too many users at once using the platform.
- blntechie 6y agoYes, I get that but wouldn’t launching new features be adding pressure during an outage which complicates the things? Capacity issue also could have been caused by a new feature probably acting greedy on the resources too.
- Groxx 6y agoWhile true in principle... those kinds of causes tend to have pretty clear cause-effect (or at least strong time correlations), and get rolled back. New call patterns / major load changes are generally pretty obvious and at least coarsely traceable if you have monitoring anywhere.
- rhizome 6y agoIf that's the cause of biweekly outages, shouldn't they be building or acquiring more capacity? I realize complexity plays a role in the rate at which capacity can be added to the system, but every two weeks is very frequent for that kind of problem. "Too many users" only counts when it's a surprise.