4 ms·
Not always that simple. 2-3 times a week is nothing! Try being on call in AWS, or any product/service at that scale. How often you get paged has less to do with
by ripper1138 4y ago
Not always that simple. 2-3 times a week is nothing! Try being on call in AWS, or any product/service at that scale. How often you get paged has less to do with your organization and more to do with the scale of your systems and business.
- dilyevsky 4y agoI’ve been a part of borg oncall at google - software that manages 90+% hardware there (and there are a lot of hardware). There were week long stretches without any pages. Dont ship garbage software and it’ll be alright at any scale.
- ripper1138 4y agoThanks for the anecdote. “it’ll be alright at any scale” is just naive.
- Tao3300 4y agoYeah, but what the hell is possibly important enough to wake up someone's family more than 2-3 times a week?
- michaelt 4y agoYou should ask your bosses to let you spend more time on bug fixing, because that's not normal, even at scale.
- morelisp 4y agoThe whole meaning of "scaling" is that you can do the same thing, but bigger. If your QoS is qualitatively different you've failed to actually scale your system. At best you've scaled a couple parts of it.
- SuperQue 4y agoNo, it really does have to do with organization priorities. You can make things reliable at scale with proper processes and automation.
- devonkim 4y agoErrors in a system are correlated with usage but a large part of our jobs as engineers is to reduce that correlation very, very hard. In organizations at even low scale I’ve had horrible levels of page outs (2-3 per night typical) but it means the system is unsustainable due to burning out workers in the end or that customers simply accept the error rates basically. At sufficiently high team size scales and error rates eventually you run out of hiring people to offset attrition which is what some people are reporting for teams at Amazon and AWS I’ve seen here and there.