4 ms·
Easiest day for engineers on-call everywhere except AWS staff. There’s nothing you can do except wait for AWS to come back online. Pour one out for the custome
by drevil-v2 1y ago
Easiest day for engineers on-call everywhere except AWS staff. There’s nothing you can do except wait for AWS to come back online.
Pour one out for the customer service teams of affected businesses instead
- codeduck 1y agoand by one I trust you mean a bottle.
- darkwater 1y agoWell, but tomorrow there will be CTOs asking for a contingency plan if AWS goes down, even if planning, preparing, executing and keeping it up to date as the infra evolves will cost more than the X hours of AWS outage. There are certainly organizations for which that cost is lower than the overall damage of services being down due to AWS fault, but tomorrow we will hear CTOs from smaller orgs as well.
- brazukadev 1y agoLots of NextJS CTOs are gonna need to think about it for the first time too
- noir_lord 1y agoThey’ll ask, in a week they’ll have other priorities and in a month they’ll have forgotten about it. This will hold until the next time AWS had a major outage, rinse and repeat.
- darkwater 1y agoIt's so true it hurts. If you are new in any infra/platform management position you will be scared as hell this week. Then you will just learn that feeling will just disappear by itself in a few days.
- noir_lord 1y agoYep, when I was a young programmer I lived in dread of an outage or worse been responsible for a serious bug in production, then I got to watch what happened when it happened to others (and that time I dropped the prod database at half past four on a Friday). When everything is some varying degree of broken at all times been responsible for a brief uptick in the background brokenness isn't the drama you think it is. It would be different if the systems I worked on where true life and death (ATC/Emergency Services etc) but in reality the blast radius from my fucking up somewhere is monetary and even at the biggest company I worked for constrained (while 100+K per hour from an outage sounds horrific - in reality the vast majority of that was made up when the service was back online, people still needed to order the thing in the end).
- swat535 1y agoThis applies to literally half of random "feature requests" and "tasks" that are urgent and needed to get done yesterday incoming from the business team..
- deleted 1y ago[deleted]
- mrits 1y agoHe will then give it to the CEO who says there is no budget for that
- ranger207 1y agoHonestly? "Nothing because all our vendors are on us-east-1 too"
- aswegs8 1y agoCan confirm, pretty chill we can blame our current issues on AWS.
- mvdtnz 1y agoNo really true for large systems. We are doing things like deploying mitigations to avoid scale-in (eg services not receiving traffic incorrectly autoscaling down), preparing services for the inevitable storm, managing various circuit breakers, changing service configurations to ease the flow of traffic through the system, etc. We currently have 64 engineers in our on-call room managing this. There's plenty of work to do.
- fisf 1y agoWell, some engineer somewhere made the recommendation to go with AWS, even tho it is more expensive than alternatives. That should raise some questions.
- array_key_first 1y agoEngineer maybe, executive swindled by sales team? Definitely.
- devjam 1y ago> Easiest day for engineers on-call everywhere I have three words for you: cascading systems failure