11 ms·
> Code quality was atrocious. We had one enormous Java method (>1000 lines) which would take care of nearly every single request coming into our service... With
by middleclick 5y ago
> Code quality was atrocious. We had one enormous Java method (>1000 lines) which would take care of nearly every single request coming into our service... With only about 7-8 unit tests
AWS is pretty reliable for the most part so I am pretty surprise that the code quality is that bad.
- colde 5y agoI think that highly depends on the service. The new App Runner service for instance is a wild ride of buggyness, lack of testing and incorrect documentation.
- salil999 5y agoI think one thing I learned from AWS is that there's so much hidden away from the customer. There definitely were (and probably still are) many issues which the customers won't actively experience. Reliability doesn't necessarily equate to good standards and good practice. But yes, from a customer point of view, AWS is pretty nice.
- cpach 5y agoReminds me of that old quote by John Godfrey Saxe: ”Laws, like sausages, cease to inspire respect in proportion as we know how they are made.”
- markus_zhang 5y ago"I just had one for breakfast." -- Jim Hacker
- wpietri 5y ago> AWS is pretty reliable for the most part so I am pretty surprise that the code quality is that bad. I'm not totally surprised because of two factors: very stable product definitions and lots and lots of users. A number of years back, I was talking with people at a famous and popular site with a broad audience. I asked them how much unit testing they did. They said that particular isolated pieces sometimes had tests. But most of the user-facing stuff didn't because they had one-button rollout and one-button rollback. Instead of bothering with unit tests, they'd just frequently release changes, watch the metrics and the customer support queue, and quickly roll back if they'd introduced a bug.
- zorked 5y agoFor very, very popular services, a second of being live will exercise more code paths and edge cases than even the most dedicated testing team could ever dream of. We hear a hell of a lot about testing but the most fundamental piece of software quality nowadays is the release strategy: running on tee'd live production traffic, canarying, metrics and alerting, quick roll backs, etc.
- bradleyjg 5y agoDepends on if you are serving ads over cat pictures or routing air traffic. Different solutions for different problems.
- sreque 5y agoIt's a very short-sighted view on testing, although I'm not surprised SREs would say it. The biggest problem with software deployment is that it is owned and managed by people who have no vested interest in developer productivity, including devops engineers. A major goal of any org should be developer productivity; otherwise you are just hemorrhaging money and talent. When I say developer productivity, I mean: How confidently and quickly can I make a shippable, rollback-free change to a unit of software? If you are the dos equis man of testing, "I don't always test my code, but when I do, I do it in production", then you can't confidently make any change without risking a production outage, so you play lots of games, like you mentioned, around canarying, rolling out to a small percentage of users, etc., but at the end of the day your developer productivity has absolutely tanked. The goal of any system maintenance should be that a developer can quickly make and test a change locally and be highly confident that the change is correct. The canarying, phased rollouts, and other such systems should not be the primary means of testing code correctness.
- wpietri 5y agoYeah, I really appreciate excellent rollout strategies, although I suspect a lot of them are more developed out of self defense by SRE teams. I see it as a series of safety nets: I'm still going to write tests for my code so that I don't have far to fall if I make a mistake. But I also want a safe rollout so if I miss the first net I don't splatter on the pavement. And I totally agree with out about developer productivity. It's just not a consideration in most places. For example, in a factory or a restaurant, meetings are things that happen rarely and in constrained time slots, because everybody realizes that production is primary. But in most software companies, actually getting work done is second priority to meetings.
- bradleyjg 5y ago> AWS is pretty reliable for the most part In telecom or traditional mainframes, for example, the compute unit itself was expected to be reliable. Individual elements of AWS are not pretty reliable in that context. Check out the single host EC2 SLA. However, today most large or even medium scale software assumes unreliable individual elements and has redundancy at the program level. For that purpose, AWS core services are pretty reliable.
- oceanplexian 5y agoWay back I used to work in telecommunications at a place that provided POTS service. They are two very completely different worlds. Software engineers act as if 5 9's is a badge of honor, when really it isn't. When you are responsible for something that people use to dial 911 and can make the difference between life and death a few minutes of downtime doesn't cut it.
- bradleyjg 5y agoRight, and even five nines would be impressive compared to: AWS will use commercially reasonable efforts to ensure that each individual Amazon EC2 instance (“Single EC2 Instance”) has an Hourly Uptime Percentage of at least 90% of the time in which that Single EC2 Instance is deployed during each clock hour (the “Hourly Commitment”). In the event any Single EC2 Instance does not meet the Hourly Commitment, you will not be charged for that instance hour of Single EC2 Instance usage. This essentially forces the use of distributed computing for even small businesses.
- ev1 5y agoEC2 is absolutely not meant for this, though. Use an abstraction layer like Heroku if you're going to not understand what you're getting into. The amount of times I've had to 'advise' small businesses that are somehow running their small business site off a single EC2 instance's ephemeral boot volume is atrocious.
- bradleyjg 5y ago
- papito 5y agoFor something heavily used, like the EC2 and load-balancing, perhaps, but I am still experiencing PTSD from my last CloudFormation encounter.
- Frost1x 5y agoAWS has the advantage of having so many engineers behind the scenes available for firefighting with a culture of pushing more than is reasonable that as a customer, that would sort of disappear. They simply have to occasionally make trade offs between unrealistic development/feature request goals and firefighting whenever the firefighting is needed. This also acts as another form of pressure to work even more to meet timeline goals. Don't let your developers know that you're expecting them to always be behind, infinitely queued up with work, and constantly in emergency mode and they won't have much time to think about what's really going on and how efficiency is being pushed at the cost of their sanity.
- kottapar 5y agoThis is indeed surprising. Any time we have slowness issues the usual recommendation would be to throw resources at the problem; increase cpu, add more memory et al. We used to lament that we should spend time debugging the problem and fix the actual issue. We then used to say that probably at places like AWS and the other biggies they'd be following some excellent best practices and we should also strive to reach that level of excellence.
- awsthro00945 5y agoAWS is very big, culturally, on making sure that all the bugginess from shitty code is not shown externally to the customer. Externally it might look like everything is fine to you, but internally AWS is a massive, leaky cargo ship with thousands of engineers running around 24/7 with duct tape and band-aids to plug the leaks.