10 ms·
We reduced the AWS costs of our streaming data pipeline
- throwaway888abc 6y ago"Eliminate unused EC2 instances" -27% of cost Haha, so cleaned the internal IT / DevOps mess and call it a day and than blog post it
- tjbiddle 6y agoEh, you're pulling a quote out of context. It's 27% reduction in EC2 use, which was only 18.5% total. So this only accounted for ~5% of total savings.
- dannyw 6y agoI mean if a quarter of your EC2 instances were unused, that is absolutely an internal devops / IT mess. The whole point of AWS is to use services on demand; it's like buying 133 conference tickets for your 100 person company.
- brianwawok 6y agoMore like ordering 133 lunches every day for your 100 employees and dumping 33 in the trash.
- billyhoffman 6y agoAnd doing this for months. Without noticing. Honestly this isn’t ultimately engineerings fault. This is a SaaS business Someone in their company is responsible for the COGS KPI. For that person to either not notice an increase in COGS, or to not be aggressively incentivizing engineering to reduce COGS, is giant red flag.
- zaphirplane 6y agoOuch this thread ruined some poor souls promotion / pay rise narrative
- jinpan 6y agoThe article did say that the motivation for doing so was because AWS credits were running out .. why prematurely optimize a free resource? :)
- rlayton2 6y agoPerhaps more like ordering 100 lunches for your 100 employees and forgetting that 20 of them moved to night shift?
- sankalp221 6y agoFun fact, 30-40% of the food in our supply chain is wasted. So, we're actually not too far from that. Sure, the meals aren't directly wasted by the company and rather through the supply chain. But, the waste is still there and is surprisingly high IMO. Source: https://www.usda.gov/foodwaste/faqs https://www.usda.gov/foodwaste/faqs
- redis_mlc 6y ago> The whole point of AWS is to use services on demand That's a decade-old misconception about how people actually use AWS. Most servers I've seen in AWS are permanent. In fact, it's an anti-pattern to wait until you need more capacity to scale up, since those servers may not be available, especially in newer instance families. Even if the needed instances available, typically ASGs don't react the way you expect without a lot of experimentation (ie. outages.) An example is if traffic increases load, your health check may consider the servers to be unhealthy and start killing them, creating a death spiral.
- tilolebo 6y agoThis is why service health endpoints should be very carefully designed, with the proper balance between health inspection and performance penalty. That being said, it's true that an ALB doesn't offer throttling capability like a true reverse proxy such as HAProxy provides, were you can cap the number of concurrent requests and give a chance to your backend to avoid death by overload. I wish there would be a way for ASGs to at least make the distinction between an unhealthy instance and an overloaded one.
- glogla 6y agoI think there's different levels of "on demand". In this case, I think this is much slower level of "on demand". I see two primary use cases of cloud: 1) You're a startup or just need something small, and want to focus on building your MVP, instead of messing around with colocated Linux servers. Cloud is much more expensive than those, but you don't care because you don't really need all that much, or maybe you are VC backed and have unlimited money. 2) You're a large company with broken internal processes. You can get server in company datacenter in three months after seven approvals (since it's capex), or you can spin up an EC2 instance. You don't care about cost since you have unlimited money. Those are kind of medium scale "on demand" - not "I need 100 new servers right this minute" but "I need server in ten minutes intead of 'when I get to buy one' or 'in three months and 37 forms'". In both cases, you're throwing money away, because time is more important for you than the extra money cloud costs.
- tjbiddle 6y agoI don't disagree with it that at all. By all means, we should be doing our best not to waste resources. Just stating that the waste was only 5% of the total savings.
- QuinnyPig 6y agoHmm. This looks to me like a lot of the savings were realized by moving away from managed services into a scenario where there’s more operator overhead. The AWS bill gets lower, but what about the cost of the engineering work?
- alharith 6y agoDoes anyone else find the costs associated with running well-tested, well developed systems overblown? Like if you know how to adjust some basic parameters, you will solve for 99% use cases (adjust memory, adjust ram). Examples I can think of is Rabbit MQ and Cassandra. But in general, we have some really battle-tested software these days that has become simpler to configure and run over time. People seem scared to run their own these days.
- owenmarshall 6y agoI vouched for this comment because it’s a valid point and I’m not sure why it was killed. I happen to disagree strongly, though: lots of engineers in my experience undervalue the work of systems administrators and underestimate the effort needed to operationalize any technology. Running your own is absolutely fine if you are willing to keep your stack small and invest time learning the tools you pick. But there are still horror stories of people thinking snapshots are backups, turning the wrong knobs and turning off fsync on their databases, ...
- chrischen 6y agoYea exactly and unless you are FB scale you can just run a single docker container and never really have to worry (granted you know how to use Docker). Most small startups are actually the ones who don’t really need SaaS services.
- ckdarby 6y ago>Yea exactly and unless you are FB scale you can just run a single docker container and never really have to worry This has not been the case at multiple employers and or consulting clients. If you're providing software to an enterprise this almost will never fly. That single docker container will have an outage when basically anything happens. The container dies, systemd fails to restart, node dies, network switch dies, data center has basically any major issue, etc. I think your comment brings value just probably biased with your own experience of running a consumer to consumer startup.
- dirtydroog 6y agoWe went through a similar process with GCP, which was annoying since GCP was sold as being cheaper than AWS.
- aritraghosh007 6y agoBack when AWS started, there would be articles about the work to master scalability and performance for the modern web but as things matured, we somehow ended up in a much larger heap of literature around AWS cost optimization.
- derex 6y agoIn some sense this is a good problem to have. With on-prem you used to have very limited resources to start with, so cost efficiency is a baked-in requirement. With cloud providers you seem to have limitless resources and the new problem of cost optimization arises. Admittedly there’s difference between optimizing fully-controlled resources and cloud provider managed services. For one, low visibility into cloud service internals makes such optimization harder.
- ooobit2 6y agoWe ended up here because Amazon can't scale. It's just uncool to admit you have to notice the pink elephant. Why? I don't know. Maybe it has to do with cred in engineering teams or for engineering teams in the broader org structure. But the problem with AWS, with a lot of the "cloud", is the pitch that remote centralization of a service scales ad infinitum. It's still subject to the same constraints as self-managed, even if those constraints appear at a higher limit. The greatest constraint is the per-unit pricing. You buy self-managed, you have huge upfront and period costs, but with remote, you see the $.03/MB price and assume that variable cost is more manageable over the long run. And it is... until price changes, overhead changes, bandwidth changes, or worse, accessibility changes. And suddenly, what you had cost-effective scaling on 18 months ago now has a massive deficit affixed to it. Because that's how most people used the platform... or because removing A or B features reduced maintenance costs or freed up bandwidth. AWS is an experiment. Does it work in many or even most use cases? Yes. For now. I love engineers. A lot. In fact, being in sales, I would give up a deal with an engineering team unless I knew for sure my ROI basis was solid. That said, I do know sales and marketing rhetoric. And having spent hundreds of hours in meetings with product, marketing and dev professionals, I wish I could record the stress-induced breakdowns I've seen in engineers and executives who had everything running buttery, "and then [provider] pushed [update]..." and they then have executives breathing on the back of their neck 16 hours a day, entire teams offline or unable to do basic tasks, etc. I just want to play that shit to people and say, "This is why you don't overpromise."
- meritt 6y agoI find it highly entertaining a two-year old company who was founded on the basis of helping slash cloud spending found so much waste in their own AWS spend. This is not an example of dogfooding, but an example of sheer incompetency and massive technical debt. I'd really like to start seeing a series of blog posts from companies who are running extremely lean and efficient tech environments by utilizing cloud in an intelligent manner and avoiding the expensive and unnecessary bullshit that's so prevalent today. The ones that can brag "How we run a $4M/yr SaaS on $40k/yr of AWS spend!" are far more interesting than "How we stopped incinerating millions of VC money by simply turning off shit we didn't need"
- tilolebo 6y agotwo-year old companies have limited resources. It might have been a deliberate trade off to focus on work that produces value to the customers. Maybe the blog post would have been "How we run a $1M/yr SaaS on $40k/yr of AWS spend!" instead of $4M?
- theatraine 6y agoInteresting idea. Does anyone do this for Azure?
- anthonysarkis 6y agoIt seems reasonable to make some of these cost comparisons more visible. ie If working on a new product or feature to understand upfront "this managed service is x% more then more bare bones" etc. essentially turning an alchemy into a science
- Cthulhu_ 6y agoAWS offers a cost calculator for just that purpose; they offer 'easier' products if you can't be arsed to dive into AWS costs and technologies yourself. I think a lot of people make the mistake of assuming AWS is just an easy off-the-shelf thing you can just grab, but if you use it seriously it's a full-time job and its own expertise. Source: I've done some AWS certifications, never was able to put them into practice though. I've also worked in multiple organizations that migrated to AWS, they all had a full-time team of people managing it. It's a full-time, specialist job and you can't just palm it off to your engineers as a background thing.
- agounaris 6y agoI am curious about the actual cost in $! Managing your own kafka or observability infra is expensive, you need a team to do this. A 67% reduction doesn't say the whole truth. They have more services to manage now, which means they need more people and more time to do this. Saving 10k from your AWS bill by hiring 2 more engineers is not cost effective.
- nojito 6y agoOr course it is. Where did we get the idea that engineers are hired to do only one thing? This has never ever been the case in my experience. Also Kafka being hard this manage is not the case. A simple look into many small companies and startups running their own clusters shows otherwise.
- agounaris 6y agoEngineers are hired to deliver and produce value. Tooling can be a part of it but if you can outsource something which is not your source of income, you have to do it. Engineering time is more valuable. I also know many startups and small companies investing 5 people and 6 months to get an observability platform up and running while they could just get datadog or new relic for half the price... and I don't get into account outages and updates to the platform. I remember a recent uber blog post on how they moved from build tool A to build tool B and a couple of weeks later, 3000 people where laid off. It's important to spend development time on revenue streams. This is some nice piece of advice https://nav.al/build-a-team-that-ships https://nav.al/build-a-team-that-ships "Outsource everything that isn’t core. Resist the urge to pick up that last dollar. Founders do Customer Service."
- cthalupa 6y ago>Where did we get the idea that engineers are hired to do only one thing? At a certain size or number of self run services, they very well might be. I used to be the guy that did the set up for these sort of self managed solutions, and ran them day to day. In some shops the workload was high enough we needed multiple people like me doing it. Or a whole team. Doing DevOps style management of them just let us do it with fewer people - it certainly didn't make it feasible for developers to do the day to day management of these services and still write code.
- tyingq 6y agoThe initial pie chart seems to indicate that either AWS glue is significantly overpriced, or that they were doing something wrong.
- brodouevencode 6y agoAs with all things AWS the more "magic" there is to it, the more expensive it is.
- csharptwdec19 6y agoHuge part of why I always try to build applications as platform agnostic as possible. If I make a .NET service or site, I know (with the tools I use) I can deploy it on any linux or windows machine without issue. I can take it anywhere that I can run any software. Sure, may need more glue for certain scenarios, but you know that you can move as soon as a provider shows it's fangs.
- brodouevencode 6y agoSpeaking from experience - not a bad idea.