8 ms·
Designing a scalable API on AWS spot instances
- ocdnix 6y agoTurns out this is about EC2 spot instances for ECS. How would it compare to ECS Fargate spot these days? I'm also missing a discussion about designing for interruption, either by not keeping state, or by being able to shed state quickly, to be picked up by other instances. Also, if you set up EC2 spot with a launch template or ASG with very differently-sized instance types (to reduce risk of running out), is there a way to even out the load coming through an ALB? The least-connections scheduling can help in some cases, but a connection might not map 1:1 to one unit of load. The ALB can use weighted balancing, but on the target group level. Dunno how easy it would be to allocate different instance sizes to different target groups and weigh them accordingly.
- kristianpaul 6y agoYou’re right , i was hoping to read more about “designing for interruption” as well because at end you can run spot instances and on-demand instances in the same ASG save money. https://docs.aws.amazon.com/autoscaling/ec2/userguide/asg-purchase-options.html https://docs.aws.amazon.com/autoscaling/ec2/userguide/asg-pu...
- ollyculverhouse 6y agoAFAIK with Fargate a lot of this is handled for you, as long as you have the auto scaling group. We have this setup with two capacity providers (FARGATE_SPOT and FARGATE) with a 75/25% split, meaning that even if there are no spot instances available we will still be up. The benefit of Fargate being that we don't need to care if certain instance sizes are not available as that is handled by AWS.
- l33tman 6y agoCool, when fargate launched they didn't have a spot possibility (AFAIK) and since we run ECS on Spot instances it would just be a massive increase in cost to switch to FG, but if it now can use underlying spot instances, it might be worth looking at again..
- Epixors 6y agoYeah spot capacity providers for Fargate only got added a few months ago, been running well for us in production.
- pluies 6y ago(Not the OP, but running a fairly similar setup, albeit for EKS nodes rather than ECS nodes :) ) Fargate Spot is about a third of the price of Fargate (at least in eu-west-1 now according to: https://aws.amazon.com/fargate/pricing/ https://aws.amazon.com/fargate/pricing/ ); so the savings are roughly identical. Re risk of running out, our current strategy is to use different-but-closely-similar instance groups; so for example we have an autoscaling group running a mix of: - m5.large - m5dn.large - m5n.large - m5ad.large - m5d.large Which are the same price on Spot instances, but I'd wager it'd be pretty rare to have all these families reclaimed at once. (We also use some on-demand only ASGs with lower priority in the cluster-autoscaler to ensure that if it _does_ happen, then we'll have a fallback)
- makkesk8 6y ago"Our backend system is built on AWS. Today I’m going to tell you how we had cut costs" This is a recurring topic here on HN and it boggles me and makes me wonder if people know that there are other platforms than aws, azure and google cloud out there that are very capable and much much cheaper. Unless any of the big 3 has a feature or certification you need I don't see any reason to use them at all due to the insane complexity and cost. So why do you or your company who uses any of the big 3 use them if you had to cut cost at some point?
- jorams 6y agoI can give you the reason we moved from Digital Ocean to Google Cloud: Spaces (object storage) was ridiculously slow and unreliable. There were several incidents per week where Spaces had bad connectivity or simply seemed to have crashed entirely, and that's not something we can deal with for production traffic. We now run a cheap CDN in front of it to cut down on the ridiculous bandwidth costs, but Google Cloud Storage has been very reliable and fast.
- saddlerustle 6y agoThis. There is no alternative for high performance object storage outside of AWS S3 and Google Cloud Storage. Even Azure's offering is wonky. Lots of providers claim to offer object storage, but try hitting them from couple thousand cores and they all tend to immediately fall over.
- speedgoose 6y agoThat's not a very common use case though. Most companies don't need to DDOS their cloud provider.
- jrockway 6y ago25Gbps is 250 cable Internet users downloading the release of your new software. What you call a DDOS is actually a woefully underwhelming instantaneous transfer rate.
- sunilkumarc 6y agoVery well written article! If someone is interested in learning all the AWS concepts, here's an awesome e-book which is written by the legend Daniel Vassallo himself. https://gumroad.com/a/238777459/MsVlG https://gumroad.com/a/238777459/MsVlG
- MrPowers 6y agoGreat article. Cutting ec2 costs is important, especially for companies with heavy data engineering / data science workflows. Spark service providers make it easy to spin up huge clusters (100+ nodes) to perform ad hoc analyses. The costs can quickly spiral out of control, even if you're getting 3x cost savings on the spot market. Some tangential thoughts: * Is there an AWS API that returns the cheapest availability zone in a region for a given instance type? Or is the GUI that's screenshot in the blog the only way to see? * I have seen the 90%+ cost savings for certain instance types * Sometimes you lose a spot instance, look at the pricing history graph to confirm the price spike, and don't see any spike that was above your bid price... it can be frustrating
- luhn 6y agoI'm not aware of any API (surprising for AWS), but you use Spot Fleet to get an array of spot instances optimized for cost. Having a bid above spot price does not guarantee you'll keep the spot instance. AWS can terminate a spot instance at any time if they need the capacity—That's the deal. It used to be more closely tied your bid price, but they've been moving away from that.
- jakozaur 6y agoFor service type workloads (e.g API service with 99.99% uptime SLA) we keep comparing on-demand vs. spot. In reality, you would like to compare Reserved Instances as you can get 60% discount. So in us-east-1: - spot costs: 35-43% of on-demand - RI 1 year standard: 60% of on-demand - RI 3 year convertible: 46% of on-demand So if you have some base load that you can commit to running for 3 years, the price gets often at spot range while not having to worry about losing capacity. In-reality sometimes combining reserved for some base capacity 60% + 40% spot for spiky seems to be the winning combination for many companies.
- kpotehin 6y agoGood point, thanks! We're looking at saving plans, probably will use them too. But when you're doing prototypes it's pretty much impossible to commit to 1 year of usage, let alone 3:)
- zerubeus 6y agoYou can still reserve classic instance for one year and when you don't use it anymore you can sell it in the market
- lyalu 6y agointeresting article!
- ditansu 6y agoVery useful article! waiting a terraform best practice
- arecurrence 6y agoThere's an excellent implementation using AWS lambda to manage spot instances at https://github.com/AutoSpotting/AutoSpotting https://github.com/AutoSpotting/AutoSpotting What's fantastic about the autospotting implementation is: 1. It dynamically replaces existing instances with spot instances just by setting a tag 2. Rather than replacing the instance with a fixed spot instance type, it will choose the cheapest that fits the requirements. 3. If there are no spot instances available that fit the requirements, it will spin up on demand instances until spot instances are available. If you have experience working with spot then you will know that these are really outstanding features that hopefully amazon will bake-in in the future.
- WatchDog 6y agoSeems like amazon has long been transitioning spot instances away from being a method of efficiently utilizing excess capacity, towards being a discounted service for less risk-averse businesses(businesses that can accept the risk of their service being terminated at any time). Even if in practice AWS never sees large spikes in compute demand and corresponding large scale instance preemption, most businesses I've worked with won't accept the risk of having OLTP systems be taken down at any time. No longer does spot seem to be a service where one can get a bargain for their compute intensive offline/batch workloads that are much more tolerant of preemption. Given that the spot prices seem to be very flat, and preemption is rare, amazon presumably have a fair bit of underutilized capacity, does anyone know if amazon uses this capacity themselves, or offers more aggressive spot pricing to select clients?