5 ms·
My take at making AWS EC2 cheaper by automating SPOT instances with AutoScaling
- kiwidrew 10y agoWhy would you bother going to all this trouble when you could just use Amazon's official Spot Fleet service/API? http://docs.aws.amazon.com/AWSEC2/latest/UserGuide/spot-fleet.html http://docs.aws.amazon.com/AWSEC2/latest/UserGuide/spot-flee...
- dmourati 10y agoHe started before the feature was released and acknowledges they came up with an in-product solution: "and it seems they now have a full fledged solution for the problem, based on pretty much a reimplementation of AutoScaling, using machine learning and with a beautiful UI and they are really successful with it. Funnily enough, they even contacted me to sell that solution to my company and we are seriously evaluating it"
- kiwidrew 10y agoAh, okay, I missed that part!
- alien_ 10y agoYes, but I actually did start a few weeks before the spot fleet was launched. The problem with the spot fleet 1) it's kind of awkward to use 2) it has statically defined capacity so you can't scale it 3) it has a static bid price, so if at some point your are outbid on all the group's bids, you end up with no capacity 4) among other things it lacks integration with the ELB so you can't really use it for so many use cases. My solution is simpler, better integrated with the rest of AWS and more resilient and once I iron out the bugs and get it production-ready, it should be a better choice.
- chanakya 10y agoHe's talking about Spotinst.com, not Amazon's fleet solution.
- falsedan 10y agoSpot prices can spike above on-demand prices for some instance types (e.g. g2.2xl). If you wanted to maintain a certain capacity of homogeneous machines, you'd have to notice the failed spot instance request & increase the size of your on-demand ASG to compensate.
- asteadman 10y agoCan anyone comment on the crazy spot prices for g2.8xl ? they spike really high (they were $26/hr earlier? 10x the on-demand price. I'm guessing someone with enough market share has a job that doesn't want to be interrupted and they bid the spot price up to 10xon-demand, which seems ridiculous since presumably they could arbitrage /w an on-demand, but i guess that once they start the job they don't want to be interrupted? i'm not exactly familiar with the intricacies of the ec2 spot market. ). Also, are they any good for ML? I'm between machines ATM and would like to experiment with some deep CNN for face recognition. At $2.6/hr for on-demand its a bit more than I would like, but if the spot price was less i'd consider it.
- IanCal 10y agoCan you checkpoint your work? Most libs I've used support that generally. That way you only lose a bit of time if kicked off, and you don't pay for the hour you get kicked off so it's even cheaper.
- dharma1 10y agoI used them for a while for ML, works ok. I also noticed the these spot price spikes with both g2.8xl and g2.2xl. Not really sure why someone (or multiple players, since the prices spike) have set their max bid so high, surely it's worth running an on-demand instance at those prices. Ended up going between two regions to avoid them but it's some hassle, just got a gtx970 in the end.
- ecesena 10y agoOne issue that we had in the past was the following - hope this is resolved now. Say you have an autoscaling group spanned across zones A and B, and say you have 1 machine in zone A and 1 in machine B. Now, the price in zone A goes up, and your machine in zone A dies. The issue (bug?) was that the autoscaling group was trying to re-instantiate a new VM in zone A. Of course, since the price was high, the new VM was basically immediately dying. And so on. Edit: issue apart, it's a great way that can save money, especially if you have a group of VMs whose computation can be interrupted/restarted relatively cheaply.
- samlin86 10y agoI agree that was an annoying issue. They've recently changed it though, such that it will retarget unfulfilled bids for spot instances in a different zone in order to reach your desired capacity. Win! See: https://forums.aws.amazon.com/ann.jspa?annID=3647 https://forums.aws.amazon.com/ann.jspa?annID=3647
- alien_ 10y agoWith my approach the AutoScaling group would always replace failed instances with the on-demand ones identical to those initially defined on the group's launch configuration. All I do is I later attempt to replace them with whatever I can buy from the spot market. On the other hand, the spot bidding implemented out of the box in AutoScaling will fail if you are outbid in all Availability Zones at the same time, since it doesn't fall back to on-demand instances. I've seen people often use a second on-demand AutoScaling group that would scale out when you get outbid on the spot one, but then you have a problem defining scaling policies so that they can scale nicely, and/or shifting the capacity between them. Someone had a nice talk at re:invent about how they do all that.
- voltagex_ 10y agoI've got https://github.com/voltagex/junkcode/tree/master/Python/spotprices https://github.com/voltagex/junkcode/tree/master/Python/spot... and https://github.com/voltagex/junkcode/tree/master/CSharp/spotprices https://github.com/voltagex/junkcode/tree/master/CSharp/spot... but I was never really happy enough with either to take them past my junkcode folder. If I could make the Python one work a bit better, I'd probably use it with Ansible - http://snappishproductions.com/2014/03/24/Spot-Instances-With-Ansible.html http://snappishproductions.com/2014/03/24/Spot-Instances-Wit...
- vgt 10y agoI'm going to plug Google Cloud's Preemptible VMs as a simpler alternative to Spot Instances: - Preemptible VMs are sold at a fixed 70% off discount, removing pricing volatility entirely - Google Cloud's Compute Engine has far fewer VM types, thus making it much easier to construct the resources you want (exception being GPUs, you don't need a "network optimized" instance to get fast network, nor you need a "storage optimized" instance to get fast/large storage - these things are modular on GCE). (disc: work on Google Cloud)
- merb 10y agoActually it would be great if your lower instances would have more memory. For low cost projects (small companies or society club) it's really odd to pay the server on something like ovh.de (more memory cheaper, nearly same performance) and the storage for GCE or AWS. I mean yeah I just realized that GCE could be cheap, still somehow the Micro instance could've had a little bit more memory. And small is already out of budget. I mean something between Small and Micro would be great, like 1.2 GB Memory with a price of 6.6 USD (which would be 7 USD with a 10 GB persistent disk).
- boulos 10y agoDoes this mean your target budget is $10/month? (I'm just asking for clarification, and to understand how you think about it). Second, what are you looking to run? A Java-based web app? Does App Engine's free tier not serve you better? Finallty, and it's perhaps poor form to point this out, but we (and AWS) will end up blowing out your budget once you include networking while OVH/Hetzner/Digital Ocean won't. Compute Engine isn't a VPS, we're giving you a small slice of a machine but it's connected to a crazy awesome network that we bill for by the byte. When you compare that to a bundled, heavily-overcommitted across tons of customers VPS networking offering the dollars don't work out. Disclosure: I work on Compute Engine and care a lot about our pricing.
- Veratyr 10y agoHopefully this comment doesn't sound too stupid but is it at all possible (or likely to be implemented) to get a worse network on GCE and pay less for it? I ask because I don't need a CDN level network like GCE provides but I do need a ton of compute and a ton of bandwidth for batch data processing on data hosted at another provider (and not a candidate for GCS). A pre-emptible network offering would open up the way for me to use compute resources, which is surely a good thing.
- flaviotruzzi 10y agoLooks like https://github.com/chaordic/tiopatinhas https://github.com/chaordic/tiopatinhas
- alien_ 10y agoI didn't know about that one, indeed, it seems pretty similar.
- alien_ 10y agoI've looked in a bit more detail and it seems to attach the nodes to the ELB used by the AutoScaling group, while I am attaching them to the group itself, which indirectly adds them to the load balancer. I'm curious how does that tool handle the scaling of the group, since the nodes are actually outside the group and can't contribute to group-wide metrics like average CPU usage, often used for scaling out.
- boulos 10y agoWhy do only update / rebalance every 30 minutes? Is that because of the per-hour billing (so replacing an on-demand one only part way through its life is a mistake). If so, it still seems like you'd want to inspect every 5 minutes or so (keep on-demand until it reaches > 60-frequency). Disclosure: I work on Compute Engine (and launched our preemptible VMs product).
- alien_ 10y agoIt's because so far I am only considering replacing the nodes slowly, one at a time, mostly in order not to hit any soft limits that may be defined on the account(like total number of instances), but also in order to give the user the chance to stop it in case things go south for whatever reason during the evaluation, since as I warned, this thing is likely full of bugs at this point.