8 ms·
This is quite a neat strategy, leveraging elastic compute costs and kubernetes "self-healing". I'm surprised I haven't heard more about this kind of technique b
by PudgePacket 7y ago
This is quite a neat strategy, leveraging elastic compute costs and kubernetes "self-healing". I'm surprised I haven't heard more about this kind of technique before.
I fully acknowledge this will only work in certain scenarios and for certain workloads, eg not ideal for long running/cache/database style services.
- toomuchtodo 7y agoThis is run of the mill for job schedulers (Nomad, Mesos/Marathon, Globus), it’s just more accessible with k8s, containers, and VM spot pricing than historically. As you mention, definitely don’t do this where persistence is paramount (cache that is expensive to backfill upon recovery, database, etc) but it’s just fine for transient workloads or workloads you can rapidly and safely preempt and resume.
- 3fe9a03ccd14ca5 7y agoIf your instance is stateless, and your app can easily self-heal, there’s lots of computing paradigms you can explore. Serverless/function is also an option. Of course, at the end of the day rarely is something ever truly stateless.
- deleted 7y ago[deleted]
- halbritt 7y agoEh, I'm currently running an old version of DCOS which does a pretty terrible job of scaling down. Much more than just a scheduler is required to make this work well.
- ec109685 7y agoYeah, your fleet can become fragmented.
- deleted 7y ago[deleted]
- gopalv 7y ago> leveraging elastic compute costs and kubernetes "self-healing" The indirect effect of building on a system like this is that the recovery mechanisms get tested on a regular basis instead of just on the odd day when things fail. Spot instances are like a natural chaosmonkey mode, with money being saved and forcing you to build failure tolerance, retries & circuit breakers early in dev.
- halbritt 7y agoGCP has a similar instance type called "preemptible", they're not quite as cheap as spot, but they don't "dry up" and they're guaranteed to go down every 24 hours. This precludes one from becoming complacent with spot instances that rarely go away.
- tuananh 7y agoYou're right. spot instances are a lot more stable than preemptible. we've seen spot instances that last a year for us.
- halbritt 7y agoThis is the second response of this sort that I'm replying to. "A lot more stable" isn't really a desirable characteristic of ephemeral compute capacity. In my experience, the less frequently the instances went away, the more complacent the operators became. Preemptible instance are stable in the sense that you know they're going away within 24 hours and must be prepared for that.
- tuananh 7y agotrue that. spot used to be that way and the price was very sensitive. but AWS tweaked it so that it's more stable. to the point, after a year or 2 of running spot instances, we don't feel the difference of spot and ondemand that much. we got complacent.
- jedifans 7y ago
- dcolkitt 7y ago> not ideal for long running/cache/database style services. Well, one question to ask yourself when considering going down this route is whether it makes more sense to move all the statefulness into managed services, like Aurora, BigTable, S3, etc. That drastically simplifies life. Now the only infrastructure directly managed by you are stateless workloads that can easily be self-healed, rolled back, scaled up/down, etc. Managed DBs are more expensive than running your own DB, but most likely the cost savings of moving the rest of the infrastructure to spot/preemptible outweighs this difference.