5 ms·
Designing to remain robust in the face of failures is compulsory for any project of any significance. Or at least it should be, though a lot of projects go on a
by defaultname 5y ago
Designing to remain robust in the face of failures is compulsory for any project of any significance. Or at least it should be, though a lot of projects go on a wing and a prayer that nothing will go awry and "save" those engineering hours until a catastrophe at some future point. It basically just prioritized what already should be a priority.
I have no doubt that fringe/niche instances have more competitive spot behaviors, though how you set your bid range dramatically impacts how you survive through competition, but I had vanilla instances last for literal years (note that by default the spot requisition has a lifespan of one year so you have to modify that) at per hour pricing somewhere in the range of 1/5th on demand.
But mileage will vary.
I don't use those spot instances anymore as my projects are much better financed now, and I have significant compute on other platforms including bare metal in colocation facilities. However when I did I stayed silent about it, feeling almost like it was a secret that would be ruined if others knew about it.
- noogle 5y agoThere it is - the hidden cost of AWS. For bare-metal the risk of hardware failure is so low that it's faster to just handle the interruption when it happens (e.g. just restart the process) than to implement interruption tolerance. The hardware fails only once in many years. The chances of that happening during the 24 hours we train a model are almost zero. On a spot instances, the risk of the same are almost 100%, requiring investment up-front. For the price of a spot instance we can get an always-on bare-metal server without having to worry for how long it will remain available.