3 ms·
Hopefully your progress gets saved in time when the spot instance inevitably gets terminated in the midst of training.
by snerbles 4y ago
Hopefully your progress gets saved in time when the spot instance inevitably gets terminated in the midst of training.
- acetabulum 4y agoIf you use Horovod Elastic, I think you can avoid this problem working across a cluster of Spot instances. https://horovod.readthedocs.io/en/stable/elastic_include.html https://horovod.readthedocs.io/en/stable/elastic_include.htm...
- belter 4y ago"Managed Spot Training..." "...Spot instances can be interrupted, causing jobs to take longer to start or finish. You can configure your managed spot training job to use checkpoints. SageMaker copies checkpoint data from a local path to Amazon S3. When the job is restarted, SageMaker copies the data from Amazon S3 back into the local path. The training job can then resume from the last checkpoint instead of restarting...." https://docs.aws.amazon.com/sagemaker/latest/dg/model-managed-spot-training.html https://docs.aws.amazon.com/sagemaker/latest/dg/model-manage...