4 ms·
At its core, this tool doesn't really care about how/where you store your data. So, one possible way to train a model on ImageNet using this tool could be to:
by subtech 5y ago
At its core, this tool doesn't really care about how/where you store your data.
So, one possible way to train a model on ImageNet using this tool could be to:
1. Store the dataset on S3
2. Have the training script download the necessary data from S3 at the beginning (for example, each process could only download the shard of data that it needs, etc.)
3. Use MLbot to run your local training script on AWS as a distributed training job with X number of nodes in your EKS cluster
In this case, when you execute "mlbot run ..." from the command-line, MLbot will package your code into a docker image and then create X number of pods within EKS (where X = the number of training nodes).
Once all of the pods are running, etcd + PyTorch Elastic begin handling all of the communication / synchronization between the various pods & distributed training happens.
So in this example, your code would be transferred to your EKS cluster (where the training would happen) and the training pods would then fetch the ImageNet data from S3.
And most importantly, all of this happens inside of your cloud infrastructure.
Now that I think about it, I think it might be helpful to add an ImageNet-based example to the repo.
- p1esk 5y agoI think it might be helpful to add an ImageNet-based example to the repo. Yes, that would be helpful. I’m not sure if S3 is fast enough to serve Imagenet one batch at a time - might become a bottleneck.