3 ms·
1. Test your code with super low batch size. Bad for convergence, good for sanity check before submitting your job to a super computer. Or you can buy a deskto
by danieldk 6y ago
1. Test your code with super low batch size. Bad for convergence, good for sanity check before submitting your job to a super computer.
Or you can buy a desktop machine for the same price as an M1 MacBook with 32GB or 64GB RAM and an RTX2060 or RTX3060 (which support mixed-precision training) and you can actually finetune a reasonable transformer model with a reasonable batch size. E.g., I can finetune a multi-task XLM-RoBERTa base model just fine on an RTX2060, model distillation also works great.
Also, there are only so many sanity checks you can do on something as weak (when it comes to neural net training). Sure, you can check if your shapes are correct, loss is actually decreasing, etc. But once you get at the point your model is working, you will have to do dozens of tweaks that you can't reasonably do on an M1 and still want to do locally.
tl;dr: why make your life hard with an M1 for deep learning, if you can buy a beefy machine with a reasonable NVIDIA GPU at the same price? Especially if it is for work, your employer should just buy such a machine (and an M1 MacBook for on the go ;)).
- jefft255 6y agoAbsolutely agree! My points were more about the benefits of running code on your own machine rather than in the cloud or on a cluster. I don’t own an M1, but if I did I wouldn’t want to use it to train models locally... When on my laptop I still deploy to my lab desktop; this adds little friction compared to a compute cluster, and as you mention we’re able to do interesting stuff with a regular gaming GPU. When everything works great and I now want to experiment at scale, I then deploy my working code to a supercomputer.