3 ms·
Sorry but even tho the big companies produce a lot of interesting research, I challenge you to not find any interesting model trained on a single GPU from recen
by caenorst 8y ago
Sorry but even tho the big companies produce a lot of interesting research, I challenge you to not find any interesting model trained on a single GPU from recent publications (the majority coming from academia). Actually it's very rare to find a paper where largely distributed training is necessary (i.e: the training would fail or would be unreasonably too long). Yes having more money help you to scale your experiments, it's nothing new and it's not something specific to AI.
- bitL 8y agoA trivial example is BERT_large; won't fit into 12/16GB and takes ~year to train from the scratch on a single 15TFlops machine. It's now a base model for transfer learning for NLP.
- caenorst 8y agoI'm not saying that there's no very big model, just saying that it's a minority of publications, for any trivia example of big model I can show you 10x trivia examples of relevant non-big models. Also you are talking about a model which is specifically designed for TPU (the dimensionality of the networks is especially fine-tuned). And even tho, BERT_large still fit in the memory of a single GPU (for very small batch), there is an implementation on Pytorch. I don't understand, are people complaining that Deep Learning is actually (reasonably) scaling ? Isn't it a good news ?
- bitL 8y agoYou need to study state-of-art a bit more. BERT_large can't reproduce results its authors achieved with TPUs on a Titan V/Tesla P100 as for getting there you need to use substantially larger batch sizes that won't fit into 12/16GB. If you get a V100/Titan RTX, it would fit, but you'd wait ~1 year for a single training session (40 epochs) to finish. MS already published another model based on BERT that is even better. It's unlikely memory x #GPUs would go down in foreseeable future; it's more like that everybody will start as large models as their infrastructure allows if they find something that improves target metrics.
- fnbr 8y agoYes, absolutely. I'm an AI researcher, and most of my colleagues just use their desktop computers with a standard GPU to do their work. Of course, long-running jobs get put on the cloud, as do large distributed jobs, but those are surprisingly rare. We just submitted a paper, for instance, that is entirely CPU based, and required running 4 CPUs for a few days to reproduce (and even then, you can reproduce 90% of the paper within minutes on a single machine).