2 ms·
I'm not saying that there's no very big model, just saying that it's a minority of publications, for any trivia example of big model I can show you 10x trivia e
by caenorst 8y ago
I'm not saying that there's no very big model, just saying that it's a minority of publications, for any trivia example of big model I can show you 10x trivia examples of relevant non-big models.
Also you are talking about a model which is specifically designed for TPU (the dimensionality of the networks is especially fine-tuned).
And even tho, BERT_large still fit in the memory of a single GPU (for very small batch), there is an implementation on Pytorch. I don't understand, are people complaining that Deep Learning is actually (reasonably) scaling ? Isn't it a good news ?
- bitL 8y agoYou need to study state-of-art a bit more. BERT_large can't reproduce results its authors achieved with TPUs on a Titan V/Tesla P100 as for getting there you need to use substantially larger batch sizes that won't fit into 12/16GB. If you get a V100/Titan RTX, it would fit, but you'd wait ~1 year for a single training session (40 epochs) to finish. MS already published another model based on BERT that is even better. It's unlikely memory x #GPUs would go down in foreseeable future; it's more like that everybody will start as large models as their infrastructure allows if they find something that improves target metrics.