3 ms·
The speed improvements are certainly interesting, the performance improvements seem decidedly not. This method has more than 2x the parameters of all but one of
by smeeth 6y ago
The speed improvements are certainly interesting, the performance improvements seem decidedly not. This method has more than 2x the parameters of all but one of the models it was compared against.
If I’m off-base here can someone explain?
- modeless 6y agoI don't care how many parameters my model has per se. What I care about is how expensive it is to train in time and dollars. If this makes it cheaper to train better models despite more parameters, that's still a win.
- sillysaurusx 6y agoThere's one important caveat, though I agree with your thrust: at GPT-3 scale, cutting params in half is a nontrivial optimization. So it's worth keeping an eye out for that concern. (Yeah, none of us are anywhere near GPT-3 scale. But I spend most of my time thinking about scaling issues, and it was interesting to see your comment pop up; I would've agreed entirely with it myself, if not for seeing all the anguish caused by attempting to train and deploy billions of parameters.)
- 6gvONxR4sf7o 6y agoSome models are still memory limited. Fewer parameters are very useful in those settings.
- dontreact 6y agoIn cases where you have to deploy the model and you are limited in terms of flops, this paper does not help much, unless it’s removal of batchnorm somehow allows a future network that is actually faster at inference time.
- modeless 6y agoThere are a lot of techniques for sparsifying or pruning or distilling models to reduce inference FLOPS, and they almost always produce better results when starting with a better model. Also, if your model is 8x faster to train at the same size then you can do 8x as much hyperparameter tuning and get a better result.
- dontreact 6y agoThis model is much more expensive than efficientnet at inference (I think the flops are about 2x?). You can use these same techniques with efficientnet.
- modeless 6y ago> they almost always produce better results when starting with a better model.
- dontreact 6y agoIf you have a flops limit this new model would first need to be shrank by 2-3x as much as Efficientnet in order to fit the same constraint. So you would be starting with a smaller model and thus lower performance. Efficientnet is still better for embedded applications most likely.
- david-gpu 6y agoBut for deployment in smaller devices you can use techniques such as distillation, quantization and sparsity. Training and inference are very different problems in practice.
- dontreact 6y agoYes but you can do that with efficientnet as well. The point is that this is an improvement only for training because it uses computations which are highly optimized on TPU