4 ms·
In cases where you have to deploy the model and you are limited in terms of flops, this paper does not help much, unless it’s removal of batchnorm somehow allow
by dontreact 6y ago
In cases where you have to deploy the model and you are limited in terms of flops, this paper does not help much, unless it’s removal of batchnorm somehow allows a future network that is actually faster at inference time.
- modeless 6y agoThere are a lot of techniques for sparsifying or pruning or distilling models to reduce inference FLOPS, and they almost always produce better results when starting with a better model. Also, if your model is 8x faster to train at the same size then you can do 8x as much hyperparameter tuning and get a better result.
- dontreact 6y agoThis model is much more expensive than efficientnet at inference (I think the flops are about 2x?). You can use these same techniques with efficientnet.
- modeless 6y ago> they almost always produce better results when starting with a better model.
- dontreact 6y agoIf you have a flops limit this new model would first need to be shrank by 2-3x as much as Efficientnet in order to fit the same constraint. So you would be starting with a smaller model and thus lower performance. Efficientnet is still better for embedded applications most likely.
- david-gpu 6y agoBut for deployment in smaller devices you can use techniques such as distillation, quantization and sparsity. Training and inference are very different problems in practice.
- dontreact 6y agoYes but you can do that with efficientnet as well. The point is that this is an improvement only for training because it uses computations which are highly optimized on TPU