3 ms·
> They don't mention how/whether they tuned the learning rates and batch sizes to optimize for each different device. All networks were trained with the same h
by bigdatarepublic 8y ago
> They don't mention how/whether they tuned the learning rates and batch sizes to optimize for each different device.
All networks were trained with the same hyperparameters. Only the batch size was increased with the amount of parallelization used (so increased 8-fold for TPU and distributed GPU).
> Like they mention, they also use a very small network that isn't something you need the power of a TPU to train quickly and may scale differently than a large network.
I agree completely here. We unfortunately didn't have access to the TPUs long enough to create more useful benchmarks on networks like Resnet-50 and with bigger datasets.
> They also don't post their code so I can't check that their problems with ADAM aren't due to using L2 regularization, which https://arxiv.org/abs/1711.05101 https://arxiv.org/abs/1711.05101 shows leads to worse performance than SGD and you should use weight decay instead.
The code is the same for all devices and in the single-GPU and CPU benchmarks we see Adam performing better than SGD. We did not use any regularization besides a dropout layer so I don't think this explains the bad Adam performance on TPU.
I'll make some effort to clean up the code so it can be shared.
- zak 8y ago+1 on open sourcing the code. If you post it somewhere, we'll take a look to see why the Adam optimizer isn't behaving as expected in your implementation. It may also be helpful to compare your code with the Fashion MNIST Colab example here: https://colab.research.google.com/github/tensorflow/tpu/blob/master/tools/colab/fashion_mnist.ipynb https://colab.research.google.com/github/tensorflow/tpu/blob...
- trishume 8y agoThe fact that you didn’t change the learning rates for different batch sizes even following a linear scaling rule let alone tuning for each batch size is pretty important. IMO without tuning the learning rates comparisons across batch sizes are pretty meaningless because it changes the effective learning rate which can make a big difference in training.
- bigdatarepublic 8y agoGood catch! I'll try to rerun the experiment. Hopefully the Google Colab TPUs give similar results to the Google Cloud ones so I can keep experimenting. Still, since Adam performs worse even for non-distributed TPU (where the batch size and learning rate are the same for all devices) this wouldn't explain what we're seeing. Anyway, I'll try to find some time tomorrow to post the code so you can all have a look at it and hopefully see what's up.