3 ms·
Good catch! I'll try to rerun the experiment. Hopefully the Google Colab TPUs give similar results to the Google Cloud ones so I can keep experimenting. Still,
by bigdatarepublic 8y ago
Good catch! I'll try to rerun the experiment. Hopefully the Google Colab TPUs give similar results to the Google Cloud ones so I can keep experimenting.
Still, since Adam performs worse even for non-distributed TPU (where the batch size and learning rate are the same for all devices) this wouldn't explain what we're seeing. Anyway, I'll try to find some time tomorrow to post the code so you can all have a look at it and hopefully see what's up.