Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
bigdatarepublic
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
by
bigdatarepublic
8y ago
Good catch! I'll try to rerun the experiment. Hopefully the Google Colab TPUs give similar results to the Google Cloud ones so I can keep experimenting. Still, since Adam performs worse even for non-distributed TPU (where the batch siz
2.
▲
by
bigdatarepublic
8y ago
> They don't mention how/whether they tuned the learning rates and batch sizes to optimize for each different device. All networks were trained with the same hyperparameters. Only the batch size was increased with the amount of