4 ms·
My personal experience with crnns and lstms was that cutting training latency from 48 hours to 24 hours didn’t make a difference in our progress. Some architec
by aey 7y ago
My personal experience with crnns and lstms was that cutting training latency from 48 hours to 24 hours didn’t make a difference in our progress. Some architectures took 4x longer to run.
48 to 1 hour would have allowed us to run a few experiment per work day. That would be huge.
If you told me that the next generation of TPUs could run the current 48 hour GPU training set in 1 hour, that would be something.
But 2x faster won’t make a difference. Generally, startups for a new chip need to demonstrate 100x improvement over the general purpose approach to get funding. So this seems like another Google vanity project.
- solidasparagus 7y agoI don't understand. Time-to-train is always a product of model and the amount of hardware you use. TPUs are fast and cheap allowing you to use more hardware for the same cost. The TPU design also makes it scale quite well. If you really wanted to get from 48 hours to multiple experiments per work day, you could (probably) do that right now by scaling horizontally. However, the cost becomes a major issue when you do that on GPUs. This is where the TPU shines since you can scale out without cost crippling you. I work on training on GPUs and the TPU is definitely an incredibly useful piece of hardware, not a vanity project.
- aey 7y agoAccording to the post, it’s only 2x cheaper. So for the same dollars spent, I would cut my time from 48 hours to 24. But, TPUs are not standard, can’t be used for any other usecase. So those savings might not actually be realizable when everything is taken into account.
- p1esk 7y agoWhy would you use cloud if you do a lot of training? Quad 2080Ti systems cost ~$7k + electricity. Assuming two such systems are equivalent in speed to 1 TPUv3 (4 chips), owning them would be more cost efficient after 4-5 months of training (depending on your electricity costs). TPUs do have the memory capacity advantage though (over 2080Ti).
- solidasparagus 7y agoThat is not what the blog is saying. The blog post does not consider cost - it simply shows that a TPU v3 Pod is twice as fast as the largest DGX-2h cluster.