4 ms·
Results are disappointing. 2x speed up over general purpose GPUs doesn’t justify a whole new hardware architecture.
by aey 7y ago
Results are disappointing. 2x speed up over general purpose GPUs doesn’t justify a whole new hardware architecture.
- anonuser123456 7y agoWhen Moores law is finally dead and buried (e.g. 5nm), new architectures will be all that's left. Seems like a great time to start down the new architecture path.
- deleted 7y ago[deleted]
- rrss 7y agoHave semiconductor companies not been on the new architecture path already? It's not like they've just been doing die shrinks this whole time.
- anonuser123456 7y agoYes, but the emphasis now is much greater than before. The old rules of turn the scaling crank that defined the industry for decades are no longer helping.
- zak 7y agoAuthor of the blog post here. I'd recommend doing a performance-per-dollar comparison before drawing this conclusion.
- aey 7y agoWhat’s the perf per watt difference? It’s hard to compare perf per dollar. Your electricity costs and GPU costs maybe vastly different from mine.
- solidasparagus 7y agoFor actual practitioners, TPUs are incredible. The cost/performance combo is unmatched. Now the real problem is, can you actually get a TPU pod in practice?
- aey 7y agoMy personal experience with crnns and lstms was that cutting training latency from 48 hours to 24 hours didn’t make a difference in our progress. Some architectures took 4x longer to run. 48 to 1 hour would have allowed us to run a few experiment per work day. That would be huge. If you told me that the next generation of TPUs could run the current 48 hour GPU training set in 1 hour, that would be something. But 2x faster won’t make a difference. Generally, startups for a new chip need to demonstrate 100x improvement over the general purpose approach to get funding. So this seems like another Google vanity project.
- solidasparagus 7y agoI don't understand. Time-to-train is always a product of model and the amount of hardware you use. TPUs are fast and cheap allowing you to use more hardware for the same cost. The TPU design also makes it scale quite well. If you really wanted to get from 48 hours to multiple experiments per work day, you could (probably) do that right now by scaling horizontally. However, the cost becomes a major issue when you do that on GPUs. This is where the TPU shines since you can scale out without cost crippling you. I work on training on GPUs and the TPU is definitely an incredibly useful piece of hardware, not a vanity project.