4 ms·
dumb question: wdym by sparse work? Is it embedding lookups? (TPUs have had BarnaCore for efficient embedding lookups since TPU v3)
by felarof 2y ago
dumb question: wdym by sparse work? Is it embedding lookups?
(TPUs have had BarnaCore for efficient embedding lookups since TPU v3)
- dekhn 2y agoMostly embedding, but IIRC DeepMind RL made use of sparsity- basically, huge matrices with only a few non-zero elements. BarnaCore existed and was used, but was tailored mostly for embeddings. BTW, IIRC they were called that because they were added "like a barnacle hanging off the side". The evolution of TPU has been interesting to watch; I came from the HPC and supercomputing space, and seeing Google as mostly-CPU for the longest time, and then finally learning how to build "supercomputers" over a decade+ (gradually adding many features that classical supercomputers have long had), was a very interesting process. Some very expensive mistakes along the way. But now they've paid down almost all the expensive up-front costs and can now ride on the margins, adding new bits and pieces while increasing the clocks and capacities on a cadence.