3 ms·
I don't think this will have anything to do with training, since it's the TPU team and trying to beat nVidia at floating-point matrix multiplication doesn't see
by highd 9y ago
I don't think this will have anything to do with training, since it's the TPU team and trying to beat nVidia at floating-point matrix multiplication doesn't seem like a great idea. Run-time can generally be done with integer operations, which is near dirt-cheap these days - I'd think the only interesting things would be operating at server scales or so low-power that you're getting 30FPS on your classification/segmentation what-have-you on your smartphone without impacting battery life. My guess is the former, though I'm not sure what the market size is for companies that need to evaluate their machine learning systems at speeds fast enough to merit specialized hardware.