4 ms·
Interesting. Plenty of work has been done with FPGAs, and a few have developed ASICs like DaDianNao in China [1]. Google though actually has the resources to de
by manav 10y ago
Interesting. Plenty of work has been done with FPGAs, and a few have developed ASICs like DaDianNao in China [1]. Google though actually has the resources to deploy them in their datacenters.
Microsoft explored something similar to accelerate search with FPGAs [2]. The results show that the Arria 10 (20nm latest from Altera) had about 1/4th the processing ability at 10% of the power usage of the Nvidia Tesla K40 (25w vs 235w). Nvidia Pascal has something like 2/3x the performance with a similar power profile. That really bridges the gap for performance/watt. All of that also doesn't take into account the ease of working with CUDA versus the complicated development, toolchains, and cost of FPGAs.
However, the ~50x+ efficiency increase of an ASIC though could be worthwhile in the long run. The only problem I see is that there might be limitations on model size because of the limited embedded memory of the ASIC.
Does anyone have more information or a whitepaper? I wonder if they are using eAsic.
[1]: http://ieeexplore.ieee.org/xpl/login.jsp?tp=&arnumber=7011421&url=http%3A%2F%2Fieeexplore.ieee.org%2Fxpls%2Fabs_all.jsp%3Farnumber%3D7011421 http://ieeexplore.ieee.org/xpl/login.jsp?tp=&arnumber=701142...
[2]: http://research.microsoft.com/pubs/240715/CNN%20Whitepaper.pdf http://research.microsoft.com/pubs/240715/CNN%20Whitepaper.p...
- hedgehog 10y agoYou can build an ASIC with fast external memory, it adds to the cost but then you can handle larger models similar to a GPU. Software support is an issue but for deep learning applications there's no reason in principle you couldn't add support to TensorFlow etc for new hardware to make it simple for application developers to adopt. Movidius has announced that they're doing this and it's likely that other ML chip vendors will do the same.
- emcq 10y agoExternal memory kills your power budget. Efficient processors, whether they be GPUs, cpus, or asics all gain efficiencies by keeping data local to computation. Having your chip focused on off chip data is difficult to optimize power for if you don't have compute bound problems.
- petra 10y ago>> I wonder if they are using eAsic. If you read the academic literature, they talk about 100x-3000x energy savings with asic vs GPU. So in that light Google's 10x improvement sounds low, and could certainly fit an eAsic story. Furthermore, eAsic has fixed wires, but logic is defined via sram configuration(AFAIK) and not fixed, so it could offer a level of programmability .