6 ms·
3 generations ahead of moore law??? I really wonder how they are accomplishing this beyond implementing the kernels in hardware. I suspect they are using spec
by jhartmann 10y ago
3 generations ahead of moore law??? I really wonder how they are accomplishing this beyond implementing the kernels in hardware. I suspect they are using specialized memory and an extremely wide architecture.
Sounds they also used this for AlphaGo. I wonder how badly we were off on AlphaGo's power estimates. Seems everyone assumed they were using GPU's, sounds like they were not. At least partially. I would really LOVE for them to market these for general use.
- reitzensteinm 10y agoBut isn't 3 generations ahead just 8x? Which doesn't sound at all unreasonable for a custom hardware.
- dnautics 10y agoThis is about right! 64-bit IEEE fp -> 16-bit IEEE-style fp[0] is a 4x bit size reduction, and multiplication is O(n^2) is silicon transistor count. [0] If google is smart, they'd ditch +/- infinity and if they were ballsy, they'd ditch zero in their FP implementation.
- afsina 10y agoActually for many neural network apps 8, 4 even 1 bit is sufficient for representing weights.
- kevinnk 10y agoI doubt FP hardware size is the limiting factor in their implementation (it's not in GPUs and definitely not in high perf CPUs). More likely they came up with higher level architectural tricks that let them specialize for machine learning (i.e. taking better advantage of locality in the application, etc).
- Symmetry 10y agoGenerally speaking GPUs are already very good at running with float32s, usually much better than they are at using float64s in fact. The big advantages of using an ASIC are mostly on the storage side but they also allow you to get away with non IEEE floating point numbers that don't necessarily implement subnormals, NaN, etc.
- Symmetry 10y agoThe rule of thumb I was taught was that going from a DSP/GPU to a custom ASIC would give you a 10X advantage in performance/power which is pretty close to this. And look at how much bitcoin mining ASICs out compete GPUs.
- BooneJS 10y agoFrom the article: "TPU is tailored to machine learning applications, allowing the chip to be more tolerant of reduced computational precision, which means it requires fewer transistors per operation."
- euyyn 10y agoDo you reckon that means it's using small floats?
- protomok 10y agoBased on recent blog posts from some Google folks regarding quantizing neural nets I'm going to guess 8-bit fixed point. For example -> https://petewarden.com/2016/05/03/how-to-quantize-neural-networks-with-tensorflow https://petewarden.com/2016/05/03/how-to-quantize-neural-net...
- pygy_ 10y agoOr small ints, possibly working in log space.
- nhaehnle 10y agoProbably. In addition, there's a lot of literature on optimizing hardware implementations of fundamental arithmetic operations like addition and multiplication. I recall seeing a paper a while ago which talked about reducing the number of gates by allowing some bounded imprecision in the results - unfortunately, I don't remember the title right now, but it sounds like that's what they may be doing.
- tcarey83 10y agoWhen I first read the blog post earlier in the day, it actually said they were using 8-bits. I remember it because it seem quite small to me.
- sputknick 10y agothese are ASICs, Application Specific Integrated Circuits, emphasis on the SPECIFIC. It's a chip built specifically for Tensor Flow. Anytime you build a chip to handle a specific application you are going to see a significant performance improvement. You can move into a new apartment using a Honda Civic, but you are going to see considerable performance improvement using a vehicle designed specifically for moving.
- j1vms 10y agoThe article confirms that "AlphaGo was powered by TPUs in the matches against Go world champion, Lee Sedol, enabling it to "think" much faster and look farther ahead between moves."
- CydeWeys 10y agoIt seems entirely reasonable to me, and there's good historical precedent for it. Let's look at SHA256 hashing as an example. The maximum number of hashes that the best GPU around can do is around 1 GHash/s. However, for the same cost, of around $600, specialized hardware can do around 5 THash/s. That's about five thousand times the performance/price. There's no reason that hardware that is super-specific to neural network computation can't similarly have large gains. Link to the specialized hardware for SHA256 hashing: http://www.amazon.com/Antminer-~4-73TH-25W-Bitcoin-Miner/dp/B014OGCP6W http://www.amazon.com/Antminer-~4-73TH-25W-Bitcoin-Miner/dp/...