7 ms·
Quantifying the performance of the TPU, our first machine learning chip
- bsamuels 10y agoso basically they're ASICs? Would love some tech details, but it seems that the paper wont be published until 5pm today
- monocasa 10y agohttps://drive.google.com/file/d/0Bx4hafXDDq2EMzRNcy1vSUxtcEk/view https://drive.google.com/file/d/0Bx4hafXDDq2EMzRNcy1vSUxtcEk...
- Cyph0n 10y agoI wonder why they didn't take the time to typeset it properly using, e.g., LaTeX. I'm hoping that the arXiv version will look more professional.
- revelation 10y agoThat's like saying "so basically it's a chip". Well yes, it is a chip.
- Cyph0n 10y agoNo, an integrated circuit (or chip) != an ASIC. An ASIC is a chip, but it is usually designed for a specific application, hence the name.
- dom0 10y agoMost accelerator-esque chips are often considered some kind of ASIC. CPUs are not ASICs because ??? ... never heard a good reason there. It's a fuzzy (but not fluffy) term.
- Cyph0n 10y agoBecause CPUs are not application-specific? ASICs are typically designed from the ground-up to do a specific thing really well.
- dom0 10y agoI'll bite: In which way is a computation engine (CPU [not ASIC], GPU [ASIC], Xeon Phi [ASIC]) more general purpose than a computation engine for AI [ASIC], GPU [ASIC] or Xeon Phi [ASIC, even though some of these are usable as a host processor?] For other parts it's even less clear: Flash or hard drive controllers are pretty normal micros with some dedicated hardware bolted on - clearly ASIC, but most of it was not designed for the "AS" part, and you could just ignore the flash and SATA interfaces and use it as a regular micro. So the distinction, if any, can't be about volume (since a lot of them are large volume parts), nor about functionality, but some fuzzy distinction by narrowness of intended use of the part (- but then again, GPUs). And how does it apply to other domains of chips? Is a TDA7000 an ASIC? :)
- Cyph0n 10y agoFirstly, on what basis are the Xeon Phi and GPUs ASICs? Secondly, I'd argue that a GPU is definitely more general-purpose than the TPU. Just look at the block diagram shown in the linked paper: a GPU is orders of magnitude more complex than that! The added complexity is a result of a GPU having to support a wider variety of workloads. As for Flash controllers etc., if they incorporate a MCU or CPU, then they are simply not ASICs? I believe that the term ASIC itself is quite overloaded. From what I've seen, many people use the term ASIC to refer to any IC that is not reconfigurable (i.e., FPGA). Going by that definition, an ASIC is any circuit that is custom designed and fabbed on a wafer. Naturally, this would include a CPU, a GPU, and whatever else you can think of. The way I like to think of it is that an ASIC is a circuit designed to perform a specific task as efficiently as possible. Note that I used to word circuit; in other words, an ASIC could be part of some larger design. Some examples off the top of my head: - Digital camera CMOS sensor - Video decryption chip (e.g., in a cable box) - Active noise cancellation chip (if custom and not a DSP) - Full-custom TPM - Full-custom RSA-2048 engine - High-performance Ethernet switch controller - CPU cache controller (ASIC that is part of a CPU)
- dicroce 10y agoIs this device optimized for forward passes or backward passes or both? It seems to me that Google engineers could use Tesla's or other high end GPU's for training and development, but then deploy those models on hardware optimized for forward passes...
- pcmonk 10y agoIf "forward passes" means inference (as opposed to training), then the post says the first generation targets that. I don't think they say anything about any future generations (other than they're working on them).
- alfalfasprout 10y agoI mean, there's no real reason they shouldn't be able to do a backwards pass assuming they're using trivially differentiable activation functions.
- wyldfire 10y agoFrom the paper: > if the TPU were revised to have the same memory system as the K80 GPU, it would be about 30X - 50X faster than the GPU and CPU. Is it "hard" to interface with GDDR5/HBM? Layout challenges? Or do they need the capacity more than the speed? Why wouldn't they have used faster memory than DDR3?
- revelation 10y agoWell HBM requires integrating the memory into the chip, presumably they didn't have the manufacturing capability or didn't want the spend required on a first version. No idea on GDDR5 vs DDR3, maybe they didn't like the latency of the first.
- dom0 10y agoMemory controllers are not so simple to do, and fast MCs also eat quite a bit of power. So a simple reason that they didn't do it could be either a) they did not want to license a more expensive, faster design b) while it would be faster, it would decrease efficiency to a point that did not meet their goals (for data centers, efficiency > absolute performance, within reasonable boundaries) c) like b) just with cost of memory d) GDDR5 and DDR3/4 have different design trade-offs. The former is optimized for sequential bandwidth (and low capacity; GDDR always was a point-to-point memory bus just to achieve the clock speeds), while DDR3/4 takes random read/write workloads into account (eh... to the amount possible with DRAM...) -- HBM requires a silicon interposer, which is basically like another complete chip (just without the FEOL parts, "just" metallization), that has to be significantly larger than the size of all chips combined. So unless you really need that performance or have a volume product it's unlikely to be a good deal.
- dom0 10y agoLooking at their block diagrams: They have two large "cache-like" structures: The unified buffer (24 MiB) and the accumulators (4 MiB). Bandwidth between these and the matrix multiply unit is high (167 GiB/s), bandwidth out of that complex is low. So it would seem that they just don't need a very large bandwidth out of that function complex.
- 10y ago
- iandanforth 10y ago"This first generation of TPUs targeted inference ..." Makes me wonder if there are more recent generations that target training.
- joe_the_user 10y agoSo since "sharing the benefits with everyone" could involve just allowing people to rent time on the Google cloud, we can still ask when/if the chips themselves will ever be available for purchase?
- pc2g4d 10y agoMaybe it's just me misunderstanding, but to me "inference" and "training" are one and the same. But the article defined it thus: This first generation of TPUs targeted inference (the use of an already trained model, as opposed to the training phase of a model, which has somewhat different characteristics) This Nvidia article treats them differently, too: https://blogs.nvidia.com/blog/2016/08/22/difference-deep-learning-training-inference-ai/ https://blogs.nvidia.com/blog/2016/08/22/difference-deep-lea... But the definition of "statistical inference" on Wikipedia says "Statistical inference is the process of deducing properties of an underlying distribution by analysis of data" which seems exactly like training.
- mjn 10y agoSome parts of stats (esp. classically) do use "inference" for the whole process of going from data -> result, especially when doing descriptive rather than predictive statistics. In most of ML the process is split into two phases, training a model on a data set, followed by using the trained model to predict labels (or whatever else it's predicting) for new data. Training is an implementation of statistical induction, from data to model, while model "use" or "evaluation" is an implementation of (probabilistic) deduction, from model + query to result. In ML, "inference" is usually a synonym for the "deduction" or "evaluation" portion. It makes sense to me intuitively if you think of it as learning things vs. inferring things from learned knowledge. But you do find constructions using it as a synonym for training too, as in phrases like "model inference" (which means inferring models from data, aka model induction or training). Inferring things from other things is a pretty general concept, so it can be slippery without context...
- DanWaterworth 10y agoInference in this context means getting the output of a neural network on some input. Training meaning adjusting a neural network to so that it's outputs are more like the training data. Training involves finding derivatives and requires higher precision calculations which is why they don't train using TPUs.
- kevinnk 10y agoTraining is inference, except over model space instead of possible outputs. Depending on what models you're looking at, the two can have very different computational properties, so treating them separately is not surprising even if what they're doing is conceptually similar.
- deleted 10y ago[deleted]
- wangqufei 10y agoThis is a very very bad idea. The so called AI is changing, far from being stable. Software can change, hardware can not.