11 ms·
Google announces a new generation for its TPU machine learning hardware
- polskibus 8y agoIs this going to be available off-the-shelf like NVidia GPUs? That's the only way to get wider and faster developer buy-in.
- dejv 8y agoThey did denied selling older generation, so I would expect TPU 3 is not going to be available as well. The only way to use those is to use Google Cloud.
- aseipp 8y agoNo, they're going to be used by Google and nobody else, so they can establish dominance in the cloud AI/AI SaaS space using their immense resources. This will be used by internal Google projects to rapidly develop their large scale models and, if you're lucky, available on Google Cloud one day (TPUv2 already is, at least). Google doesn't need "developer buy in" for these to make sense -- they need better hardware for training and deploying deep learning models for their products, which is the overwhelming motivation. And, if they offer these to you, it's only because you're willing to pay for faster TensorFlow iteration.
- option_greek 8y agoModels are the new bitcoins and no one is willing to sell the hardware directly. Weirdly enough exact same thing happened in mining (specialized hardware passed nvidia gpus in performance at some point).
- jacksmith21006 8y agoWould doubt it and not really make much sense in 2018.
- shaklee3 8y agoNo sense at all. Nvidia's stock increased 14x since 2014 by running a closed system in a public cloud. Oh wait...
- jacksmith21006 8y agoNot sure what it has to do with Nvidia stock. The future is cloud providers using their own silicon. Google has set the bar where others will have to match. Also looks like Google is now 2 generations ahead of Nvidia.
- shaklee3 8y agoNvidia is on their 5th generation of datacenter ASIC. Google is on their third. Nvidia has a faster interconnect than anyone else (nvswitch), which allows close to the same access bandwidth to HBM on a remote GPU than a local GPU. If Google has that, they haven't announced it, and all I can see is they operate over PCIe v3 since they are slaves to the CPU manufacturers. If the future was cloud providers using their own silicon, Google would have made a CPU to compete with Intel. Instead, they use Intel exclusively and are evaluating POWER. They made a single ASIC that is very simple relatively speaking, that can perform a single task and cannot be accessed generically. They are not 2 generations ahead. They chose a path that's much easier for many reasons, but mostly because it allows their own internal training to be accelerated.
- jacksmith21006 8y agoAll comes down to the cost. We can see that the TPUs were about 1/2 as much as using Nvidia before the TPU 3.0. This is based on doing the same task with the TPUs versus Nvidia with AWS. But it might be a bigger difference and Google taking margins. Could have been the other way but think WaveNet proves not that direction. Would hope we get an even bigger price difference with this generation. BTW, you are mixing up graphic generations and ML generations. Google never did or will graphics. For ML Google is 2 generations ahead. Hard to imagine Nvidia ever catching up as so much of the AI breakthroughs are coming from Google. A perfect example is WaveNet. Google is using a NN for audio in real time at 16k cycles a second. That is competing with using the old way of doing this. The Google approach gets a better result but people are only willing to pay too much more for a better result. So Google had to price WaveNet competitively to the old approach which they have. I suspect that would have been impossible using Nvidia. The cost in running would have been just too high. The computation required for WaveNet is huge and the old way it is minimal. Also we are going to need to really separate training with inference. They have different requirements and the costs in running are different. Has Nvidia done an inference only solution? Or do they still always mix them? "Google would have made a CPU to compete with Intel. " I do a lot of surfing and have to say one of the more silly things to read. There is little gains today with CPUs. Why on earth would you invest in doing your own? Google has ported all their stuff to Power so they have a backup and not reliant on Intel. But CPUs are NOT the future and would be crazy to invest into them at this point. Processing is moving to TPU type processors and we will see more and more traditional things. A fantastic paper on this from Jeff Dean at Google I suggest you read. https://arxiv.org/abs/1712.01208 https://arxiv.org/abs/1712.01208 One thing I love about Google is they do NOT reinvent the wheel. If there is something that can work like the Linux kernel they use. Versus having an ego of having to do yourself. They did all their own network silicon by hiring the Lanai team several years ago because there was nothing that could work for their needs. They then created a determinate network stack so they could do Spanner. There was nothing off the shelve. Heck they laid their own fiber under the ocean to make possible. Spanner is the first and only horizontally scalable RDBMS solution to exist. Google beat the speed of light by using their custom silicon, custom stack that is determinate, their own fiber. Then using atomic clocks with GPS to remove the latency from speed of light out of the equation. Highly recommend the papers including the Spanner paper. https://ai.google/research/pubs/pub39966 https://ai.google/research/pubs/pub39966 The network uses custom silicon inside and then they put the intelligence on out edge. Then they control exactly the traffic so they get a determinate result. This also allows them to use far cheaper hardware as do not need to over provision as they never drop packets on the ground or need much memory for buffers. Google would have used Nvidia in a second if they could do what they need. But what is being done in AI (ML) today requires unique silicon. Just not possible without as the power footprint would be prohibitive. I do expect Google more and more to leverage on client ML where needed with their PVC. So it is NOT just the TPUs but for somethings you are going to need on client. Take the new voices as an example. They are doing 6 voices as each has a different model. The cost in switching models is prohibitive in the cloud. So you will see them eventually move to some things being done on the client. But they are so far ahead they have tons of time. I would not be surprised if we do not seen anyone else do Wavenet for a couple of years. What Nvidia should be working on is finding a way to do WaveNet at a reasonable cost for Amazon and others. But the long term is the big cloud providers will do their own silicon. Google set the path with the TPUs and now the others will have to do the same to keep up. But just makes sense as the cloud providers have the data to iterate. MS is using a FPGA which was foolish, IMO. Never going to get there with that route. BTW, you need to realize how this works for inference. You keep the model in memory. Google invented Capsule networks that use dynamic routing and are going to cause different patterns in memory access and would be curious if the TPU 3.0 were optimized for such? They have a huge advantage of inventing the algorithms and then get to make their own silicon optimized. Something Nvidia really needs to be able to better compete, IMO. BTW2, Google using Power did include the excellent Nvidia memory interconnect as part of the spec. Not sure did Google license or does Nvidia give away? The spec is called Open Power if not aware? https://www.theregister.co.uk/2016/04/07/open_power_summit_power9/ https://www.theregister.co.uk/2016/04/07/open_power_summit_p...
- thwd 8y agoIt's 'Tensor Processing Unit' as far as I know.
- tlb 8y agoTitle corrected from "Tensorflow Processing Unit"
- frisco 8y ago100 petaflops per what? All of google’s deployment? Per rack? 100 Pflop / chip can’t be right.
- blattimwind 8y agoPer pod (256 TPUs)
- deepnotderp 8y agoThey increased the pod size by 4X, also each "TPU" has 4 tpu chips on it IIRC.
- raverbashing 8y agoCan you allocate a whole pod to do your computations? How much does that cost?
- mlazos 8y agoYeah I love this, 8x PFlop improvement per pod, but the pod size has increased 4x. That 8x speed up is now mystically halved.
- Voloskaya 8y ago*mystically quartered.
- deleted 8y ago[deleted]
- sabalaba 8y agoThis is great news. It’s extremely important to the research community that large companies enter into the DL silicon space to contend with NVIDIA’s monopoly. NVIDIA is now exerting pricing power to the point where they’ve decided to train their sales people to disregard the metric that is most important to customers: cost of training. Talk with one of their enterprise sales people and you’ll find they’ll say things like “FLOPS / $ doesn’t matter” to justify a 10x increase in price for their TESLA line. As history has shown, a monopolist sows the seeds of their own destruction and, by disregarding the metrics that matter, they alienate their customers. Here are open source projects that you can contribute to to break the monopoly: ROCm: https://rocm.github.io/ https://rocm.github.io/ MIOpen: https://github.com/ROCmSoftwarePlatform/MIOpen https://github.com/ROCmSoftwarePlatform/MIOpen TensorFlow: https://github.com/tensorflow/tensorflow https://github.com/tensorflow/tensorflow
- twtw 8y agoInstead of being able to buy your own nvidia gpu and run any cuda or opencl on it, now you can run only tensor flow on a tpu only in Google cloud. How fantastic.
- chewxy 8y agoAnd only using Tensorflow.
- sabalaba 8y agoChoice in the marketplace is still choice. And, Google might eventually decide to sell this to outside companies. Eventually the desire to build competitive advantages for Google Cloud might be surpassed by the desire to compete with NVIDIA. Commoditizing the complement is a powerful strategy.
- throwaway2048 8y agoYeah seems like a way more abusive position to me frankly, a Google monopoly would be way worse than a NVIDIA one.
- 8y ago
- jamesblonde 8y agoWhat will be interesting to see is if they are going the hardware specialization route, like Nvidia with their support for efficient 4x4 fused FP16 matmuls with FP32 matrix output (they call them 'tensor cores' - hah). I suspect with the liquid cooling, they are just dialing up the matmul speed, which is probably the right way to go, IMO. We are doing data-center experiments - using oil cooling. It's suprising to see circuitboards dumped in ordinary oil and working away, no short circuits.
- deepnotderp 8y agoYou mean immersion cooling? As long as the fluid isn't conductive you're good to go. If you're interested, Alibaba has been doing some work along the lines of production immersion cooling.
- jamesblonde 8y agoYes, it's not my research - a colleagues. He has some big vats of oil. It's a sureal experience to see servers just being oil-boarded (can i say that? :) ) and coming out ok!
- dig1 8y agoCooling via oil insulation isn't something new - it is used for years inside high voltage transformers to keep temperature stable. What kind of oil are you using?
- polvs 8y agoI'd love to discuss your research. We've been working low-profile during the last years doing a lot of R&D and experimenting with different fluids and components. Mineral oil is ok for experimentation, but for long term material compatibility and fire risk I wouldn't recommend it. FWIW I co-founded https://submer.com https://submer.com where we've developed an all-embedded computing immersion cooling solution that is virtually compatible with any kind of hardware (even fiber optics) and it's orders of magnitude more efficient than traditional data center cooling technologies.
- Voloskaya 8y ago> Google CEO Sundar Pichai said the new TPU is eight times more powerful than last year Are we sure about this? He specifically said that a "pod" would be 8 times faster than last year, not the TPU itself. And the picture in the background showed what looked like to be 8 racks of 64 TPUs (or maybe 32?). Until now a "pod" was a single rack of 64 TPUs. So if the new definition for a "pod" is 8 times as many TPUs as it was before, the result is less impressive... Is there any actual spec released?
- twtw 8y agoGoogle started playing this game with the tpuv2, wherein they define a "TPU" as whatever they want to make it sound suitably impressive (I.e. Cloud TPU is four chips). This in turn led to nvidia calling the DGX-2 "the world's largest GPU
- jacksmith21006 8y agoWhat does it matter? What we care about is the cost. The TPU gen 2 are about half as much as using Nvidia in the Amazon cloud for the same work. To me that is what matters. If the TPU 3 further drops the cost fantastic. How many chips sounds a lot more like a pissing contest. What if it was a giant chip with actually a bunch of chips inside? Who cares?
- deepnotderp 8y agoYou care from a technical standpoint. But, if you insist on going by the cost metric, you've already lost because you can buy nVidia GPUs and that's a lot cheaper than any of the TPU instances :)
- jacksmith21006 8y agoCost is to do some task.
- p1esk 8y ago
- ai2323 8y agoWelcome to ASIC land....training going the route of crypto mining
- Maybestring 8y agoIt's just an arithmetic logic unit for tensors. It's not at all like a crypto miner that implements a single algorithm.
- jacksmith21006 8y agoExactly. Drives me crazy when people compare to a single alogrithm chip. It is not.
- deepnotderp 8y agoBut the TPU actually is: it's a systolic array matrix multiply ASIC.
- jacksmith21006 8y agoThe TPUs support all kinds of NN. Not specific to one type.
- deepnotderp 8y agoIndeed, all neural nets that use matrix multiplication. Out just so happens to be all the main types right now.
- jacksmith21006 8y agoGoogle is the leader on AI algorithms. Where GANs and Capsules and so many others came from for example. So things change and Google can implement in silicon far earlier than anyone else as needed. But right now they have what they need. BTW, I would be curious on the TPU 3 design considerations to better support dynamic routing with Capsule networks. I suspect creates different memory access patterns and where the big power savings are at today.
- esmi 8y agoDoes the Google Cloud API for TensorFlow expose any hardware details? I am just curious why Google announces these hardware details at all.
- jacksmith21006 8y agoLook at this thread. Because people want to know as just what "techies" do. But what I wanted to know is performance in terms of joules compared to TPU 2s. My hope is we get a paper on the TPU 2 now the 3 is released. We got the TPU 1 paper as they were releasing the TPU 2 I suspect.
- ai2323 8y agoThis very typical google. V100 comes out, they don't deploy it in their cloud immediately. Then they launch their tpu cloud. Spend 2 months touting cost/speed in benchmarks. Then a week before io they make v100 available in gce, nearly six months after aws. Then at io they announce tpu3. This is supposed to be nvidia's most captive customer, and now looks like going to be their biggest problem. Would love to see google spend with them and how its changed since they ramped tpu2.
- oh-kumudo 8y agoLower GPU price for all, I am for it.
- ai2323 8y agoFor all the google hate, nvidia is too blame as well. Their ceo running around saying 'we saving you money every time you spend 10k on a gpu' is getting old. The should have been far more aggressive price wise with this biz considering the resources of their target customers. Almost every 'choose not to use a gpu' use case being touted by msft/goog etc is about cost. Meanwhile nvda busy tweeting stuff like msft seeing ai app for blind powerd by nvda gpu. Soon micron will be tweeting they powering all ai as well, and then arista...all the way down to the power companies. Watching goog today its not looking good for the competition, by the time all these ml/dl accelerators are commoditized google will have one hell of a moat.
- jacksmith21006 8y agoSpend with them? Who?
- ai2323 8y agoTheir spend on nvidia gpu's...how it's shifted as they have progressed tpu wise. I think you go back 18 months google was their biggest here.
- jacksmith21006 8y ago
- zackmorris 8y agoJust to play devil's advocate for a moment: I'm excited that there will finally be some viable competition for GPUs, but am disappointed that this isn't a general-purpose multiprocessor. We're long overdue for a general-purpose CPU with say 1024 cores, that avoids a central main memory, where each core can be independently programmed just like any other CPU. Google's may count as a somewhat general-purpose DSP, which is definitely a step forward. But no matter how mature or mainstream a framework like TensorFlow gets, it can never replace full programmability. Without seeing the internals, I'm going to have to give this a nay vote for now. There are many other rather exciting problems that need to be opened up to a new generation of tinkerers. Off that top of my head, it's things like: content-addressable memory to provide high data locality (evolving various interconnects instead of hardwiring them), exploring other types of general vector processing like the kind MATLAB/Octave uses, and exploring other hill-climbing algorithms than backpropagation/neural nets. I picture something more like network topology-agnostic Docker containers programmed in Elixer/Erlang/Go that can act as semi-autonomous agents and switch into various modes in order to solve the problem at hand. I just find that a much simpler metaphor to work with than OpenCL/CUDA/TensorFlow. Yes it would take more silicon and would probably violate YAGNI, but only full programmability gives us the freedom to explore the problem space at the level that's going to be required to implement artificial general intelligence.
- tim333 8y agoThe announcement bit on video https://youtu.be/ogfYd705cRs?t=1h36m https://youtu.be/ogfYd705cRs?t=1h36m