13 ms·
Cloud TPUs in Beta
- Talyen42 9y agoHow does this compare to Nvidia GPUs on AWS price/perf-wise? The article makes it sound like this is a new thing...
- bloudermilk 9y agoGoogle claims[0] the TPU is many times faster for the workloads they've designed it for. > On our production AI workloads that utilize neural network inference, the TPU is 15x to 30x faster than contemporary GPUs and CPUs. As far as I know this will be the first opportunity for the public to prove those claims, as until now they've not been available on GCP. I don't mean to sound skeptical–I'm quite confident they're not exaggerating. [0]: https://cloudplatform.googleblog.com/2017/04/quantifying-the-performance-of-the-TPU-our-first-machine-learning-chip.html https://cloudplatform.googleblog.com/2017/04/quantifying-the...
- phamilton 9y agoI wonder how these would compare with Amazon's FPGA instances with a comparable core running.
- deleted 9y ago[deleted]
- indescions_2018 9y agoThe reserve TPU button has been available on the dashboard for the last few months. But I assume instances have been prioritized for large customers such as Two Sigma. From the paper: "Despite low utilization for some applications, the TPU is on average about 15X - 30X faster than its contemporary GPU or CPU, with TOPS/Watt about 30X - 80X higher. Moreover, using the GPU's GDDR5 memory in the TPU would triple achieved TOPS and raise TOPS/Watt to nearly 70X the GPU and 200X the CPU." In-Datacenter Performance Analysis of a Tensor Processing Unit https://arxiv.org/abs/1704.04760 https://arxiv.org/abs/1704.04760 Price is about 5x cloud nvidia gpu instance on an hourly basis.
- twtw 9y agoIt will be interesting to see some benchmarks that compare TPUs to V100, since all previously published comparisons from Google compare TPU to K80 (3 GPU architectures ago).
- xiphias 9y agoOne thing to keep in mind that Google was using Tensorflow for comparision, which is heavily optimized for TPU, and GPU is just a second class citizen. Of course it was a great strategy on Google's side, and TPUs perform better than GPUs, but this is a little bit of cheating in the benchmarks.
- vomjom 9y agoKeep in mind that what you linked refers to TPUv1, which is built for quantized 8-bit inference. The TPUv2, which was announced in this blog post, is for general purpose training and uses 32-bit weights, activations, and gradients. It will have very different performance characteristics.
- bloudermilk 9y agoThanks for pointing that out!
- ebikelaw 9y agoPerf per watt matters to Google but not you. You should only think of it on a perf/$ basis, right?
- gh02t 9y agoThey're closely related though, since if the perf per watt is lower then Google can charge you less doller per perf. The price they charge you is ultimately tied to the operating cost.
- headmelted 9y agoI would imagine that (by design) they're not directly comparable. I suspect that we'll see more information about the ASICs over time, but it'll take time to really understand their characteristics vs a Nvidia GPU - which are at least right now a bit better understood.
- perfmode 9y agoIt is. TPUs perform calculations on weights using low-precision floating point and integer types. This saves a ton of computation, but doesn't matter much for training models.
- bicubic 9y agoBut GPU is also able to use lower resolution types. There must be more to the TPU advantage.
- Cyph0n 9y agoGPUs are much more complex (general-purpose) and therefore cannot be optimized beyond a certain point due to timing requirements and PVT (process, temperature, voltage) variations. In other words, the more stuff you have on an ASIC, the more careful you have to be ensure a margin of tolerance for variations.
- saguro 9y agoSo the only advantage of the TPU is it's a simpler and more specialized asic? Google didn't break any new ground in terms of training perf?
- Cyph0n 9y ago> So the only advantage of the TPU is it's a simpler and more specialized asic And everything that entails: lower energy consumption, higher throughput, lower cost at volume, higher profits for GCP, etc. > Google didn't break any new ground in terms of training perf? Relative to GPUs, sure, but I can't say how well they stack up against other custom ASICs for DL applications.
- dgacmu 9y agoSo, a way to think of this is: The speed (and therefore, cost) of training a machine learning model depends on (a) the ML techniques (how rapidly the model converges and to what accuracy); and (b) how quickly the processor executes the operations involved in the ML techniques. The TPU is only an improvement in (b). It's not going to result in a big-O style speedup, because the same training algorithms and architectures will run on it that we run on CPUs & GPUs today. I'm not sure what counts as "breaking new ground" - is that 10%? 100%? 1000? :-) The things to watch out for in benchmarks will be: (a) Perf/$. This is actually a big deal - one of my students recently blew through $5000 of Google Cloud credits running Imagenet experiments, in a week. And we didn't finish them! As this cost really drops, it enables things like Neural Architecture Search, which uses tons of compute capability to explore architectural variants automatically. (b) Absolute perf. (c) Performance scaling. To what degree will the fast, 2D torroidal mesh allow a full pod of Cloud TPUs to scale nearly-linearly? Absolute training times matter from a user productivity standpoint. Waiting 30 minutes for a result is very different from waiting 12 hours (you can do one of these while you sneak out to go running! :-). The NIPS'17 slides have more technical context for some of this: https://supercomputersfordl2017.github.io/Presentations/ImageNetNewMNIST.pdf https://supercomputersfordl2017.github.io/Presentations/Imag...
- briffle 9y agoThis is a new thing. Google also has Nvidia GPUs. these are new custom designed ASICs google has designed for certain ML tasks.
- jakozaur 9y agoThat may be Google Cloud competitive edge for AI startups. Both in terms of development cycle and cost efficiency. Hard to replicate by competitors: AWS and Azure.
- minimaxir 9y agoThat $6.50/hr rate might be the big deal here. Amazon does offer instances with a V100 GPU (https://aws.amazon.com/ec2/pricing/on-demand/ https://aws.amazon.com/ec2/pricing/on-demand/, the P3 instances), but if you're training something like ImageNet, you'll want the biggest image (p3.16xlarge) at $24.48/hr. Attaching a VM of similar power to a TPU on Google Compute Engine is much cheaper (https://cloud.google.com/compute/pricing https://cloud.google.com/compute/pricing, n1-highmem-64, +$3.78/hr to the TPU cost for $10.28/hr total). Per recent benchmarks for training ImageNet (https://dawn.cs.stanford.edu/benchmark/ https://dawn.cs.stanford.edu/benchmark/), training ImageNet on a p3.16xlarge cost $358, when this post claims it'll cost less than $200. (EDIT: never mind; the benchmark uses ImageNet-152, and Google compares TPU performance against ImageNet-50) Interesting.
- twtw 9y agoWhy does it make any sense to compare the price/hour for a single TPU (4 ASICs) to the price/hour for p3.16xlarge, which has 8x V100? Also, that benchmark cost of $358 is for Resnet-152, not Resnet-50.
- minimaxir 9y agoWhoops, I misread, added edit.
- douglasfshearer 9y agoP3.16x benchmark is ResNet-152, TPU cost of $200 was for ResNet-50. Tensorflow benchmarks show ResNet-152 resulting in 2.4x lower throughput than ResNet-50. [0] [0] https://www.tensorflow.org/performance/benchmarks https://www.tensorflow.org/performance/benchmarks
- Cyph0n 9y agoA better comparison would be the f1.16xlarge[1] instance @ ~$4/hr. It comes with 8 FPGAs (12 Gbps link) and 64 vCPUs. [1]: https://aws.amazon.com/ec2/instance-types/f1/ https://aws.amazon.com/ec2/instance-types/f1/ Edit: I'm genuinely curious about why this comment is getting downvotes.
- tejasmanohar 9y agoThis is exciting. There are lots of specific reasons to choose Google Cloud over AWS (and vice versa), but proprietary hardware is surely an advantage that is going to be hard to replicate / compete with. If TPUs hold up to the hype, GCloud may become the de facto for ML/AI startups.
- brd 9y agoHaving had the chance to attend a fireside chat with leadership from Google and SAP, I get the sense that the hype is likely to hold up. There are a lot of big bets happening in the Enterprise space around this notion of efficient, easy to implement ML.
- karpodiem 9y agoCan you describe a line of business function that makes novel use of ML?
- 52-6F-62 9y agoFrom a media standpoint: frontline comment moderation. It would take a lot of the legwork out of filtering for advertisements, uncivil discussion, attacks, off topic posts, and trolling. I believe NYT does this already, but using minimal oversight to prevent any edge case misses or false positives. Presently there’s not much in the way of suitable options for large media that build their modules in house. At the same time media tends to prefer to not invest too heavily in hardware if they don’t have to. Convincing leadership of using a cloud service to train an AI/ML model sounds leaner and lets them tick off even more buzzwords for the executive, etc. That said, results from efforts in the aforementioned application sound promising.
- obmelvin 9y agoFor those who want to read more: https://www.nytimes.com/2017/06/13/insider/have-a-comment-leave-a-comment.html https://www.nytimes.com/2017/06/13/insider/have-a-comment-le... [not particularly techincal, but given the GP seemed to be skeptical about real world use I think this is still appropriate]
- deleted 9y ago[deleted]
- a_imho 9y agotensor processing unit https://en.wikipedia.org/wiki/Tensor_processing_unit https://en.wikipedia.org/wiki/Tensor_processing_unit
- twtw 9y agoSome things: A "single TPU" is 4 ASICs. It is not clear if it makes sense to compare a "single TPU" to a "single GPU." As a point of reference, NVIDIA's numbers are 6 hours for Resnet-50 on Imagenet when training with 8xV100. From a naive extrapolation, 4xV100 would probably take ~12 hours and 1xV100 about two days. Google has previously only compared TPUs to K80, so it will be interesting to see some benchmarks that compare TPUs to more recent GPUs. K80 was released in 2014, and the Kepler architecture was introduced in 2012.
- jacksmith21006 9y agoThe comparison was the first generation TPUs not the second generation which is what these are. But ultimately it comes down to the cost to complete some amount of work. Google also offers Nvidia GPUs in their cloud for training and should be able to compare the cost of using one over the other as both are supported by TF. That is the ultimate guide on how good or not good the TPUs really are.
- p1esk 9y agoOn 4x1080Ti it takes 2 days to train ResNet-50. 4xASICs doing it in a day is not that impressive.
- jlebar 9y ago> A "single TPU" is 4 ASICs. It is not clear if it makes sense to compare a "single TPU" to a "single GPU." Why does the number of chips matter? Put another way, suppose Google tomorrow announced Cloud TPU v3 which was one ASIC identical in all ways to four v2 ASICs glued together. Would that be notable in any way? Seems like it would be a nop to me. I think what matters is, how fast can you train a model, and at what cost? Doesn't really matter if it's one chip or 10,000 behind the scenes.
- twtw 9y agoIt doesn't matter in the ways you are considering. The ultimate comparisons are going to be time, cost, and power to complete some benchmark, just as you say. I only mention the number of chips because loads of people are comparing the "single TPU" to a single V100 with the assumption that it is meaningful. I don't know the TDP, die size, etc. of the TPUv2 chip, so it may well make more sense for ballpark comparisons to compare "single TPU" to 4xV100. For example, a "single TPU" has 64 GB of memory, whereas a "single GPU" has 16 GB (V100). Is this meaningful? I don't know. It just seems like something worth noting. I could buy a DGX1-V with 8xV100, rebrand it as the TWTW TPU, and then go around and tell everyone how my TPU is 8x faster than GPUs. It appears that everyone is normalizing by marketing unit until benchmarks come out, which is potentially flawed.
- yazr 9y agoWhat are the chances of TensorFlow code gradually optimizing for TPUs over GPUs?! (Yes TF is OSS, but realistically Google is putting much more resources into it)
- adyavanapalli 9y agoThere is already preliminary support for TPU devices in the TF API.
- dgacmu 9y agoVery low. A lot of the performance on GPUs comes from Nvidia's optimizations in CuDNN -- it's mostly a matter of making sure TensorFlow feeds the right formats/etc. to CuDNN for core NN ops. TF should run well on CPUs, GPUs, TPUs, and likely future embedded accelerators (via tensorflow lite, which already supports the Android Neural Networks API). (I'm part time on Brain, but, of course, this isn't some kind of Official Statement(tm)).
- DannyBee 9y agoTF funds one of my teams explicitly just to optimize CPUs and GPUs. Every discussion i've had with them tells me they care about making customers succeed, period. So i'm going to with "pretty low".
- gcp 9y agoInterestingly, GCP now appears to be available to individuals in Europe. It wasn't like that before, no idea when that policy got changed. Before, GCP wasn't even a consideration compared to AWS (which always handled that).
- kyrra 9y agoMore details: https://cloud.google.com/billing/docs/resources/vat-overview https://cloud.google.com/billing/docs/resources/vat-overview
- gcp 9y ago"You can’t change the tax status of your Google Cloud Platform billing account." I think this is what tripped me up before. I closed my business years ago but it was completely impossible to get Google to fix this. Now it fixed it "by itself". Just a warning to everyone before signing up with your main Google account :-)
- ramshanker 9y agoSomeone at Dell/HPE headquarter - When can we start selling "Integrated TPU" machines. ;) Google aspiring to be leader in Cloud machine learning. Let's do On Premise.
- mobileexpert 9y agoI assume Azure and AWS have some buddying up with Intel/Nervana and Nvidia counterstroke to Google TPUs. I can’t quite imagine what it will be though.
- jacksmith21006 9y agoAmazon announced today they are working on their own TPU type chips.
- thousandx 9y agoDo you have a link for that?
- jacksmith21006 9y ago"Amazon is reportedly following Apple and Google by designing custom AI chips for Alexa" https://www.theverge.com/2018/2/12/17004734/amazon-custom-alexa-echo-ai-chips-smart-speaker https://www.theverge.com/2018/2/12/17004734/amazon-custom-al...
- otterley 9y agoI'm puzzled by the phrase "differentiated performance per dollar." Is it more performant, or less? If it's less performant, why mention it at all? If it's more performant, why not simply say "better performance per dollar"?
- surajrmal 9y agoIt is both more performant overall as well as per dollar.
- deleted 9y ago[deleted]
- aw4y 9y agodoes anyone thought about cryptocurrency mining?
- polskibus 9y agoWhen is off the shelf edition coming?
- DannyBee 9y agoMy guess: Never
- polskibus 9y agoI hope that's not true, for the sake of progress. Todays clouds wouldn't have happened if AMD and Intel had restricted cloud use of their processors.
- DannyBee 9y agoAmong other things, it would be expensive (in a ton of ways), a digression, require providing direct end user support in a way they aren't good at. It also would have significant export restrictions: Neural network related asics are very tightly export controlled: https://www.bis.doc.gov/index.php/forms-documents/pdfs/1245-category-3/file https://www.bis.doc.gov/index.php/forms-documents/pdfs/1245-... (search for neural network) My 2c: It would be an expensive waste of time for Google :) Though certainly, not gonna disagree it would be cool for the sake of progress.
- singularity2001 9y agoif there ever was a chance for a hardware start-up to become the next big thing it's entering this space. unfortunately Nervana sold out to intel.
- danjoc 9y agoReading the TOS it seems like this is a really great deal for Google: "When you upload, submit, store, send or receive content to or through our Services, you give Google (and those we work with) a worldwide license to use, host, store, reproduce, modify, create derivative works (such as those resulting from translations, adaptations or other changes we make so that your content works better with our Services), communicate, publish, publicly perform, publicly display and distribute such content. The rights you grant in this license are for the limited purpose of operating, promoting, and improving our Services, and to develop new ones." All your training data are belong to us. We can use your models to improve ours. The terms will prevent me from using it. I can't grant Google permission to redistribute HIPAA PHI.
- deleted 9y ago[deleted]
- barrus 9y agoCloud TPU product manager here. The TOS you are quoting only refers to the information you provide in the survey. Here are the Google Cloud TOS: https://cloud.google.com/terms/ https://cloud.google.com/terms/ if you're interested in what Cloud does with customers data. 5.2 Use of Customer Data. Google will not access or use Customer Data, except as necessary to provide the Services to Customer. Your training data and models are secure.
- danjoc 9y agoThis URL isn't on the TPU beta signup page. The Google TOS is. Perhaps you can see the confusion? I would be reluctant to trust random 37 karma guy on Hacker News message board on this particularly important consideration.
- whataretensors 9y agoI don't like it. Google is mixing too many things. No way to buy a TPU. No competition from other cloud providers. Proprietary hardware and vendor lock-in.
- puzzle 9y agoThis is really Tensorflow as a service. You get an IP address and a port you send gRPC requests to: https://github.com/tensorflow/tpu/blob/master/tools/diagnostics/diagnostics.py#L103 https://github.com/tensorflow/tpu/blob/master/tools/diagnost... Presumably, there's a whole server behind that address that has all the right drivers and libraries: details you don't need to care about. The only partial lock-in is that not all ops are supported and you need to figure if there are any parts of the graph in the critical part that will run on the CPU instead. There's a tool for that: https://cloud.google.com/tpu/docs/cloud-tpu-tools#tpu_compatibility_checker https://cloud.google.com/tpu/docs/cloud-tpu-tools#tpu_compat... Competitors could launch something similar that uses GPUs tomorrow. Now, if you don't already use TF and don't want to switch, that's another story.
- whataretensors 9y agoThat's my point. Competitors are largely moated out by high costs of TPU production and proprietary drivers.
- puzzle 9y agoWhy are proprietary drivers a blocker? As long as you expose the same gRPC interface, your customers don't need to know what happens behind the scenes. You could have an FPGA or a Beowulf cluster of Raspberry Pis hiding.
- deleted 9y ago[deleted]
- whataretensors 9y agoI should clarify. I like all the individual pieces(hardware, cloud services, grpc interface) I just wish you could opt into them independently.
- bufferoverflow 9y agoWe really need a standard easy to run benchmark.
- boulos 9y agoDisclosure: I work on Google Cloud. I want to highlight this paragraph from the post: > Here at Google Cloud, we want to provide customers with the best cloud for every ML workload and will offer a variety of high-performance CPUs (including Intel Skylake) and GPUs (including NVIDIA’s Tesla V100) alongside Cloud TPUs. We fundamentally want Google Cloud to be the best place to do computing. That includes AI/ML and so you’ll see us both invest in our own hardware, as well as provide the latest CPUs, GPUs, and so on. Don’t take this announcement as “Google is going to start excluding GPUs”, but rather that we’re adding an option that we’ve found internally to be an excellent balance of time-to-trained-model and cost. We’re still happily buying GPUs to offer to our Cloud customers, and as I said elsewhere the V100 is a great chip. All of this competition in hardware is great for folks who want to see ML progress in the years to come.
- riku_iki 9y ago> high-performance CPUs (including Intel Skylake) Any plans for ryzen?
- boulos 9y agoWe’re always exploring the best hardware for the dollar. We’re a founding member of OpenPOWER and to your question about AMD parts, we’ve previously (publicly) run Opterons when they were the best choice. At this time, we don’t have any announcements to make :). But I’d like to note that even if we were to use parts internally at Google (or not!), that for Cloud what matters is market demand. If there really was enormous customer demand for say ARM64, then we would look into it, even if the rest of Google wasn’t interested.
- singularity2001 9y agohow much did NVIDIA pay your boss to make that statement?
- sabalaba 9y agoAny plans to support AMD GPUs and the Radeon Open Compute project? The AI/ML community really needs viable alternatives to NVIDIA, otherwise they will continue to flex pricing power. Google, via TensorFlow, is in a phenomenal position to promote open source alternatives to the proprietary Deep Learning software ecosystem that we see today with CUDA/CuDNN.
- tempay 9y agoIs there any way to use these for applications other than tensorflow/machine learning?
- deleted 9y ago[deleted]
- bcheung 9y agoThis seems a bit pricey compared to other offerings. Wouldn't an ASIC make things more economical? Seems like in terms of cost per performance, both AWS P3 spot instances and Paperspace v100 offerings are more economical. Are these prices expected to become more competitive once it is out of beta?
- make3 9y agoisn't the tpu kind of a deep learning asic?
- utopcell 9y agoGame-changer.
- tveita 9y agoIs this just go-faster-juice for Tensorflow code or does it have other implications? If you train on TPUs can you still run the model efficiently elsewhere?