17 ms·
Google's First Tensor Processing Unit: Architecture
- tibbydudeza 3y agoI listened to a talk by Jim Keller from Tens torrent and their different approach to making AI cores - 5 Risc V cores one core for loading data, one for uploading data and the rest dedicated to performing matrix operations. He did mention Google's TPU and the fact it was like programming a VLIW and they had about 500 people dedicated to their compiler.
- Donaldwide 3y ago[dead]
- layer8 3y ago> However, although tensors describe the relationship between arbitrary higher-dimensional arrays, in practice the TPU hardware that we will consider is designed to perform calculations associated with one and two-dimensional arrays. Or, more specifically, vector and matrix operations. I still don’t understand why the term “tensor” is used if it’s only vectors and matrices.
- ralusek 3y agoIf nothing else, the term "tensor" is shorter than "vectors and matrices," and then has the added benefit of representing n-dimensional arrays.
- layer8 3y agoHow is that an added benefit if the hardware doesn’t actually support n-dimensional arrays (other the n = 1 and 2)? And, strictly speaking, a vector can be considered a 1xn (or nx1) matrix, so Matrix Processing Unit would have been fine.
- thatguysaguy 3y agoAt the end of the day all the arrays are 1 dimensional and thinking of them as 2 dimensional is just an indexing convenience. A matrix multiply is a bunch of vector dot products in a row. Higher tensor contractions can be built out of lower-dimensional ones, so I don't think it's really fair to say the hardware doesn't support it.
- whimsicalism 3y agoit’s an abstraction, just like 2d arrays
- layer8 3y agoI’d say it’s more like calling an ALU that can perform unary and binary operations (so 1 or 2 inputs) an “array processing unit” because it’s like it can process 1- and 2-element arrays. ;)
- whimsicalism 3y agowhat? the ml framework can support n-dimensional arrays. that’s what i mean by an abstraction
- deleted 3y ago[deleted]
- necroforest 3y agoIt's branding (see: TensorFlow); also, pretty much anything (linear) you would do with an arbitrarily ranked tensor can be expressed in terms of vector ops and matmuls
- nxobject 3y ago"Fixed-Function Matrix Accelerator" just doesn't have the same buzzy ring to it.
- sillysaurusx 3y agoI was confused as hell for a long time when I first got into ML, until I figured out how to think about tensors in a visual way. You're right: fundamentally ML is about vector and matrix operations (1D and 2D). So then why are most ML programs 3D, 4D, and in a transformer sometimes up to 6D (?!) One reasonable guess is that the third dimension is time. Actually not. It turns out that time is pretty rare in ML, and it's only (relatively) recently that it's been introduced into e.g. video models. Another guess is that it's to represent "time" as in, think of how transformers work: they generate a token, then another given the previous, then a third given the first two, etc. That's a certain way of describing "time". But it turns out that transformers don't do this as a 3D or 4D dimension. It only needs to be 2D, because tokens are 1D -- if you're representing tokens over time, you get a 2D output. So even with a cutting edge model like transformers, you still only need plain old 2D matrix operations. The attention layer creates a mask, which ends up being 2D. So then why do models get to 3D and above? Usually batching. You get a certain efficiency boost when you pack a bunch of operations together. And if you pack a bunch of 2D operations together, that third dimension is the batch dimension. For images, you typically end up with 4D, with the convension N,C,H,W, which stands for "Batch, Channel, Height, Width". It can also be N,H,W,C, which is the same thing but it's packed in memory as red green blue, red green blue, etc instead of all the red pixels first, then all the green pixels, then all the blue pixels. This matters in various subtle ways. I have no idea why the batch dimension is called N, but it's probably "number of images". "Vector" wouldn't quite cover all of this, and although "tensor" is confusing, it's fine. It's the ham sandwich of naming conventions: flexible, satisfying to some, and you can make them in a bunch of different varieties. Under the hood, TPUs actually flatten 3D tensors down into 2D matrix multiplications. I was surprised by this, but it makes total sense. The native size for a TPU is 8x128 -- you can think of it a bit like the native width of a CPU, except it's 2D. So if you have a 3x4x256 tensor, it actually gets flattened out to 12x256, then the XLA black box magic figures out how to split that across a certain number of 8x128 vector registers. Note they're called "vector registers" rather than "tensor registers", which is interesting. See https://cloud.google.com/tpu/docs/performance-guide https://cloud.google.com/tpu/docs/performance-guide
- layer8 3y agoThanks for the background! I still don’t think it’s appropriate to call a batch of matrices a tensor.
- shrubble 3y agoTensor is from mathematics and was popularized over a century ago.
- layer8 3y agoI know what a tensor is mathematically. However, as far as I can see, ML isn’t based on tensor calculus as such.
- phkahler 3y agoSomething similar happens on Wikipedia, where topics that use math inevitably get explained in the highest level math possible. It makes topics harder to understand than they need to be.
- jiggawatts 3y agoAs a helpful Wiki editor just trying to make sure that we don't lead people astray, I've made some small changes to clarify your statement: In the virtual compendium of Wikipedia, an extensive repository of human knowledge, there is a discernible proclivity for the hermeneutics of mathematically-infused topics to be articulated through the prism of esoteric and sophisticated mathematical constructs, often employing a panoply of arcane lexemes and syntactic structures of Greek and Latin etymology. This phenomenon, redolent of an academic periphrasis, tends to transmute the exegesis of such subjects into a crucible of abstruse and high-order mathematical discourse. Consequently, this modus operandi obfuscates the intrinsic didactic intent, thereby precipitating an epistemological chasm that challenges the layperson's erudition and obviates the pedagogical utility of the exposition.
- xarope 3y agoscarily, I actually understood this.
- whimsicalism 3y agomultidimensional arrays are multilinear mappings, and that is how they are used in ml usually. it seems fine to me
- thatguysaguy 3y agoWell, in the transformer forward pass there are a bunch of 4-dimensional arrays being used.
- deleted 3y ago[deleted]
- smilekzs 3y agoCame in to say this. The Einsum notation makes it desirable to formulate your model/layer as multi-dimensional arrays connected by (loosely) named axes, without worrying too much about breaking it down to primitives yourself. Once you get used to it, the terseness is liberating.
- jeffhwang 3y ago(I think) technically, all of these mathematical objects are tensors of different ranks: 0. Scalar numbers are tensors of rank 0. 1. Vectors (eg velocity, acceleration in intro high school physics) are tensors of rank 1. 2. Matrices that you learn in intro linear anlgebra are tensors of rank 2. Nested arrays 1 level deep, aka a 2d array. 0. Tensors numbers are tensors of rank 3 or higher. I explain this as ‘nested arrays’ to people with programming backgrounds as nested arrays of arrays with 3dimensions of arrays or higher. But I’m mostly self-taught in math so ymmv.
- a_wild_dandan 3y agoEvery tensor is just a stack of vectors wearing a trench coat.
- WhitneyLand 3y agoIt says: tensors describe the relationship between high-d arrays It does not say: tensors “only” describe the relationship between high-d arrays The term “tensor” is used because it covers all cases: scalars, vectors, matrices, and higher-dimensional arrays. Tensors are still a generalization of vectors and matrices. Note the context: In ML and computer science, they are considered a generalization. From a strict pure math standpoint they can be considered different. As frustrating as it seems one is not really more right and context is the decider. There are lots of definitions across STEM fields that change based on the context or field they’re applied to.
- adrian_b 3y agoThe word tensor has become more ambiguous during the time. Before 1900, the use of the word tensor was consistent with its etymology, because it was used only for symmetric matrices, which correspond to affine transformations that stretch or compress a body in certain directions. The square matrix that corresponds to a general affine transformation can be decomposed into the product of a tensor (a symmetric matrix which stretches) and a versor (a rotation matrix, which is antisymmetric and which rotates). When Ricci-Curbastro and Levi-Civitta have published the first theory of what now are called tensors, they did not define any new word for the concept of a multidimensional array with certain rules of transformation when the coordinate system is changed, which is now called tensor. When Einstein has published the Theory of General Relativity during WWI in which he used what is now called tensor theory, for an unknown reason and without any explanation for this choice he has begun to use the word "tensor" with the current meaning, in contrast with all previous physics publications. Because Einstein has become extremely popular immediately after WWI, his usage of the word "tensor" has spread everywhere, including in mathematics (and including in the American translations of the works of Ricci and Levi-Civita, where the word tensor has been introduced everywhere, despite the fact that it did not exist in the original). Nevertheless, for many years the word "tensor" could not be used for arbitrary multi-dimensional arrays, but only for those which observe the tensor transformation rules with respect to coordinate changes. The use of the word "tensor" as a synonym for the word "array", like in ML/AI, is a recent phenomenon. Previously, e.g. in all early computer literature, the word "array" (or "table" in COBOL literature) was used to cover all cases, from scalars, vectors and matrices to arrays with an arbitrary number of dimensions, so no new words are necessary.
- adrian_b 3y agoI do not know which is the real origin of the fashion to use the word tensor in the context of AI/ML. Nevertheless, I have always interpreted it as a reference to the fact that the optimal method of multiplying matrices is to decompose the matrix multiplication into tensor products of vectors. The other 2 alternative methods, i.e. decomposing the matrix multiplication into scalar products of vectors or into AXPY operations on pairs of vectors, have a much worse ratio between computation operations and transfer operations. Unfortunately, most people learn in school the much less useful definition of the matrix multiplication based on scalar products of vectors, instead of its definition based on tensor products of vectors, which is the one needed in practice. The 3 possible methods for multiplying matrices correspond to the 6 possible orders for the 3 indices of the 3 nested loops that compute a matrix product.
- IncreasePosts 3y agoSigh...learning about TPUs a decade ago made me invest heavily in $GOOG for the coming AI revolution...got that one 100% wrong. +400% over 10 years isn't bad but I can't help but feel shortchanged seeing nvidia/etc
- smallmancontrov 3y ago+400% over 10 years isn't bad.
- layer8 3y agoIt’s almost 15% per year, quite a lot.
- fragmede 3y agoyeah, but nvda is up like 500% in 2 years, so if you’re naive enough to think you can time the market, you’d have fomo over having invested in the “wrong” thing.
- bongodongobob 3y agoSeeing the difference between GPT2 and GPT3 made me run to NVDA immediately. One of the few bets in my life I've ever been confident about. I think NVDA was a pretty reasonable bet on AI like 5+, maybe 10 years ago when deep learning was ramping up.
- deleted 3y ago[deleted]
- genidoi 3y agoI don't think anybody in 2014 believed that the performance of GPT-4/Claude Opus/... was 10 years away. 25 years maybe, 50 years probably, but not 10.
- bongodongobob 3y ago
- nl 3y agoOn the podcast interview now Groq CEO Jonathon Ross did[1] he talked about the creation of the original TPUs (which he built at Google). Apparently originally it was a FPGA he did in his 20% time because he sat near the team who was having inference speed issues. They got it working, then Jeff Dean did the math and the decided to do an ASIC. Now of course Google should spin off the TPU team as a separate company. It's the only credible competition NVidia has, and the software support is second only to NVidia. [1] https://open.spotify.com/episode/0V9kRgNS7Ds6zh3GjdXUAQ?si=qE3MLNuFQLCFmSnvhG87jQ&nd=1&dlsi=3752b373bdfd4f2c https://open.spotify.com/episode/0V9kRgNS7Ds6zh3GjdXUAQ?si=q...
- summerlight 3y ago> Now of course Google should spin off the TPU team as a separate company. Given the size of the market and its near-monopoly situation, I strongly think this has the potential to (almost immediately) surpass the Pixel hardware business. But the problem here is that TPU is a relatively scarce computing resource even inside Google and it's very likely that Google has a hard time to meet its internal demands...
- xrd 3y agoThis article really connected a lot of abstract pieces together into how they flow through silicon. I really enjoyed seeing the simple CISC instructions and how they basically map on to LLM inference steps.
- Donaldwide 3y ago[dead]
- rhelz 3y agoQuote from the OP: "The TPU v1 uses a CISC (Complex Instruction Set Computer) design with around only about 20 instructions." chuckle CISC/RISC has gone from astute observation, to research program, to revolutionary technology, to marketing buzzwords....and finally to being just completely meaningless sounds. I suppose it's the terminological circle of life.
- dmoy 3y agoIdk maybe it's just me, but what I was taught in comp architecture was that cisc vs risc has more to do with the complexity of the instructions, not the raw count. So TPU having a smaller number of instructions can still be a cisc if the instructions are fairly complex. Granted the last time I took any comp architecture was a grad course like 15 years ago, so my memory is pretty fuzzy (also we spent most of that semester dicking around with Itanium stuff that is beyond useless now)
- cowsandmilk 3y agoYou’re seeming to imply the number of instructions available is what distinguishes CISC, but it never has been.
- LelouBil 3y agoThe fact that it's opposed to RISC (Reduced Instruction Set) adds to the confusion.
- rhelz 3y agoGuys....what are the instructions? The on-chip memory they are talking about is essentially...a big register set. So we have load from main memory into registers, store from registers into main memory, multiply matrices--source and dest are stored in registers.... We have a 20 instruction, load-store cpu....how is this not RISC? At least RISC how we used the term in 1995?
- dmoy 3y agoI think the "multiply matrices" instruction is the one that makes it a cisc
- formercoder 3y agoGoogler here, if you haven’t looked at TPUs in a while check out the v5. They support PyTorch/JAX now, makes them much easier to use than TF only.
- tw04 3y agoWhere can I buy a TPU v5 to install in my server? If the answer is “cloud”: that’s why NVidia is wiping the floor.
- deleted 3y ago[deleted]
- ShamelessC 3y agoYou probably can't even rent them from Google if you wanted to, in my experience.
- inhumantsar 3y agohttps://cloud.google.com/tpu https://cloud.google.com/tpu
- jedberg 3y agoI think OPs point was Google claims to have TPUs in their cloud but in reality they are rarely available.
- foota 3y agoHow many people are out there buying H100s for their personal use?
- Workaccount2 3y agoProbably many orders of magnitude greater than those buying TPU's for personal use...
- kleton 3y agoWhich ocean creature name is the current TPU?
- hipadev23 3y agoHow is it that Google invented the TPU and Google Research came up with the paper on LLM and NVDA and AI startup companies have captured ~100% of the value
- beachy 3y agoFor historical precedent see Xerox Parc.
- readyplayernull 3y agoIBM, Intel, Apple's Newton.
- chillfox 3y agoKodak
- Forgeties79 3y agoMan I remember my last semester of college taking a history of photography course that was only offered every 3-4 years by a pretty legendary professor. The the day before the first day of class (or super close), Eastman Kodak declared bankruptcy after what? 110 years? He scrapped his day 1 lecture and threw together a talk - with photos of course - about Kodak and how an intrepid engineer developed then the company foolishly hid the first digital camera because it would compete with their film line. Incredible lecturer for sure haha
- klodolph 3y agoThe story I like to tell for the Newton is that it was launched before the technology was ready yet. Like the Sega Game Gear. Old video phones. All those tablets that launched before the iPad. They’re good ideas, but they shipped a few years too early, and the technology to make them work well at a good price point wasn’t available until later. Like, the Sega Game Gear had a cool active matrix LCD screen, but it took six AA batteries and the batteries only lasted like four hours.
- brcmthrowaway 3y agoBroadcom did the TPU
- HarHarVeryFunny 3y agoNot the whole design - the core processing part (systolic array - matrix multiplier) was designed by Google, but Broadcom designed all the highspeed chip I/O and mapped the design onto TSMCs tools/rules.
- tmp5120 3y ago[dead]
- uptownfunk 3y agoWhat Google really needs to do is get into the 2nm EUV space and go sub 2nm. When they have the electro lithography (or whatever tech ASML has that prints on the chips) then you have something really dangerous. Probably a hardcore Google X moonshot type project. Or maybe they have 500mm sitting around to just buy one of the machines. If their tpu are really that good - maybe it is a good business - especially if they can integrate all the way to having their own fab with their own tech
- ejiblabahaba 3y agoThis is frankly infeasible. Between the decades of trade secrets they would first need to discover, the tens- or maybe hundreds- of billions in capital needed to build their very first leading edge fab, the decade or two it would take for any such business to mature to the extent it would be functional, and the completely inconsequential volumes of devices they'd produce, they would probably be lighting half a trillion dollars on fire just to get a few years behind where the leading edge sits today, ten or more years from now. The only reason leading edge fabs are profitable today is because of decades of talent and engineering focused on producing general purpose computing devices for a wide variety of applications and customers, often with those very same customers driving innovation independently in critical focus areas (e.g. Micron with chip-on-chip HDI yield improvements, Xilinx with interdie communication fabric and multi chip substrate design). TPUs will never generate the required volumes, or attract the necessary customers, to achieve remotely profitable economies of scale, particularly when Google also has to set an attractive price against their competitors. If Google has a compelling-enough business case, existing fabs will happily allocate space for their hardware. TPU is not remotely compelling enough.
- ThinkBeat 3y agoGiven what seems to be an enormous demand for fab space, when Microsoft or Google create a proprietary chip and need it produced how do they get to the front of the line? Are they simple enough that "older outdated less in demand" fabs can produce them? I know Apple and Nvidea has a lock on a lot of fab space?
- cavisne 3y agoThey operate on outdated fab's (roughly state of the art - 1) https://en.wikipedia.org/wiki/Tensor_Processing_Unit#Products https://en.wikipedia.org/wiki/Tensor_Processing_Unit#Product... They also do have a serious presence/spend on things like HBM, semianalysis has some good pieces on this.
- ThinkBeat 3y agoThis is probably a dumb question, that just shows my ignorance but I keep hearing on the consumer end of things that the M1-M4 chips are good at some AI. The most important for me these days would be Photoshop, Resolve etc, and I have seen those run a lot faster on Apple new proprietary chips than on my older machines. That may not translate well at all to what this chip can do or what a H100s can do. But does it translate at all? Of course Apple is not selling their propritary chips either so for it to be practical Apple would have to release some from of an external, server something stuffed with their GPUs and AI chips
- singhrac 3y agoI’m also not quite an expert, but have benchmarked an M1 and various GPUs. The M* chips have unified memory and (especially Pro/Max/Ultra) have very high memory bandwidth even compared eg to a 1080 (an M1 Ultra has memory bandwidth between 2080 and 3090). At small batch sizes (including 1, like most local tasks), inference is bottlenecked by memory bandwidth, not compute ability. This is why people say the M* chips are good for ML. However H100s are used primarily for training (at enormous batch sizes) and require lots of interconnect to train large models. At that scale, arithmetic intensity is very high, and the M* chips aren’t very competitive (even if they could be networked) - they pick a different part of the Pareto power/efficiency curve than H100s which guzzle up power.
- sroussey 3y agoI wonder how hardware will change if LLMs quantized to -1,0,1 really take off.
- 11101010001100 3y agoBut where is the puck going?