10 ms·
How Transformers Work
- ggm 3y agoVery well done animation. In no way does that change my mind about the eggregious continuing use of words like "learn" for what happens inside the software system. To me, the english-language choices in how LLM and Transformers are described as working is part and parcel of the "sell" of steps-to-AGI because of normalising use of words which imply intentionality, introspection (as humans know it), thinking, cognition. So I both love and loathe this animation. It re-inforces people's naieve tendency to "its alive" when in fact, its doing a really good job of explaining how weighted sums and statistics inform the parse, and the consequences of modulation of the products of the parse. The programmers were very smart people. They've done well and have codified a huge body of knowledge which helps systems with producing outputs from input strings, in line with our requirements. Sounds exciting doesn't it!
- ben_w 3y ago> and have codified a huge body of knowledge which helps systems with producing outputs from input strings That is specifically what the programmers have not done.
- famouswaffles 3y agoYeah. Maybe it's because we use the word "train" and that conjures up a mental model of some artisan distilling his process to an an apprentice ?
- ggm 3y agoThis artisan hasn't been trained in how to make steel. It's been trained in how to talk about steel making plausibly.
- famouswaffles 3y agoMy man you're completely off base here. No one is programming anything into neural networks and we have no idea how models like GPT-4 work at a high level. The whole point of neural networks is that we don't know shit and we don't know shit about how to distill the mechanics of some very real, very important processes so we figured out a way to make the machine figure this out on its own.
- ggm 3y agoYou're saying that the billions of lines of text have not informed a statistical model of relationships of words in grammer? You think the "body of knowledge" has to be in the meaning of the words? No. It's in the weighted maths of the likelihood of each element of the phrase. That statistical model is knowledge. Some semantic intent is there too, to disambiguate who flies like an arrow or flies like a banana.
- ben_w 3y ago> That statistical model is knowledge Glad you agree. None of the free parameters of these statistical models were selected by programmers.
- ggm 3y agoHmm. Is adjusting weights metaprogramming? It's based on observations of the outputs of the system. It's certainly not directed programming intent in the normal sense. It shares one property of traditional programming: garbage in garbage out. The fruits of their labour indirectly is a model. It's proving valuable. You want to ascribe the value to .. the bit nobody claims to understand? I'm sticking to the labour theory of value here, either the robot makers imbued value, or we have to redefine the value.
- deleted 3y ago[deleted]
- xcv123 3y agoHumans are not adjusting the weights. That is done automatically by the algorithm. The humans only provide training input.
- ggm 3y agoSo when people told the staff the models were racist you think they just "asked it to be less racist" I think they did a little more than just adjust the input data, they also modulated the logic behind the Weighting methods. No?
- xcv123 3y agoSo the term "Machine Learning" is egregious misuse of language? These models are not explicitly programmed. They actually learn a model from data automatically. The term is technically correct.
- ggm 3y agoThe "learning" part has always caused me grief. Always. Unfortunately it's beyond recovery. But "hallucination" and other new coined terms.. they definitely have problems.
- xcv123 3y agoI don't understand how that would cause grief. Seems totally irrational to me.
- ggm 3y agoThese words are applied to living people. Applied to machines they create states of belief outside the cognoscenti what is going on. It's as if we called programming "magic" and programmers "wizards" and then kids ask "is Harry potter real" and somebody says "yes" It's not hallucinating. It's an analogy. But analogies are not definitional. We're going to wind up in conversations about AGI and use of imprecise analogistic language will alter how people discuss this problem. Sorry you think I'm irratinal. I think you're being obtuse and not considering real world consequences of the ontologies used outside of the field.
- nercury 3y agoDon't "run" programs, don't let them "spawn" threads, don't allow CPU to "follow" instructions, don't "read" or "write" files, should I keep going? All these words used to mean humans doing things, we have been anthropomorphising the machines since the beginning. You are fighting a loosing battle here, better to just calm down and relax.
- 3y ago
- KirillPanov 3y agoHonestly this single image does a better job of explaining transformers than anything else I've ever seen: https://people.idsia.ch/~juergen/fastweights754x288.png https://people.idsia.ch/~juergen/fastweights754x288.png It's pretty damn simple: a linearized transformer is a "slow" neural net whose outputs determine the weights of another ("fast") neural net. It's a NN that can use tools, subject to one major restriction: the tool must be another NN. The restriction is is "the trick" that lets you backpropagate gradients through both NNs so you can train the slow NN based on an error function of the fast NN's outputs. The only difference between linearized transformers and the kind that OpenAI uses is adding a softmax operation in one place.
- raylad 3y agoIt seems to explain them well because you already know how they work. To most people, that diagram would not explain very much.
- KirillPanov 3y agoNo, I didn't! Before I came across that image (and the paper Linearized Transformers are Secretly Fast Weight Programmers) I very much did not understand transformers. I had spent at least 20 hours with Attention Is All You Need, and probably another dozen hours with https://e2eml.school/transformers.html https://e2eml.school/transformers.html and wasn't getting much of anywhere. I have no ML/AI background; just a typical undergraduate-CS-level familiarity with neural networks -- the basic stuff that hasn't changed since the 1990s. I do have some experience with linear algebra, but that isn't the hard part of any of this. Frankly, most people who publish in this field go out of their way to obfuscate the key insights. Mediocre physicists (but not the truly brilliant ones) do the same thing. It's very annoying.
- xcv123 3y agoSo the diagram makes sense after you spent at least 32 hours studying LLMs.
- nurettin 3y agoThis is more like "what transformers do" rather than "how they work".
- famouswaffles 3y agoYes well nobody really knows the "how they work" part.
- marktani 3y agoSo the title of this post could be changed to "What Tranformers Are" or "What Transformers Do" instead
- gnabgib 3y agoThat seems to be the (uneditorialized) title: "Generative AI exists because of the transformer (This is how it works)". Maybe the context of 'it' was lost on the submitter.
- trumbitta2 3y agoI came here hoping to geek out about Autobots and Decepticons.
- agentgumshoe 3y agoYeah, this was definitely not more than meets the eye.
- bigstrat2003 3y agoI would argue that by subverting OP's expectations, it was more than meets the eye. It definitely isn't robots in disguise, though. :(
- jacomoRodriguez 3y agoI really enjoyed the beginning of this page, where it talked about embeddings and such. But starting with the self attention, it gets a bit fussy. How, based on what does the attention mechanism work? The similarity of embeddings can't be all of it - "it", "dog" and "bone" will alwats have the same similarity, even if the surrounding sentence changes. Can someone in simple terms describe how thus works? Extra: the embedding explanation works great with words, but with actual tokens it gets way weirder... e.g. in German tokens are often just 1-3 characters, which do not contain any meaning in the normal sense. So the fact, that this works for example with the "hu" from "Hund", "fl" from "Flasche" and "fl" from "Fleisch" is interesting, as "Fleisch" is probably more similar to "Hund" then "Flasche". (Tokens are just examples, not sure how this wird's are broken down). I guess the solution to this is related to the question above? For this kind of embedding to be useful it is important to look at more than just the two tokens, but their complete surroundings?
- ma2rten 3y agoAttention takes in all tokens in the sequence and outputs a new representation of the current token in context. Each layer of the transformer adds more context to the token. I haven't read this explanation in detail and although they have some nice animations, I wouldn't go to FT to explain machine learning concepts. Here are two well known explanations that might be better: http://jalammar.github.io/illustrated-transformer/ http://jalammar.github.io/illustrated-transformer/ http://nlp.seas.harvard.edu/annotated-transformer/ http://nlp.seas.harvard.edu/annotated-transformer/.
- lawlessone 3y agoSo is it analogous to how a CNN starts with fragments of images and further up the chain assembles these into objects?
- ma2rten 3y agoYes, I think that is a reasonable way to think about it, in my opinion. However, with the language modeling objective it predicts the next token and because of the residual connections each intermediate layer is in the same space. So, maybe it would be more accurate to say that it is an increasingly accurate representation of the next token.
- Leimi 3y agoI wonder how these "visual story telling" articles are created, they are really great. Like, what tools do the authors have. Are the content authors super tech savvy or not. How much specific code must be created for each article. How long does everything take compared to a normal, mostly-text page. How many people work on one article. etc. Must be pretty interesting.
- revskill 3y agoYes, the author doesn't understand that, user is interested more on how article is created than the article content itself.
- lurquer 3y agoAgree. Ironically, I turned to ChatGPT to ask how one would design a webpage such as that!
- alsodumb 3y agoDistill was a new take at publishing research/ideas in deep learning in a visual way: https://distill.pub/ https://distill.pub/ I love their articles and while it was hard to sustain, the quality of the ones in their are pretty good. They provide some tips and templates on how to develop such visual storytelling articles.
- srvmshr 3y agoMostly a ton of JavaScript. As for the concepts explained visually, ICYMI there is a small footnote at the end of the article: > To generate the 50D word embeddings we used the GloVe 6B 50D pre-trained model and converted to Word2Vec format. To generate the 2D representation of word embeddings we used the BERT large language model and reduced dimensionality using UMAP. The self-attention values and the probability scores in the beam search section are conceptual.
- samlearner 3y agoHey, I'm one of the graphics journalists/authors on the piece This is not very helpful, but the answers to most of your questions is "it depends" Our team is made up of reporters, designers, and graphics journalists, but the specific makeup of the team on a given project or who else gets drawn into it varies a lot depending on the topic/scope of the story For our stories, always lot of React and headache-inducing CSS transition stuff, but the tools/libraries beyond that depend a lot on the needs of the project - On some stories, there's a lot of blender/threejs work, like this one on quantum computing: https://ig.ft.com/quantum-computing/ https://ig.ft.com/quantum-computing/ - For others, like this one, there's a lot of mapping/data work: https://ig.ft.com/ukraine-war-food-insecurity/ https://ig.ft.com/ukraine-war-food-insecurity/ Some stories take a couple of weeks, some take a couple of months and feature fairly large codebases with 500+ commits
- light_hue_1 3y agoAnother horrible and misleading description that I cannot imagine had a single machine learning person in the loop. TLDR: They say that transformers are pretrained word embeddings with one round of self attention that predict the next word sequence by beam search. They aren't any of those things. Why start with word embeddings? The authors think that you first train word embeddings and then use them as the input for your Transformer. Nope. > A word embedding can have hundreds of values, each representing a different aspect of a word’s meaning. Just as you might describe a house by its characteristics — type, location, bedrooms, bathrooms, storeys — the values in an embedding quantify a word’s linguistic features. No. That's literally not what a word embedding is! It's completely the wrong intuition. It mixes up a semantic representation with the distributional hypothesis. These are two completely different ideas. Word embeddings are not describing words by their characteristics, exactly the opposite! They abandon that idea. To say that this is true but we just don't know what the characteristics being described are is just a total confusion about what is happening. > Self-attention looks at each token in a body of text and decides which others are most important to understanding its meaning. Why even introduce the word token here? When it's not explained at all. Just say word. And no, self attention does not decide which words are important for the meaning of this word, it's which other inputs are important to carry out some task. "meaning" is not a thing. Also, that's just attention. Where's the self part friends!? It sounds like they think self attention refers to a word itself paying attention. Things are getting dicey. > With self-attention, the transformer computes all the words in a sentence at the same time. Capturing this context gives LLMs far more sophisticated capabilities to parse language. Well, for one thing there's masked attention. For another, this isn't a property of self-attention. It's a property of how Transformers split apart attention and the positional feed forward layers. But I'm sure we'll talk about that (spoiler: we won't, things fall apart dramatically shortly; the authors don't understand anything). > Transformers process an entire sequence at once — be that a sentence, paragraph or an entire article — analysing all its parts and not just individual words. They just talked about an RNN that process words. And all of their parts. Also there are bidirectional RNNs. This comparison is nonsense. They should just not talk about RNNs if they can't say anything reasonable. > This allows the software to capture context and patterns better, and to translate — or generate — text more accurately. This simultaneous processing also makes LLMs much faster to train, in turn improving their efficiency and ability to scale. Ironically, they managed to identify the one part of LLMs that is slowest and most problematic as the part that is the fastest. A whole lot of papers focus on making attention faster because it's the computational bottleneck of Transformers; very funny to call that what makes them faster to train. Also, faster to train than what!? RNNs just don't scale up well (traditionally, let's not get into newer models). So it's not a matter of faster, it's a matter of, we couldn't train large models before at all. Then there's a whole lot of "it does X, Y, and Z", no idea what in their story would make anyone think the models should ever do that. > From this enormous corpus of words and images, the models learn how to recognise patterns and eventually predict the next best word. Huh? How'd we sneak this one in? So the story they're telling is: we learn word embeddings, then we self attend to them which means looking at how other words related to them, and then that model learns to predict the next best word? Wow, it's like watching a blender process ideas into goo. > After tokenising and encoding a prompt, we’re left with a block of data representing our input as the machine understands it, including meanings, positions and relationships between words. tokenising!? encoding a prompt!? Where'd this come from. I love how we start the sentence with "tokenizing" and end it with "words", just a total confusion of ideas. A block of data encoding the text as the machine understands it. Look at that beautiful block of data! A beautiful block of data that comes from.. word embeddings + positional encoding (huh? never talked about this one either) + self attention. Because, that's all there is to transformers. Just some self attention, one round of it. On top of word embeddings. That's the message. That's total and complete gibberish. > At its simplest, the model’s aim is now to predict the next word in a sequence and do this repeatedly until the output is complete. What is there to predict? In the story they've told I have pretrained word embeddings and self attention. > To do this, the model gives a probability score to each token, which represents the likelihood of it being the next word in the sequence. A "probability score"!? My inner statistician just took their own life. The next word in the sequence!? They're mixing up decoders with masked word prediction. > With beam search, the model is able to consider multiple routes and find the best option. Why!? Why talk about beam search? Who told them about beam search? Most models do something called top-p not beam search these days. But that's such a minor and worthless detail. > This produces better results, ultimately leading to more coherent, human-like text. I wish human-like text was coherent. So there you have it. Transformers are pretrained word embeddings with one round of self attention that predict the next word sequence by beam search. cry
- doublerabbit 3y agoThis is about ML. I was expecting a subject based on how power station transformers works
- moomoo3000 3y agoLiving in a cave huh?
- doublerabbit 3y agoAt this point in time living in a cave sounds lovely as for the dystopian future we are brewing.
- zwieback 3y agoIn a cave with power brought to us by TRANSFORMERS, the real kind, made out of iron and copper.
- bigstrat2003 3y agoNot knowing terms of art for machine learning isn't exactly living under a cave.
- hoosieree 3y agoLet's ROLL OUT!
- deleted 3y ago[deleted]
- Imnimo 3y agoThis feels like it's a mish-mash of concepts that are used in different settings (pre-trained word embeddings, unmasked attention, beam search, ...) but that aren't really used in things like ChatGPT. I guess if you're reading this kind of article, the details don't really matter that much, though.
- CaffeinatedDev 3y agoI was so enamored with the clean styling animations that I couldn't process the article.
- est 3y agoI am interested in how these work, and most importantly why others failed to work. e.g. why LLM's capabilities where "unlocked" by prompts like "step-by-step"? How does CoT work?