6 ms·
Why the original transformer figure is wrong, and some other tidbits about LLMs
- amelius 3y agoAre there any still-human-readable pictures where the entire transformer is shown in expanded form?
- dnw 3y agoNot quite what you’re asking for but perhaps in the direction https://openai.com/research/language-models-can-explain-neurons-in-language-models https://openai.com/research/language-models-can-explain-neur...
- minihat 3y agoTry this: https://jalammar.github.io/illustrated-transformer/ https://jalammar.github.io/illustrated-transformer/ Attention is explained separately. I have not seen an all-in-one diagram and cannot imagine one being helpful, since there's too much going on.
- chaxor 3y agoDo you mean a graph that contains all neurons from the network in one structure? Similar to these: https://gfycat.com/BonyTotalArthropods https://gfycat.com/BonyTotalArthropods https://gfycat.com/BitesizedWeeBlacklemur https://gfycat.com/BitesizedWeeBlacklemur That would be wonderful and I have been trying to do this. However, unfortunately some 'assumptions'/shortcuts have to be made. For example, the attention matrix is not known without input, so if just the structure of the network (weighted by the weights) is wanted, you have to put in some value 'p' ('1', '-1', w/e) to these edges. Also skip connections have to be dealt with explicitly instead of just adding them to a block diagonal matrix as one would with an MLP. I am very interested if someone has a good solution that already has done these things though.
- amelius 3y agoThose are nice diagrams. Yes, well actually I'm interested in anything that's between the abstracted form in the paper and the fully expanded form where you see all neurons.
- chaxor 3y agohttps://miro.medium.com/v2/resize:fit:1100/format:webp/0*Y4G0mSaguuNEWFI2.png https://miro.medium.com/v2/resize:fit:1100/format:webp/0*Y4G... Does that work?
- amelius 3y agoAnother cool diagram. But one thing it misses is that you can't follow the arrows from input to output. For example, you might be tempted to think that the Keys or Queries are inputs to the neural net since there is no arrow going into them.
- visarga 3y agoAs someone who started reading ML papers 10 years ago, I find transformers pretty simple compared to most architectures. For example Google's Inception was a complex mesh of various sized convolutions, other models have an internal algorithm like Non Maximum Suppression for object detection, or Connectionist Temporal Classification for OCR, GANs use complicated probability theory for the loss function. Even LSTM is more complicated. If anything, we have been abandoning exotic neural nets in favour of a single architecture and that one is pretty simple, just linear layers (vector-matrix product), key-value products (matrix-matrix product), softmax (fancy normalisation), weighted averaging (a sum of products) and skip connections (an addition). Maybe it's become hard for me to see what is complicated about it, I'd be curious to know what part is difficult. Is it the embeddings, masking, multiple heads, gradient descent, ...? Embeddings have been famous for 10 years, ever since the king - man + woman = queen paper. You don't need to be able to derive the gradients for the network by hand to understand it. In short, a transformer is mixing information between tokens in a sequence and computing updates. The mixing part is the "self attention" or "cross attention". The updating part is the feed-forward sublayer. It has skip connections (adds the input to the output) in order to keep training stable.
- chaxor 3y agoEven when someone understands the architecture very well, there's still great utility in having a full graph representation of the full NN architecture. This could be for teaching students, or for doing analyses on the structure of the full network, etc.
- Buttons840 3y agoSimple is good. Especially in machine learning where a bug usually means that it kinda works, but not as well as it could. Also, when an off-the-shelf algorithm half works, it's good to be able to add you own tweaks to it, and again, this requires simplicity. For a complicated architecture to succeed, it's going to need to reliably achieve state of the art performance on everything without requiring any adjustment or tweaks.
- rasbt 3y agoAgreed, compared to other architectures, transformers are actually quite straight-forward. The complicated part comes more from training it in distributed setups, making the data loading and tensor parallelism work due to the large size etc. Like the vanilla architecture is simple, but the practical implementation for large-scale training can be a bit complicated.
- aptitude_moo 3y agoI find this one as the most complete [1] [2] [1]: https://github.com/ajhalthor/Transformer-Neural-Network/blob/main/Transformer_Architecture_complete.png https://github.com/ajhalthor/Transformer-Neural-Network/blob... [2]: https://www.youtube.com/watch?v=Nw_PJdmydZY&list=PLTl9hO2Oobd97qfWC40gOSU8C0iu0m2l4&index=10 https://www.youtube.com/watch?v=Nw_PJdmydZY&list=PLTl9hO2Oob...
- fn-mote 3y agoThis note contains four papers for "historical perspective"... which would usually mean "no longer directly relevant", although I'm not sure that's really what the author means. You might be looking for the author's "Understanding Large Language Models" post [1] instead. Misspelling "Attention is All Your Need" twice in one paragraph makes for a rough start to the linked post. [1] https://magazine.sebastianraschka.com/p/understanding-large-language-models https://magazine.sebastianraschka.com/p/understanding-large-...
- visarga 3y ago> which would usually mean "no longer directly relevant" Or it could mean the lesson from these papers has been assimilated and spread wide and far, thus they are no longer "news". The pre-layernorm is one.
- homarp 3y agoalso one of these papers is from Schmidhuber and https://news.ycombinator.com/item?id=23649542 https://news.ycombinator.com/item?id=23649542 gives some context to the "For instance, in 1991, which is about two-and-a-half decades before the original transformer paper above ("Attention Is All You Need")"
- rasbt 3y ago> Misspelling "Attention is All Your Need" twice in one paragraph makes for a rough start to the linked post. 100%! LOL. I was traveling and typing this on a mobile device. Must have been some weird autocorrect/autocomplete. Strange. And I didn't even notice. Thanks!
- canjobear 3y agoThe original Transformer wasn’t user in an LLM.
- andreyk 3y agoThe actual title "Why the Original Transformer Figure Is Wrong, and Some Other Interesting Historical Tidbits About LLMs" is way more representative of what this post is about... As to the figure being wrong, it's kind of a nit-pick: "While the original transformer figure above (from Attention Is All Your Need, https://arxiv.org/abs/1706.03762 https://arxiv.org/abs/1706.03762) is a helpful summary of the original encoder-decoder architecture, there is a slight discrepancy in this figure. For instance, it places the layer normalization between the residual blocks, which doesn't match the official (updated) code implementation accompanying the original transformer paper. The variant shown in the Attention Is All Your Need figure is known as Post-LN Transformer."
- rasbt 3y agoSo weird, I posted it with almost the original title (only slightly abbreviated to make it fit: "Why the Original Transformer Figure Is Wrong, and Some Interesting Tidbits About LLMs". Not sure what happened there. Someone must have changed it! So weird! And I agree that the current title is a bit awkward and less representative.
- trivialmath 3y agoI wonder if for example a function is an example of a transformer. So the phrase "argument one is cat" and argument two is dog and operation is join so the result is the word catdog is operated by the transformer as the function concat(cat,dog). Here the query is the function and the keys are the argument for the function and the value is a function from word to words.
- visarga 3y agoThey can intelligently parse the unstructured input into a structured internal form, apply a transform, and then format the result back into unstructured text. Even the transform itself can be an argument.
- ijidak 3y agoHas anyone bought his book: "Machine Learning Q and AI"? Is it a helpful read as a cliff notes for the latest in Generative AI?
- YetAnotherNick 3y agoThis is the commit that changed it: https://github.com/tensorflow/tensor2tensor/commit/d5bdfcc85fa3e10a73902974f2c0944dc51f6a33 https://github.com/tensorflow/tensor2tensor/commit/d5bdfcc85...