8 ms·
Transformers from Scratch (2021)
- metalloid 3y agoThe author of the article should had provided an implementation of the transformer using only numpy or pure C++.
- jaymody 3y agoI wrote a minimal implementation in NumPy here (the forward pass code is only 40 lines): https://github.com/jaymody/picoGPT https://github.com/jaymody/picoGPT And also a related blog post: https://news.ycombinator.com/item?id=34726115 https://news.ycombinator.com/item?id=34726115 Although this is for a decoder-only transformer (aka GPT) and doesnt include the encoder part.
- newhouseb 3y agoAnd I adapted Jay's work to Typescript (without the numpy obviously, just raw typescript/javascript): https://github.com/newhouseb/potatogpt https://github.com/newhouseb/potatogpt
- toyg 3y agoMORE THAN MEETS THE EYE! ... oh, not those Transformers. Meh.
- zabzonk 3y agothat was my first glance reading too - shape-changing toys written in the scratch language! how cool could that be?
- MisterTea 3y agoI was hoping it was an article on designing and building an electrical transformer complete with pictures of a home made, hand wound transformer. I was very disappointed.
- quickthrower2 3y agoso I’m on the same journey of trying to teach myself ML and I do find most of the resources go over things very quickly and leave a lot you to figure out yourself. Having had a quick look at this one, it looks very beginner, friendly, and also very careful to explain things slowly, so I will definitely added to my reading list. Thanks to the author for this!
- KyeRussell 3y agoUltimately because it’s such a hot topic the “market” is flooded with people that want to crank out content without understanding what they’re talking about.
- quickthrower2 3y agoTrue but I weed that stuff out. I have been doing university courses shared online or very reputable courses.
- bambax 3y ago[flagged]
- Reason077 3y ago[flagged]
- stared 3y agoThank you for sharing! For the "from scratch" version, I recommend "The GPT-3 Architecture, on a Napkin" https://dugas.ch/artificial_curiosity/GPT_architecture.html https://dugas.ch/artificial_curiosity/GPT_architecture.html, which was there as well (https://news.ycombinator.com/item?id=33942597 https://news.ycombinator.com/item?id=33942597). Then, to actually dive into details, "The Annotated Transformer", i.e. a walktrough "Attention Is All You Need", with code in PyTorch, https://nlp.seas.harvard.edu/2018/04/03/attention.html https://nlp.seas.harvard.edu/2018/04/03/attention.html.
- ziyunli 3y agoThere is a newer version of the second article https://nlp.seas.harvard.edu/annotated-transformer/ https://nlp.seas.harvard.edu/annotated-transformer/
- dsubburam 3y agoAn early explainer of transformers, which is a quicker read, that I found very useful when they were still new to me, is The Illustrated Transformer[1], by Jay Alammar. A more recent academic but high-level explanation of transformers, very good for detail on the different flow flavors (e.g. encoder-decoder vs decoder only), is Formal Algorithms for Transformers[2], from DeepMind. [1] https://jalammar.github.io/illustrated-transformer/ https://jalammar.github.io/illustrated-transformer/ [2] https://arxiv.org/abs/2207.09238 https://arxiv.org/abs/2207.09238
- driscoll42 3y agoThe Illustrated Transformer is fantastic, but I would suggest that those going into it really should read the previous articles in the series to get a foundation to understand it more, plus later articles that go into GPT and BERT, here's the list: A Visual and Interactive Guide to the Basics of Neural Networks - https://jalammar.github.io/visual-interactive-guide-basics-neural-networks/ https://jalammar.github.io/visual-interactive-guide-basics-n... A Visual And Interactive Look at Basic Neural Network Math - https://jalammar.github.io/feedforward-neural-networks-visual-interactive/ https://jalammar.github.io/feedforward-neural-networks-visua... Visualizing A Neural Machine Translation Model (Mechanics of Seq2seq Models With Attention) - https://jalammar.github.io/visualizing-neural-machine-translation-mechanics-of-seq2seq-models-with-attention/ https://jalammar.github.io/visualizing-neural-machine-transl... The Illustrated Transformer - https://jalammar.github.io/illustrated-transformer/ https://jalammar.github.io/illustrated-transformer/ The Illustrated BERT, ELMo, and co. (How NLP Cracked Transfer Learning) - https://jalammar.github.io/illustrated-bert/ https://jalammar.github.io/illustrated-bert/ The Illustrated GPT-2 (Visualizing Transformer Language Models) - https://jalammar.github.io/illustrated-gpt2/ https://jalammar.github.io/illustrated-gpt2/ How GPT3 Works - Visualizations and Animations - https://jalammar.github.io/how-gpt3-works-visualizations-animations/ https://jalammar.github.io/how-gpt3-works-visualizations-ani... The Illustrated Retrieval Transformer - https://jalammar.github.io/illustrated-retrieval-transformer/ https://jalammar.github.io/illustrated-retrieval-transformer... The Illustrated Stable Diffusion - https://jalammar.github.io/illustrated-stable-diffusion/ https://jalammar.github.io/illustrated-stable-diffusion/ If you want to learn how to code them, this book is great: https://d2l.ai/chapter_attention-mechanisms-and-transformers/index.html https://d2l.ai/chapter_attention-mechanisms-and-transformers...
- cuuupid 3y agoThis is cool, I highly recommend Jay Alammar’s Illustrated Transformer series to anyone wanting to get an understanding of the different types of transformers and how self-attention works. The math behind self-attention is also cool and easy to extend to e.g. dual attention
- Buttons840 3y agoThis article describes positional encodings based on several sine waves with different frequencies, but I've also seen positional "embeddings" used, where the position (the position is an integer value) is used to select an differentiable embedding from an embedding table. Thus, the model learns its own positional encoding. Does anyone know how these compare? I've also wondered why we add the positional encoding to the value, rather than concatenating them? Also, the terms encoding, embedding, projection, and others are all starting to sound the same to me. I'm not sure exactly what the difference is. Linear projections start to look like embeddings start to look like encodings start to look like projections, etc. I guess that's just the nature of linear algebra? It's all the same? The data is the computation, and the computation is the data. Numbers in, numbers out, and if the wrong numbers come out then God help you. I digress. Is there a distinction between encoding, embedding, and projection I should be aware of? I recently read in "The Little Learner" book that finding the right parameters is learning. That's the point. Everything we do in deep learning is focused on choosing the right sequence of numbers and we call those numbers parameters. Every parameter has a specific role in our model. Parameters are our choice, those are the nobs that we (as a personified machine learning algorithm) get to adjust. Ever since then the word "parameters" has been much more meaningful to me. I'm hoping for similar clarity with these other words.
- giovannibonetti 3y ago> I recently read in "The Little Learner" book that finding the right parameters is learning. That's the point. Everything we do in deep learning is focused on choosing the right sequence of numbers and we call those numbers parameters. Every parameter has a specific role in our model. Parameters are our choice, those are the nobs that we (as a personified machine learning algorithm) get to adjust. Be careful not to mistake parameters for hyperparameters. - Parameters are the result of the training phase, as you mentioned. They start with random values and are discovered by the training algorithm; - Hyperparameters, on the other hand, are the knobs you tweak to make the training process arrive at the "right" parameters. You can think of them as meta-parameters; Also, it is important to think on the ML architecture - transformers, neural networks, random forests and so on - as the parameters change completely depending on which one you're using.
- adriantam 3y agoIf you want a TensorFlow implementation, here it is: https://machinelearningmastery.com/building-transformer-models-with-attention-crash-course-build-a-neural-machine-translator-in-12-days/ https://machinelearningmastery.com/building-transformer-mode...
- erwincoumans 3y agoAndrej Karpathy's 2 hour video and code is really good to understand the details of Transformers: "Let's build GPT: from scratch, in code, spelled out." https://youtube.com/watch?v=kCc8FmEb1nY https://youtube.com/watch?v=kCc8FmEb1nY
- lucidrains 3y agobesides everything that was mentioned here, what made it finally click for me early in my journey was running through this excellent tutorial by Peter Bloem multiple times https://peterbloem.nl/blog/transformers https://peterbloem.nl/blog/transformers highly recommend
- deleted 3y ago[deleted]
- dingosity 3y agoDid anyone make the obvious "Robots in Smalltalk" joke yet? Okay... here goes... When I first read that title I thought the author was talking about Robots in Smalltalk.
- JackFr 3y agoRead this as "Transformers in Scratch" at first and was very curious. Obviously implementing transformers in Scratch is likely impossible, but has anyone built a Scratch-like environment for building NN models?
- pmoriarty 3y agoSo how practical is learning to create your own transformers if you can't afford a giant amount of resources to train them?
- almost 3y agoUnderstanding how things work is useful and worthwhile on its own. Also while you probably can’t afford to train your own LLM you probably can afford to fine tune an existing one or to join one to another mode or lots of other things like that.
- sachinkalsi 3y agoCheck this out https://youtu.be/73gTEub2e3I https://youtu.be/73gTEub2e3I
- dang 3y agoRelated: Transformers from Scratch - https://news.ycombinator.com/item?id=29315107 https://news.ycombinator.com/item?id=29315107 - Nov 2021 (17 comments) also these, but it was a different article: Transformers from Scratch (2019) - https://news.ycombinator.com/item?id=29280909 https://news.ycombinator.com/item?id=29280909 - Nov 2021 (9 comments) Transformers from Scratch - https://news.ycombinator.com/item?id=20773992 https://news.ycombinator.com/item?id=20773992 - Aug 2019 (28 comments)
- leobg 3y agoCan somebody explain to me the sinus wave positional encoding thing? The naïve approach would be to just add number indices to the tokens, wouldn’t it?
- tipsytoad 3y agoI'm no expert, but I think it's so that the model can learn the relative position wrt other tokens. They use indices for models like vision transformers with a fixed number of patches but for variable length context I think it's more beneficial to use encodings that can also capture the relative distance.
- ftxbro 3y agoAccording to https://kazemnejad.com/blog/transformer_architecture_positional_encoding/ https://kazemnejad.com/blog/transformer_architecture_positio... they explain it as a clever way to satisfy the following criteria: - It should output a unique encoding for each time-step (word’s position in a sentence) - Distance between any two time-steps should be consistent across sentences with different lengths. - Our model should generalize to longer sentences without any efforts. Its values should be bounded. - It must be deterministic. Your example contradicts the 'values should be bounded' criterion as it generalizes to longer sentences.