6 ms·
Some other good resources: [0]: The original paper: https://arxiv.org/abs/1706.03762 https://arxiv.org/abs/1706.03762 [1]: Full walkthrough for building a GPT
by techbruv 3y ago
Some other good resources:
[0]: The original paper: https://arxiv.org/abs/1706.03762 https://arxiv.org/abs/1706.03762
[1]: Full walkthrough for building a GPT from Scratch: https://www.youtube.com/watch?v=kCc8FmEb1nY https://www.youtube.com/watch?v=kCc8FmEb1nY
[2]: A simple inference only implementation in just NumPy, that's only 60 lines: https://jaykmody.com/blog/gpt-from-scratch/ https://jaykmody.com/blog/gpt-from-scratch/
[3]: Some great visualizations and high-level explanations: http://jalammar.github.io/illustrated-transformer/ http://jalammar.github.io/illustrated-transformer/
[4]: An implementation that is presented side-by-side with the original paper: https://nlp.seas.harvard.edu/2018/04/03/attention.html https://nlp.seas.harvard.edu/2018/04/03/attention.html
- itissid 3y ago[1] is thoroughly recommended.
- revskill 3y agoRecommended for who ?
- basedbertram 3y agoFrom looking at the video probably someone who has good working knowledge of PyTorch, familiarity with NLP fundamentals and transformers, and somewhat of a working understanding of how GPT works.
- deleted 3y ago[deleted]
- quickthrower2 3y agoMasochists! In a good way! I recommend you do the full course not jump into that video. I did the full course, paused to do some of a university course around lecture 2 to really understand some stuff then came back and finishing it off. Bu the end of you would have done stuff like hand working out back-propagation though sums, broadcasting, batchnorm etc. Fairly intense for a regular programmer!
- arthurcolle 3y agoit really is amazing. to be fair if you actually are following along and writing the code yourself, you have to stop and playback quite frequently, and the parts around turning the attention layer into a "block" is a little hard to grok because he starts to speed up around 3/4 through, but yeah this is amazing. I went through it week before starting as lead prompt engineer at an AI startup, and it was super useful and honestly a ton of fun. Reserve 5 hours of your life and go through it if you like this stuff! It's an incredibly great crash course for any interested devs
- basedbertram 3y agoThere's a newer version of [4]: http://nlp.seas.harvard.edu/annotated-transformer/ http://nlp.seas.harvard.edu/annotated-transformer/
- quickthrower2 3y agoDone 1. It is a drawdropper! Especially if you have done the rest of the series and seen results of older architectures. And I was like “where is the rest of it, you ain’t finished!” … and then … ah I see why they named the paper attention is all you need. But even the crappy (small 500k param IIRC) Transformer model trained on a free colab in a couple of minites was relatively impressive. Looking at only 8 chars back and train on a HN thread it got the structure / layout of the page pretty good, interspersed with drunken looking HN comments.
- samvher 3y agoI found this lecture and the one following it very helpful as well: https://www.youtube.com/watch?v=ptuGllU5SQQ&list=PLoROMvodv4rOSH4v6133s9LFPRHjEmbmJ&index=9 https://www.youtube.com/watch?v=ptuGllU5SQQ&list=PLoROMvodv4...
- copirate 3y agoAnd also the ones before that explain the attention mechanism: https://youtu.be/wzfWHP6SXxY?t=4366 https://youtu.be/wzfWHP6SXxY?t=4366 https://youtu.be/gKD7jPAdbpE https://youtu.be/gKD7jPAdbpE (up to 25:42)
- leminimal 3y agoMaybe this is more of a general ML question but I faced it when transformers became popular. Do you know of a project-based tutorial that talks more about neural net architecture, hyperparameters selection and debugging? Something that walks through getting poor results and make explicit the reasoning for tweaking? When I try to use transformers or any AI thing on a toy problem I come up with, it never works. And there's this blackbox of training that's hard to debug into. Yes, for the available resources, if you pick the exact problem, the exact NN architecture and exact hyperparameters, it all works out. But surely they didn't get that on the first try. So what's the tweaking process?
- t-vi 3y agoThere is A. Karpathy's recipe for training NNs but it is not a walkthrough with an example: https://karpathy.github.io/2019/04/25/recipe/ https://karpathy.github.io/2019/04/25/recipe/ but the general idea of "get something that can overfit first" is probably pretty good. In my experience getting the data right is probably the most underappreciated thing. Karpathy has data as step one, but in my experience, also data representation and sampling strategy does quite the miracle. In Part II of our book we do an end-to-end project including e.g. a moment where nothing works until we crop around "regions of interest" to balance the per-pixel classes in the training data for the UNet. This has been something I have pasted into the PyTorch forums every now and then, too.
- leminimal 3y agoThanks for linking me to that post! Its much better at expressing what I'm trying to say. I'll have a careful read of it now. I think I'm still at a step before the overfit. It doesn't converge to a solution on its training data (fit or overfit). And all my data is artificially generated so no cleaning is needed (though choosing a representation still matters). I don't know if that's what you mean by getting the data right or something else. Example problems that "don't work": fizzbuzz, reverse all characters in a sentence.