9 ms·
Implementing a ChatGPT-like LLM from scratch, step by step
- AndrewKemendo 3y agoWriting a technical book in public is a level of anxiety I can’t imagine, so kudos to the author!
- rasbt 3y agoIt kind of is, but it's also kind of motivating :)
- waynesonfire 3y agoIt's actually less risky. The author may be able to reap the benefits of writing a book without actually finishing it. Ideally, maybe not much more than Chapter 1.
- _giorgio_ 3y ago[flagged]
- waynesonfire 3y agoWhat refund? Nobody even mentioned a transaction taking place. The contents are on a github repo. Manning has zero risk here, except that if they keep getting burned by half finished books they may want to re-evaluate these contracts.
- _giorgio_ 3y agoAnd here's the crazy one explaining to Manning how to do business. Incredible.
- rasbt 3y agoI'd say that I've finished all of my previous books, and I have no intention of doing anything different here. Of course, there's always the chance that I get run over by a bus or equivalent, but in that case, I assume that Manning would find a replacement (as per contract) who finishes the book. I don't think there are any benefits to be reaped from not finishing.
- kif 3y agoLooks like just the kind of book I'd want to read. I bought a copy :)
- rasbt 3y agoGlad to hear and thanks for the support. Chapter 3 should be in the MEAP soonish (submitted the draft last week). Will also upload my code for chapter 4 to GitHub soonish, in the next couple of days, just have to type up the notes.
- intalentive 3y agoThe model architecture itself is really not too complex, especially with torch. The whole process is pretty straightforward. Nice feasible project.
- turnsout 3y agoThis looks amazing @rasbt! Out of curiosity, is your primary goal to cultivate understanding and demystify, or to encourage people to build their own small models tailored to their needs?
- rasbt 3y agoI'd say my primary motivation is an educational goal, i.e., helping people understand how LLMs work by building one. LLMs are an important topic, and there are lots of hand-wavy videos and articles out there -- I think if one codes an LLM from the ground up, it will clarify lots of concepts. Now, the secondary goal is, of course, also to help people with building their own LLMs if they need to. The book will code the whole pipeline, including pretraining and finetuning, but I will also show how to load pretrained weights because I don't think it's feasible to pretrain an LLM from a financial perspective. We are coding everything from scratch in this book using GPT-2-like LLM (so that we can load the weights for models ranging from 124M that run on a laptop to the 1558M that runs on a small GPU). In practice, you probably want to use a framework like HF transformers or axolotl, but I hope this from-scratch approach will demystify the process so that these frameworks are less of a black box.
- turnsout 3y agoThanks for such a thoughtful response. I'm building with LLMs, and do feel uncomfortable with my admittedly hand-wavy understanding of the underlying transformer architecture. I've ordered your book and look forward to following along!
- clueless 3y agoare the code for chapter 4 through 8 missing?
- rasbt 3y agoIt's in progress still. I have most of the code working, but it's not organized into the chapter structure, yet. I am planning to add a new chapter every ~month (I wish I could do this faster, but I also have some other commitments). Chapter 4 will be either uploaded by the end of this weekend or by the end of next weekend.
- _giorgio_ 3y agoDepending on your level, it could take a lot of weeks to go through the already available material (code and pdf), so I'd suggest to purchase it anyway... It makes no sense to wait until the end, if you're interested in the subject.
- malermeister 3y agoHow does this compare to the karpathy video [0]? I'm trying to get into LLMs and am trying to figure out what the best resource to get that level of understanding would be. [0] https://www.youtube.com/watch?v=kCc8FmEb1nY https://www.youtube.com/watch?v=kCc8FmEb1nY
- _giorgio_ 3y agoYou can't understand it unless you already know most of the stuff. I've watched it many times to understand well most of it. And obviously you must already know pytorch really well, including the matrix multiplication, backpropagation etc. He speaks very fast too...
- tayo42 3y agoHe has like 4 or 5 videos that can be watched before that one where all of that is covered. He goes over stuff like writing back prop from scratch and implementing layers without torch.
- _giorgio_ 3y agoI know... That material isn't for beginners.
- mikeiavelli 3y ago...but then, what material did you expect as a beginner?
- _giorgio_ 3y agoHow can Karpathy videos defined for beginners when you have to know: programming, python, pytorch, matrix multiplication, derivatives...
- hadjian 3y agoDid you really watch all videos in the playlist? I am at video 4 and had no background in PyTorch or numpy. In my opinion he covers everything needed to understand his lectures. Even broadcasting and multidimensional indexing with numpy. Also in the first lecture you will implement your own python class for building expressions including backprop with an API modeled after PyTorch. IMHO it is the second lecture I can recommend without hesitation. The other is Gilbert Strang on linear algebra.
- SushiHippie 3y agofyi probably qualifies as an "Show HN:"
- npalli 3y agoimport torch From the first code sample, not quite from scratch :-)
- rasbt 3y agoLol ok, otherwise it would probably be not very readable due to the verbosity. The book shows how to implement LayerNorm, Softmax, Linear layers, GeLU etc without using the pre-packaged torch versions though.
- notso411 3y ago[dead]
- nerdponx 3y agoI don't think implementing autograd is relevant or in-scope for learning about how transformers work (or writing out the gradient for transformer by hand, I can't even imagine doing that).
- PheonixPharts 3y agoAutomatic differentiation is why we are able to have complex models like transformers, it's arguably the key reason (in addition to large amounts of data and massive compute resources) that we have the revolution in AI that we have. Nobody working in this space is hand calculating derivatives for these models. Thinking in terms of differentiable programming is a given and I think certainly counts as "from scratch" in this case. Any time I see someone post a comment like this, I suspect the don't really understand what's happening under the hood or how contemporary machine learning works.
- schneems 3y agoI’m very comfortable with AI in general but not so much with Machine Lesrning. I understand transformers are a key piece of the puzzle that enables tools like LLMs but don’t know much about them. Do you (or others) have good resources explaining what they are and how they work at a high level?
- bosky101 3y agoHow was the process of pitching to Manning?
- rasbt 3y agoThat was pretty smooth. They reached out whether I was interested in writing a book for them (probably because of my other writings online), I mentioned what kind I book I want to write, submitted a proposal, and they liked that idea :)
- deleted 3y ago[deleted]
- wslh 3y agoI jumped to Github thinking this is would be a free resource (with all due respect to the author work). What free resources are available and recommended in the "from scratch vein"?
- larme 3y agohttps://jaykmody.com/blog/gpt-from-scratch/ https://jaykmody.com/blog/gpt-from-scratch/ for a gpt2 inference engine in numpy then https://www.dipkumar.dev/becoming-the-unbeatable/posts/gpt-kvcache/ https://www.dipkumar.dev/becoming-the-unbeatable/posts/gpt-k... for adding a kv cache implementation
- larme 3y agoI'd like to add that most of these text only talking about inference part. This book (I also purchased the draft version) has training and finetuning in the TOC. I assume it will include materials about how to do training and finetuning from scratch.
- rasbt 3y agoI added notes to the Jupyter notebooks, I hope they are also readable as standalone from the repo.
- natrys 3y agoNeural Networks: Zero to Hero[1] by Andrej Karpathy [1] https://karpathy.ai/zero-to-hero.html https://karpathy.ai/zero-to-hero.html
- villedespommes 3y ago+1, Andrey is an amazing educator! I'd also recommend his https://youtu.be/kCc8FmEb1nY?si=mP0cQlQ4rcceL2uP https://youtu.be/kCc8FmEb1nY?si=mP0cQlQ4rcceL2uP and checking out his github repos. MinGPT, for example, implements a small gpt model that's compatible with HF API, whereas more modern nanoGPT shows how to use newer features such as flash attention. The quality of every video, every blog post is just so high.
- whartung 3y agoCan I use any of the information in this book to learn about reinforcement learning? My goal is to have something learn to land, like a lunar lander. Simple, start at 100 feet, thrust in one direction, keep trying until you stop making craters. Then start adding variables, such as now it's moving horizontally, adding a horizontal thruster. next, remove the horizontal thruster and let the lander pivot. Etc. I just have no idea how to start with this, but this seems "mainstream" ML, curious if this book would help with that.
- Buttons840 3y agoI enjoyed "Grokking Deep Reinforcement Learning"[0]. It doesn't include anything about transformers though. Also, see Python's gymnasium[1] library for a lunar lander environment, it's the one I focused on most while I was learning and I've solved it a few different ways now. You can also look at my own notebook I used when implementing Soft Actor Critic with PyTorch not too long ago[2], it's not great for teaching, but maybe you can get something out of it. [0]: https://www.manning.com/books/grokking-deep-reinforcement-learning https://www.manning.com/books/grokking-deep-reinforcement-le... [1]: https://gymnasium.farama.org/environments/box2d/ https://gymnasium.farama.org/environments/box2d/ [2]: https://github.com/DevJac/learn-pytorch/blob/main/SAC.ipynb https://github.com/DevJac/learn-pytorch/blob/main/SAC.ipynb
- thatguysaguy 3y agoTry OpenAI's spinning up: https://spinningup.openai.com/en/latest/ https://spinningup.openai.com/en/latest/
- Buttons840 3y agoThis is a good and short introduction to RL. The density of the information in Spinning Up was just right for me and I think I've referred to it more often than any other resource when actually implementing my own RL algorithms (PPO and SAC). If I had to recommend a curriculum to a friend I would say: (1) Spend a few hours on Spinning Up. (2) If the mathematical notation is intimidating, read Grokking Deep Reinforcement Learning (from Manning), which is slower paced and spends a lot of time explaining the notation itself, rather than just assuming the mathematical notation is self-explanatory as is so often the case. This book has good theoretical explanations and will get you some running code. (3) Spend a few hours with Spinning Up again. By this point you should be a little comfortable with a few different RL algorithms. (4) Read Sutton's book, which is "the bible" of reinforcement learning. It's quite approachable, but it would be a bit dry and abstract without some hands-on experience with RL I think.
- LoveYourBooks 3y ago[dead]
- Karupan 3y agoBought a copy. Good luck rasbt!
- rasbt 3y agoThanks :)
- ijustwanttovote 3y agoWow, great info. Thanks for sharing.
- theogravity 3y agoPurchased the book. Really excited to read it!
- rasbt 3y agoThanks! And please don't hesitate to reach out via the Forum or the GitHub Discussions if you have any feedback or questions.
- photon_collider 3y agoBought a copy! Looking forward to reading it. :) Is there a way for readers to give feedback on the book as you write it?
- canyon289 3y agoFor an additional resource I'm writing a guide book, though its in various stages of completion The fine tuning guide is the best resource so far https://ravinkumar.com/GenAiGuidebook/language_models/finetuning.html https://ravinkumar.com/GenAiGuidebook/language_models/finetu...
- czechdeveloper 3y agoSuch a great source of information. Thank you.
- canyon289 3y agoOf course! Is there's anything in particular you're interested in or a topic you want me to cover let me know. This tech is powerful by itself. Hoping to empower people with knowledge of all this works too :)
- iamcreasy 3y agoThank you for this endeavour. Do you have an ETA for the completion of the book?
- rasbt 3y agoThe ETA for the last chapter is August if things continue to go well. It's usually available in the MEAP a few weeks after that, some time in September. And print version should be available early 2025 I think.
- iamcreasy 3y agoI'll definitely buy it once released. In the meantime, do you know any other free/paid resource that comes close to what you are trying to achieve with this book?
- rasbt 3y agoUnfortunately, I am not aware of any other resource that delves into these topics. However, as others commented above, Karpathy has a 2h YouTube video that is probably worthwhile watching. Based on skimming the YT video, it has some overlap with chapters 3 & 4, but the book has a much larger scope. I am not sure how to link to other comments on HN, so let me just copy & paste it here: > How does this compare to the karpathy video [0]? I'm trying to get into LLMs and am trying to figure out what the best resource to get that level of understanding would be. [0] https://www.youtube.com/watch?v=kCc8FmEb1nY https://www.youtube.com/watch?v=kCc8FmEb1nY > Haven't fully watched this but from a brief skimming, here are some differences that the book has: - it implements a real word-level LLM instead of a character-level LLM - after pretraining also shows how to load pretrained weights - instruction-finetune that LLM after pretraining - code the alignment process for the instruction-finetuned LLM - also show how to finetune the LLM for classification tasks - the book it overall has a lots of figures. For Chapter 3, there are 26 figures alone :) The video looks awesome though. I think it's probably a great complementary resource to get a good solid intro because it's just 2 hours. I think reading the book will probably be more like 10 times that time investment.
- two_in_one 3y agoAs it's still work in progress may I suggest? It would be nice if you go beyond what others have already published and add more details. Like different position encodings, MoE, decoding methods, tokenization. As it's educational easy to use should be a priority, of course.
- rasbt 3y agoThanks, comparing positional encodings, MoEs, kv-caches etc are all good topics that I have in mind for either supplementary material and/or a follow-up book. The reason why it probably won't land in this current book is the length and time line. It's already going to be a big book as it is (400-500 pages). And I also want to be a bit mindful of the planned release date. However, these are indeed good suggestions.
- Buttons840 3y agoQuestion for the author: I'm not interested in language models specifically, but there are techniques involved with language models I would like to understand better and use elsewhere. For example, I know "attention" is used in a variety of models, and I know transformers are used in more than just language models. Will this book help me understand attention and transformers well enough that I can use them outside of language models?
- rasbt 3y agoThe attention mechanism we implement in this book* is specific to LLMs in terms of the text inputs, but it's fundamentally the same attention mechanism that is used in vision transformers. The only difference is that in LLMs, you turn text into tokens, and convert these tokens into vector embeddings that go into an LLM. In vision transformers, instead of regarding images as tokens, you use an image patch as a token and turn those into vector embeddings (a bit hard to explain without visuals here). In both text or vision context, it's the same attention mechanism, and it both cases it receives vector embeddings. (*Chapter 3, already submitted last week and should be online in the MEAP soon, in the meantime the code along with the notes is also available here: https://github.com/rasbt/LLMs-from-scratch/blob/main/ch03/01_main-chapter-code/ch03.ipynb https://github.com/rasbt/LLMs-from-scratch/blob/main/ch03/01...)
- corethree 3y agoNowadays anyone can probably put together a good book about this topic by using an LLM.
- towelpluswater 3y agoBought a copy! Your posts and newsletter content has been such a huge inspiration for me throughout 2023 - good luck, this is a huge effort!
- rasbt 3y agothanks for the kind words!