3 ms·
Haven't fully watched this but from a brief skimming, here are some differences that the book has: - it implements a real word-level LLM instead of a character
by rasbt 3y ago
Haven't fully watched this but from a brief skimming, here are some differences that the book has:
- it implements a real word-level LLM instead of a character-level LLM
- after pretraining also shows how to load pretrained weights
- instruction-finetune that LLM after pretraining
- code the alignment process for the instruction-finetuned LLM
- also show how to finetune the LLM for classification tasks
- the book it overall has a lots of figures. For Chapter 3, there are 26 figures alone :)
The video looks awesome though. I think it's probably a great complementary resource to get a good solid intro because it's just 2 hours. I think reading the book will probably be more like 10 times that time investment.
- malermeister 3y agoThank you for the answer! What is the knowledge that your book requires? If I have a lot of software dev experience and sorta kinda remember algebra from uni, would it be a good fit?
- rasbt 3y agoGood question. I think a Python background is strongly recommended. PyTorch knowledge would be a nice to have (although I've written a comprehensive 40 page intro for the Appendix, which is also already available). From a math perspective, I think it should be gentle. I'm introducing dot products in Chapter 3, but I also explain how you could do the same with for-loops. Same with matrix multiplication. I'm bad at estimating requirements, but I hope this should be sufficient.