Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
rasbt
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
31.
▲
by
rasbt
3y ago
That's a good point. I may briefly mention RAG-like systems and add some literature references on this, but I am bit hesitant to give general advice because it's heavily project-dependent in my opinion. It usually also comes down
32.
▲
by
rasbt
3y ago
This is a good point. It's currently not in the TOC, but I may add this as supplementary text.
33.
▲
by
rasbt
3y ago
Unfortunately, I am not aware of any other resource that delves into these topics. However, as others commented above, Karpathy has a 2h YouTube video that is probably worthwhile watching. Based on skimming the YT video, it has some overlap
34.
▲
by
rasbt
3y ago
The attention mechanism we implement in this book* is specific to LLMs in terms of the text inputs, but it's fundamentally the same attention mechanism that is used in vision transformers. The only difference is that in LLMs, you turn
35.
▲
by
rasbt
3y ago
Glad to hear! I was thinking hard whether to write an intro to PyTorch for this book and am glad that this was useful!
36.
▲
by
rasbt
3y ago
Thanks! And please don't hesitate to reach out via the Forum or the GitHub Discussions if you have any feedback or questions.
37.
▲
by
rasbt
3y ago
Good question. I think a Python background is strongly recommended. PyTorch knowledge would be a nice to have (although I've written a comprehensive 40 page intro for the Appendix, which is also already available). From a math perspect
38.
▲
by
rasbt
3y ago
Thanks :)
39.
▲
by
rasbt
3y ago
Thanks for the support! There's the official Manning Forum for the book, but you are also welcome to use the Discussions page on the GitHub page.
40.
▲
by
rasbt
3y ago
The ETA for the last chapter is August if things continue to go well. It's usually available in the MEAP a few weeks after that, some time in September. And print version should be available early 2025 I think.
41.
▲
by
rasbt
3y ago
I'd say Chapter 1 would be the high-level intro to transformers and how they relate to LLMs.
42.
▲
by
rasbt
3y ago
On that note, I have a relative comprehensive intro to PyTorch in the Appendix (~40 pages) that go over automatic differentiation etc. The alternative, if you want to build something truly from scratch, would be to implement everything in C
43.
▲
by
rasbt
3y ago
Thanks for your support, I hope you'll get something useful out of this book!
44.
▲
by
rasbt
3y ago
Thanks, comparing positional encodings, MoEs, kv-caches etc are all good topics that I have in mind for either supplementary material and/or a follow-up book. The reason why it probably won't land in this current book is the lengt
45.
▲
by
rasbt
3y ago
Yeah, I don't think creating educational materials makes sense from an economical perspective, but it's one of my hobbies that gives me joy for some reason :). Hah, and 'insane amount of work' is probably right -- lots o
46.
▲
by
rasbt
3y ago
Sorry, in that case I would rather recommend a dedicated RL book. The RL part in LLMs will be very specific to LLMs, and I will only cover what's absolutely relevant in terms of background info. I do have a longish intro chapter on RL
47.
▲
by
rasbt
3y ago
I added notes to the Jupyter notebooks, I hope they are also readable as standalone from the repo.
48.
▲
by
rasbt
3y ago
That was pretty smooth. They reached out whether I was interested in writing a book for them (probably because of my other writings online), I mentioned what kind I book I want to write, submitted a proposal, and they liked that idea :)
49.
▲
by
rasbt
3y ago
Lol ok, otherwise it would probably be not very readable due to the verbosity. The book shows how to implement LayerNorm, Softmax, Linear layers, GeLU etc without using the pre-packaged torch versions though.
50.
▲
by
rasbt
3y ago
Haven't fully watched this but from a brief skimming, here are some differences that the book has: - it implements a real word-level LLM instead of a character-level LLM - after pretraining also shows how to load pretrained weights - i
51.
▲
by
rasbt
3y ago
It's in progress still. I have most of the code working, but it's not organized into the chapter structure, yet. I am planning to add a new chapter every ~month (I wish I could do this faster, but I also have some other commitment
52.
▲
by
rasbt
3y ago
I'd say my primary motivation is an educational goal, i.e., helping people understand how LLMs work by building one. LLMs are an important topic, and there are lots of hand-wavy videos and articles out there -- I think if one codes an
53.
▲
by
rasbt
3y ago
Glad to hear and thanks for the support. Chapter 3 should be in the MEAP soonish (submitted the draft last week). Will also upload my code for chapter 4 to GitHub soonish, in the next couple of days, just have to type up the notes.
54.
▲
by
rasbt
3y ago
It kind of is, but it's also kind of motivating :)
55.
▲
Implementing a ChatGPT-like LLM from scratch, step by step
(github.com)
739 points
by
rasbt
3y ago
|
98 comments
56.
▲
by
rasbt
3y ago
During training, it's more efficient than full finetuning because you only update a fraction of the parameters via backprop. During inference, it can ... 1) ... be theoretically a tad slower if you add the LoRA values dynamically duri
57.
▲
by
rasbt
3y ago
I think the main use case remains behavior changes: instruction finetuning, finetuning for classification, etc. Knowledge addition to the weights is best done via pretraining. Or, if you have an external database or documentation that you w
58.
▲
by
rasbt
3y ago
Yeah, the LoRA part is from scratch. The LLM backbone in this example is not, this is to provide a concrete example. But you could apply the exact same LoRA from scratch code to a pure PyTorch model if you wanted to: E.g. class Multil
59.
▲
by
rasbt
3y ago
Hah, yeah that's LoRA as in Low-Rank Adaptation :P
60.
▲
LoRA from scratch: implementation for LLM finetuning
(lightning.ai)
339 points
by
rasbt
3y ago
|
78 comments
More ›