6 ms·
Hi HN the main (more detailed) article is here https://github.com/karpathy/llm.c/discussions/481 https://github.com/karpathy/llm.c/discussions/481 Happy to an
by karpathy 2y ago
Hi HN the main (more detailed) article is here
https://github.com/karpathy/llm.c/discussions/481 https://github.com/karpathy/llm.c/discussions/481
Happy to answer questions!
- 1024core 2y agoThank you, from an appreciative reader!
- ngiyabonga 2y agoHi Andrej! First, thank you for your teaching, it has helped me a lot, didn't think I'd ever have the chance to say thank you, but here you are and I hope this gets to you! Question - what's a relevant (05-2024) baseline to compare the performance of c code to? Back when you made nanoGPT you were seeing "the file train.py reproduces GPT-2 (124M) on OpenWebText, running on a single 8XA100 40GB node in about 4 days of training". So twice the memory on the c node, but unsure of data size /epochs, any other details I may be missing. I.e. what's the net uplift of running c vs "legacy" torch code? Thanks again for everything.
- karpathy 2y agoThe baseline is definitely PyTorch (or JAX), and indeed something like nanoGPT. I just never got nanoGPT "past the finish line" of really crossing the t's and dotting the i's and reproducing the models with as much care as I did now and here in llm.c, and getting to the point where it's a single launch command that just does the thing. I think I'll try to develop the `train_gpt2.py` inside llm.c to be that, so that we have the two implementations exactly side by side, and it's all nice and comparable. The C/CUDA code is currently a little bit faster than PyTorch (last time I measured ~2 weeks ago it was about 6% faster), and I think we can push this further. This is done by manually hard-coding a bunch of fusions/optimizations that are non-trivial for torch.compile to find (e.g. our FusedClassifier). But PyTorch has some pending work/PRs that will also speed up their side a lot. Ultimately my interest in llm.c is to have a nice, clean, minimal, super dependency-light repo in direct C/CUDA implementation, which I find aesthetically pleasing. And on top of that, educational, i.e. using all of the above as an endpoint of an intro LLM course.
- ilaksh 2y agoJust out of curiosity, how do you feel about Tinygrad? They just released 0.9 and are also on the HN home page today.
- raymond_goo 2y agoMaybe talk to MasterClass...
- sturza 2y agoDo you think grokking leads to proper generalized reasoning? https://arxiv.org/abs/2405.15071 https://arxiv.org/abs/2405.15071
- espadrine 2y agoHow big of a perf improvement would result from using the architectural tweaks that Llama3 and others have put in place since GPT-2?
- karpathy 2y agoMy understanding and suspicion is mostly less than you think. Llama 3 architecture has the following changes on GPT-2: 1. delete the absolute positional encoding and replace with RoPE 2. delete all biases in all layers (in LayerNorms, they turn into RMSNorm) 3. GeLU -> SwiGLU non-linearity in the MLP 4. longer context length 5. architecture hyperparameter changes, e.g. slightly different aspect ratios And there was a paper that I can't find the reference to anymore that claimed that if you train long enough, the gap becomes even lower. Possibly because the absolutely positional encoding has enough time to train more fully, where as the RoPE layer benefits from the "inductive bias" it adds in the earlier stages of training. But I don't have full confidence on the above claim, maybe someone has tried or has better/concrete reference.
- jorlow 2y agoNote llama's feed forward is a bit different too: self.w2(F.silu(self.w1(x)) * self.w3(x)) I.e. the nonlinearity is a gate. https://github.com/meta-llama/llama3/blob/14aab0428d3ec3a9596f1dea06d9c564f9c0e35f/llama/model.py#L219 https://github.com/meta-llama/llama3/blob/14aab0428d3ec3a959...
- soraki_soladead 2y agoFwiw, that's SwiGLU in #3 above. Swi = Swish = silu. GLU is gated linear unit; the gate construction you describe.
- lagrange77 2y agoThank you for the effort you put in your educational work, it helped me and others a lot! In fact, i'm training my nanoGPT version right now. :) > Ultimately my interest in llm.c is to have a nice, clean, minimal, super dependency-light repo in direct C/CUDA implementation, which I find aesthetically pleasing. Also, it's awesome that you spend your time on your passion. Any plans on making a video series on llm.c? :D
- karpathy 2y agoYes definitely. Related tweet of mine: https://x.com/karpathy/status/1760388761349927356?lang=en https://x.com/karpathy/status/1760388761349927356?lang=en 1. Build the thing 2. Build the ramp Currently on step 1 :). It helps to build it first so you know where you are going, and then you can more easily re-build it when you're vector pointed at the end result.
- lagrange77 2y agoThat's fantastic. My gradient field is pointing towards it. Thank you again!
- htrp 2y agoEverytime you take gardening leave, you build something new and interesting!
- LorenzoGood 2y agoI love when you leave your job.
- 363849473754 2y agoYou might have covered this topic before, but I'm curious about the main performance differences between nanoGPT and llm.c. I'm planning to take your "Zero to Hero" course, and I'd like to know how capable the nanoGPT chatbot you'll build is. Is its quality comparable to GPT-2 when used as a chatbot?
- karpathy 2y agoZero To Hero doesn't make it all the way to a chatbot, it stops at pretraining, and even that at a fairly small scale or character-level transformer on TinyShakespeare. I think it's a good conceptual intro but you don't get too too far as a competent chatbot. I think I should be able to improve on this soon.
- 363849473754 2y agoThanks! So, you are considering expanding the Zero to Hero series to include building a basic GPT-2 toy chatbot? I believe you mentioned in one of the early lectures that you planned to include building a toy version of Dalle. Do you still have plans for that as well?
- maskil 2y agoPlease do! It's a fantastic series!
- dang 2y agoOk, we've changed the URL to that from https://twitter.com/karpathy/status/1795484547267834137 https://twitter.com/karpathy/status/1795484547267834137 above. Thanks!
- karpathy 2y agosounds good. both work, (though) I think HN has a bit of an anti-twitter bias.
- pests 2y agoFirst, love the videos and other work you've been doing. The micrograd videos are a great way to show people this is all math in the end, and I've linked to specific timestamps in that video and others more times than I can count. For why I think we have a anti-twitter bias... Twitter doesn't show replies or any further context without being logged in. Most people will have accounts but I know a lot here deleted theirs or refuse to use it for one reason or another. Also IMO most here are going to want to read the full source so it just cuts out the middleman. This would usually fall under the "Please submit the original source. If a post reports on something found on another site, submit the latter." guideline which is a little different since the source is yourself, but still the Twitter post doesn't add anything new or novel.
- karpathy 2y agofwiw I totally understand the sentiment! it's actually a bit sad to me that so much of our content is moving from the shared, open web to platforms like twitter, unfortunately there seems to be too much value add around built-in discoverability, comments, ease of authoring, for many people revenue sharing, etc.
- pests 2y agoYes, definitely. I had to double check your age (apologies! feels rude somehow) and yep, we're basically the same age. The web was different back then. Maybe not better; maybe that's nostalgia. But never before has more creators had as many tools and avenues to promote and monotonize their work as they do now.
- m11a 2y agoWhy write in CUDA and not just use PyTorch etc? if performance, how much faster is it, out of curiosity?
- kgwgk 2y ago> Why write in CUDA and not just use PyTorch etc? “LLM training in simple, pure C/CUDA. There is no need for 245MB of PyTorch or 107MB of cPython. […] A few more words on what I want this repo to be: First, I want llm.c to be a place for education.”
- simonw 2y ago> Keep in mind that here we trained for 10B tokens, while GPT-3 models were all trained for 300B tokens. [...] GPT-3 actually didn't change too much at all about the model (context size 1024 -> 2048, I think that's it?). Andrej, based on that do you have a rough cost estimate for what it would take to train a GPT-3 Ada (350M)? Do you plan to get there with llm.c ?
- karpathy 2y agoThe 350M model I trained last night was 30B tokens, 14 hours, ~$200. Conveniently, 300B is exactly 10X the tokens so ~$2K would be the estimate. You'd have to wait 140 hours on one box though. Getting an H100 box instead of A100 will already cut the time latency down probably by a factor of 2-3X, for free, even without going to fp8 (which we do plan to support). So TLDR at this model scale, llm.c is already there functionally, I think, it's a matter of the compute resources and patience. I currently have this one box from Lambda and I have to look around for a few more boxes and merge the pending PR for multi-node training support. Getting all of this into a nice, stable state is probably a good chunk of the pending work right now.
- deleted 2y ago[deleted]
- localhost 2y agoHow large is the set of binaries needed to do this training job? The current pytorch + CUDA ecosystem is so incredibly gigantic and manipulating those container images is painful because they are so large. I was hopeful that this would be the beginnings of a much smaller training/fine-tuning stack?
- karpathy 2y agoThat is 100% my intention and hope and I think we are very close to deleting all of that. Right now on master, I am already only using Python for the tokenization preprocessing. In principle the requirements for llm.c should be extremely minimal. I think this a few days of work that is high on my mind. Biggest problem right now is finding a place that can host the 135GB of tokens for FineWeb100B. Will probably use S3 or something. Related see: https://github.com/karpathy/llm.c/issues/482 https://github.com/karpathy/llm.c/issues/482
- metadat 2y agoCould this be a good case for a torrent?
- deleted 2y ago[deleted]
- dekhn 2y agoWould you consider switching your interest to protein structure prediction? In particular, the current most advanced model is a closed-source, closed-weights system that was trained on a proprietary hardware. It is intentionally kept that way for now to enable deepmind to commercialize their product. The goal here isn't to make the best performing model: it's ablation. How much can we remove from protein structure prediction (such as multiple sequence alignments and molecular dynamics, which were two improvements in AF3), while still having a generalized model that can predict novel folds. Then focus on teaching the minimal necessary math and code to reproduce the results to the larger biological community. All I can say about AF3 is that it literally taught me that everything I learned about protein structure prediction in the last 30 years was misguided, or outright wrong. Don't worry about drug discovery or any of the hard stuff. Just continue to show that all that's required to predict novel structures is the existing PDB.
- treme 2y agolol I appreciate your effort to guide his genius towards 'max human good'
- wizzwizz4 2y ago> switching your interest That's not usually how it works. > Just continue to show that all that's required to predict novel structures is the existing PDB. Sounds like you know a lot about this topic. You should do it!
- dekhn 2y agoYes I already published several papers in the area, but I don't work on it any more.
- jonesn11 2y agoLike the FAQ, you correctly anticipated my questions.
- 0x1ceb00da 2y agoHi. Is it possible to somehow run llm.c on an amd gpu?
- anthonix1 2y agoYeah, I just reproduced the GPT2 from scratch results in 8.75 hours on 4x 7900 XTX. The fork is here: https://github.com/anthonix/llm.c https://github.com/anthonix/llm.c
- sytelus 2y agoSo, NanoGPT took 1.8 days on 8xA100 for 124M model training on 30.7B tokens using flash attention. This would translate to 14.4hr for 10B tokens. With llm.c it is ~1.5 hr which is almost 10X speedup! Does this look ballpark correct? Is there any summary of where majority of this improvement comes from?