7 ms·
LLM from scratch, part 28 – training a base model from scratch on an RTX 3090
- DeathArrow 10mo agoI think this is a very valuable exercise if you try to understand how LLMs work and if you have the time.
- rvnx 10mo agoSadly to go beyond an exercise, having the money is really what you need if you actually want LLMs now, not time. Nowadays training very powerful LLMs is easy because all the tooling, source-codes, training datasets, and teaching agents are available. Getting access to dozens of millions of USD or more is not easy, and for big players this is a just drop in their ocean.
- contrast 10mo agoYou seem to be talking about a production-grade model rather than building an LLM as an exercise? Or if not, why do you disagree with the article's example of building a small LLM for $100?
- rvnx 10mo agoI think I should have replied as a totally separate comment. This is my mistake. It is nice that the author shared the results of his exercise / experiment. Just got sad as I was reminded (when the 100 USD were mentioned) that all this game is 90%+ about money and hardware rather than skills. That being said I really like the initiative of the author.
- DeathArrow 10mo agoThat is true for many kinds of software where you need a big amount of resources. No matter how skilled I am, I cannot build Facebook, Google, Photoshop alone. But a tiny version of it just to learn? Why not!
- victorbjorklund 10mo agoYou could 100% build Facebook. You don’t need any hardcore hardware before you have many users.
- meehai 10mo agoit's skills first and then money and hardware for scale A more skilled person that understands all the underlying steps will always be more efficient in scaling up due to knowing where to allocate more. basically... you always need the skills and the money is the fine tuning.
- jbs789 10mo agoI understand the emotional aspect of feeling like it’s out of reach for you. Thing is, if you focus on your own skill development and apply it at even a small scale, very few people do that. Then you go for a job and guess what, the company has resources you can leverage. Then you do that, and ultimately you could be in a position to have the credibility to raise your own capital. Play the long game and do what you can do now.
- rvnx 10mo agoNot at all. The majority with the current AI craze not really about credibility or skills. It's like a kitchen. Take a genius chef but give him rotten ingredients. He sweats, he tries, but the meal is barely edible. That's the $100 exercise, but only experts recognize the talent behind. Take an unskilled cook but give him A5 Wagyu and prepared truffles. The result tastes amazing to the average person who will claim the chef is great (the investors). It's about access to capital and selling a story ('ex'-Googler doesn't make you competent), not skills. Great chefs in dark alleys go unnoticed. Mediocre tourist traps near the Eiffel Tower are fully booked. Look at Inflection AI. Average results, yet massive funding. They have the "location" and the backing, so they win. It's not about who cooks better; it's about who owns the kitchen but who sells a dream that tomorrow the food will be better. We don't talk about small funding, we talk about 1.3 billion USD, just for that specific example, yet a tourist trap (using name-dropping / reputation instead of talent) Snake-oil is rewarded as much as, or even more than real talent; a lot of people cannot see the difference between a chef and the ingredients, this is what I think is sad.
- YouAreWRONGtoo 10mo ago[dead]
- victorbjorklund 10mo agoTotally. While the LLM:s today are amazing it is a bit sad that you can’t build SOTA models on your own (vs a few years ago where someone with the skills and access to a dataset could build a state of art models)
- Chabsff 10mo agoIn the grand scheme of things, we've only had about a quarter century where you needed a *very* specific kind of problem where prosumer hardware wasn't adequate across computer science as a whole. It's kind of amazing we got that at all for a while.
- djmips 10mo agoIf you discard the early days of gigantic expensive computers. I guess it's come full circle after a fashion.
- ducktective 10mo agoAre off-shelf GPUs (like one 3090) suitable for modern academic research on current AI advancements or is it better to rent some cloud compute?
- i5heu 10mo agoIt depends on what you want to do in this gigantic field.
- ACCount37 10mo agoResearch runs on a variety of scales - but "check if this new idea/method/architecture isn't completely dumb on small scale before trying to scale up" is a common enough pattern. And most of those fail on small scale.
- htrp 10mo agodepressingly enough, things that work on small scale architectures often don't work at larger scales
- ACCount37 10mo agoYep, most of what's remaining fails to scale. But it's still a very solid filter. Sure, there are things that don't work on small scale and then work on large scale. But they're rare, and they sure are going to be expensive to find and validate.
- lynndotpy 10mo agoIf you're seriously doing deep learning research, it's very very nice to own your own GPU. For four years of AI PhD research I worked with a 1050Ti on a personal laptop and a 2060 on a personal desktop. You can do a lot of validation and development on consumer GPUs. That said, the OP does not train an LLM from scratch on a 3090. That would not be feasible
- joefourier 10mo ago
- Havoc 10mo ago> When you’re looking at a pre-training dataset in the frontier lab and you look at a random internet document, it’s total garbage. I don't even know how this works at all. It’s [stuff] like stock tickers, symbols, it's a huge amount of slop and garbage from like all the corners of the internet Seems like there would be low hanging fruit in heavier pre processing then? Something deterministic like a reading level score. Or even a tiny model trained for the task to pick out good data?
- haolez 10mo agoIf you can create this filtering model, you have created Skynet and solved AGI :D
- ACCount37 10mo agoData filtering. Dataset curation. Curriculum learning. All already in use. It's not sexy, it's not a breakthrough, but it does help.
- Havoc 10mo ago> All already in use. At the big labs that makes sense. Bit more puzzled by why it isn’t used in the toy projects. Certainly more complexity but seems like it would make a big difference
- famouswaffles 10mo agoCurriculum learning is not really a thing for these large SOTA LLM training runs (specifically pre-training). We know it would help, but ordering trillions of tokens of data in this way would be a herculean task.
- ACCount37 10mo agoI've heard things about pre-training optimization. "Soft start" and such. So I struggle to believe that curriculum learning is not a thing on any frontier runs. Sure, it's a lot of data to sift through, and the time and cost to do so can be substantial. But if you are already planning on funneling all of that through a 1T LLM? You might as well pass the fragments through a small classifier before you do that.
- RagnarD 10mo agoI really like this article. I hadn't thought that an RTX 3090 would be capable of generating a sort-of decent small LLM from scratch in a reasonable time, but he shows how in detail.
- billylo 10mo agoIf you are curious about doing something similar with TPU, Google has an article. https://developers.googleblog.com/train-gpt2-model-with-jax-on-tpu/ https://developers.googleblog.com/train-gpt2-model-with-jax-...
- nullbound 10mo agoI love the level of detail ( probably, because I see it less and less these days ). It genuinely makes me wonder if anyone tried training LLMs on their own writings ( assuming those bigger than 100+ pages ) and what the results were.
- jadbox 10mo agoI just want to chime in here about the importance of taking notes and having a journal. These things are now more important than ever as they can literally help fine-tune agents to help assist you using your personal style.
- trial3 10mo ago> These things are now more important than ever oh definitely. i agree here. can't wait to read the rest of the sentence, probably saying something meaningful about the creative benefits of unstructured writing, or the importance of relying on your own thoughts and language and unique voice in the era of LLMs > as they can literally help fine-tune agents to help assist you using your personal style. oh
- jadbox 10mo agoI get it. Both things can be true. Unstructured writing can help you develop as a person. It can also teach your own model the 'real raw human train of thoughts' of your personal journey. Personally I love the idea of booting up great-great-grandpa-model that'll have been trained on his 40 years of almost daily journaling. We are not trying to 'remake him' to be clear- we are talking about being have to have an interaction chat with his personality-vibe as it was recorded by his own hand and in his own words.
- SecretDreams 10mo agoIs this what tool and die makers used to feel when going to LOC to train their replacements? Personally, I do not want my likeness to persist after my death, nor do I wish for a company to be able to leverage my likeness after I leave said company.
- lepicz 10mo agocool, i was looking for something like this to try on my own puny hw - thanks!
- BubbleRings 10mo ago> …reused its embedding matrix as the weights for the linear layer that projects the context vectors from the last Transformers layer into vocab space to get the logits. At first glance this claim sounds airtight, but it quietly collapses under its own techno-mythology. The so-called “reuse” of the embedding matrix assumes a fixed semantic congruence between representational space and output projection, an assumption that ignores well-known phase drift in post-transformer latent manifolds. In practice, the logits emerging from this setup tend to suffer from vector anisotropification and a mild but persistent case of vocab echoing, where probability mass sloshes toward high-frequency tokens regardless of contextual salience. Just kidding, of course. The first paragraph above, from OP’s article, makes about as much sense to me as the second one, which I (hopefully fittingly in y’all’s view) had ChatGPT write. But I do want to express my appreciation for being able to “hang out in the back of the room” while you folks figure this stuff out It is fascinating, I’ve learned a lot (even got a local LLM running on a NUC), and very much fun. Thanks for letting me watch, I’ll keep my mouth shut from now on ha!
- jcims 10mo agoThe turbo encabulator lives on.
- empath75 10mo agoIt's a 28 part series. If you start from the beginning, everything is explained in detail.
- ekropotin 10mo agoI have no idea what you’ve just said, so here is my upvote.
- tomrod 10mo agoDisclaimer: working and occasionally researching in the space. The first paragraph is clear linear algebra terminology, the second looked like deeper subfield specific jargon and I was about to ask for a citation as the words definitely are real but the claim sounded hyperspecific and unfamiliar. I figure a person needs 12 to 18 months of linear algebra, enough to work through Horn and Johnson's "Matrix Analysis" or the more bespoke volumes from Jeffrey Humpheries to get the math behind ML. Not necessarily to use AI/ML as a tech, which really can benefit from the grind towards commodification, but to be able to parse the technical side of about 90 to 95 percent of conference papers.
- chiengineer 10mo agoOff topic question since im not a regular here if its ok Is anyone here actually using the 200$ a month subscriptions with chat gpt or the google 150$ per month ? Is it worth it for more code generation ? Or spend my money on a couple gpus and go local
- Taek 10mo agoI used the $200/mo OpenAI subscription for a while, but cancelled when Gemini 3 came out. It was useful for the deep research credits until the Web search gpt got sufficiently good on it's own
- esafak 10mo agoTo answer the last question: What kind of programming do you do? You are not going to be able to run a model competitive with the SOTA yet; use the cloud. Since you have the budget I'd suggest getting a $20 subscription of each (Claude, Gemini, ChatGPT) so you can lean on their respective strengths.
- magicalhippo 10mo agoI got a free month of the Premium tier with Google[1], YMMV. Been pleasantly surprised about Gemini 3 Pro. Got ChatGPT Business at work to compare it to. That said, Google's VSCode integration was terrible, kept logging me out and just didn't work well. [1]: https://one.google.com/about/plans https://one.google.com/about/plans
- logicallee 10mo agoyou can train an LLM in the browser, see this demonstration: https://taonexus.com/mini-transformer-in-js.html https://taonexus.com/mini-transformer-in-js.html It's a very simple neural network with two attention heads that runs right in the browser in pure Javascript, you can view source on this implementation. Even after training for a hundred epochs it really doesn't work very well (you can test it in the Inference tab after training it), but it doesn't use any libraries, so you can see the math itself in action in the source code.
- spi 10mo agoThis is a very nice, detailed post! I have a few minor comments though (maybe a few are discussed somewhere, it's a _long_ article and I can't claim 100% coverage :-) ): Calling it "training LLM" is a bit misleading. This is a small GPT-2-sized model (~160M params), while the "L" in "LLM" stands for large... The early discussion and worries about truncating strings look a bit weird. The author then realizes they're anyway not even going to use 30% of the total available data, so who cares if for each given string we're only using the first 1024 tokens? (And anyway, even if doing more epochs, he doesn't discuss the obvious solution to avoid throwing away data, i.e. not clipping always the tail but starting from a random point each epoch - maybe after a punctuation or something) At this level of simplicity, setting up a validation loop might be an unneeded complication (for the autoregressive pretraining part, not the instruction-tuning of course). That's because anyway the model is training for < 1 epoch, so no data is seen twice (*). One might as well just track the training loss, it's slightly less "clean" because it's evaluated each time on different data, but the sheer size of it makes up for the issue. The final plot shows that the two curves are similar - train is noisier of course, but nothing a bit of rolling smoothing couldn't solve. The choice to load all tokenized text into RAM feels odd... it works, and it's possibly slightly faster than loading on-the-fly, but only if you have enough RAM to "waste". PyTorch loads data on separate processes in a non-blocking way, so it feels like having it on disk and loaded on-the-fly would be safer and not make any hit on runtime. But well, if it fits, it's certainly easier that way (although, as the author remarks, it only works if you can store it as a numpy array or torch tensor of some internally supported dtypes like int or float; if they are any Python "object" types, they get replicated per dataloader worker, and OOM is guaranteed) The choice to concatenate everything into a long string is a bit outdated nowadays. Because it trains with attention between different sentences that have nothing to do with each other, and could cause a bias or anyway suboptimal results. Nowadays people use masked attention ("document masking"), which is so popular it's even supported by FlashAttention: https://github.com/Dao-AILab/flash-attention/issues/654 https://github.com/Dao-AILab/flash-attention/issues/654 (*) Of course, the data is dirty enough that there _will_ be some duplicated stuff here or there, but the same is true for a random train/validation split. Also such a small model would have very little risk to memorize, even if some data were replicated.*
- BoxOfRain 10mo ago
- kburman 10mo agoAnyone interested can also follow these amazing playlists: 1. Building LLMs from scratch - https://www.youtube.com/playlist?list=PLPTV0NXA_ZSgsLAr8YCgCwhPIJNNtexWu https://www.youtube.com/playlist?list=PLPTV0NXA_ZSgsLAr8YCgC... 2. Reasoning LLMs from Scratch - https://www.youtube.com/playlist?list=PLPTV0NXA_ZSijcbUrRZHm6BrdinLuelPs https://www.youtube.com/playlist?list=PLPTV0NXA_ZSijcbUrRZHm... 3. Build a SLM from Scratch - https://www.youtube.com/playlist?list=PLPTV0NXA_ZShuk6u31pgjHjFO2eS9p5EV https://www.youtube.com/playlist?list=PLPTV0NXA_ZShuk6u31pgj... 4. Build DeepSeek from Scratch - https://www.youtube.com/playlist?list=PLPTV0NXA_ZSiOpKKlHCyOq9lnp-dLvlms https://www.youtube.com/playlist?list=PLPTV0NXA_ZSiOpKKlHCyO...
- youngNed 10mo agoThese all look great, I'm very interested in hearing from anyone who has followed any of these. How did you find it, what did you get from it?
- spi 10mo agoA separate comment about conclusions about why they are worse than OpenAI GPT2 - which to me feel to be missing the point. One main point is batch size - I'd agree with Gemini here. Batch size <= 5 with 1024 seq len is really tiny. Nowadays models are trained with effective batch size of millions of tokens in total. Of course, this won't fit into memory, one uses gradient accumulations to that purpose, again as mentioned by Gemini. Training duration is definitely also a reason - models do get better over time, otherwise people wouldn't train so long wasting millions :-) just how long for optimality is unclear, but certainly < 2 days is not optimal even at this "small" scale. The optimizer could also play a role. As the author mentions, a fixed learning rate is hardly optimal, it is typically both increased in the beginning ("warm up", but that's for stability, if training works without, that's not an issue) and scaled down at the end ("cool down" - that is, annealing, with cosine as mentioned in the article). This generally squeezes out a bit more performance. Also, while it's true that dropout was used back then (might be useful for many epochs, likely only harmful for < 1 epoch), using _both_ dropout _and_ weight_decay > 0, as the author does, is probably wrong and makes training too slow & careful to get good results. Also, even if used, a "good" implementation of weight decay should skip some layers like embeddings and biases (GPT2 did that, and it's relatively important to do so). On the other hand, I'm pretty sure that using mixed precision and TF32 has absolutely no downsides. It's really standard nowadays to use either mixed precision (FP16 gradients + FP32 base weights) or directly BF16 ("brain" float 16, a bit like the TF32 described there, but with only 16 bits) and I have almost never seen either one fail... and when it does, it typically fails spectacularly, with NaN losses or the model degenerating to trivial performance.
- gpjt 10mo agoOP here -- thanks! I'm in the process of doing some trains using the same code plus DDP on big Lambda Labs machines, and (within the bounds of what I can afford) will hopefully have some interesting results about all of those shortly.
- gpjt 10mo agoOK, early indicators support both you and Gemini quite strongly re: batch size. On my (somewhat ad-hoc) test dataset, I get losses like this: * OpenAI medium weights: 3.231 * OpenAI small weights: 3.500 * My locally trained model, FineWeb Chinchilla, batch size 6: 3.944 * My locally trained model, FineWeb-Edu Chinchilla, batch size 6: 4.167 * My locally trained model, FineWeb-Edu double Chinchilla, batch size 6: 4.135 * My cloud trained model, FineWeb Chinchilla, batch size 13 \* 8 = 104: 3.674 That last one was trained on an 8x A100 machine with 40 GiB per GPU, with the same code as before, just converted to DDP. It certainly looks like the much larger batch size has improved the model significantly. I'll be trying on larger machines. No gradient accumulation yet, but it's certainly looking like a valuable lever to pull for local training runs (and, I suspect, might also be useful on "small" cloud machines like the one I used -- will have to see what things look like with the bigger mini-batches I can squeeze onto 80 GiB and 160 GiB GPUs).
- roschdal 10mo agoNow this is cool. and can be used for evil AI.
- nico 10mo agoHas anyone done something like this but with apple silicon instead of a graphics card? Training a small LLM on an M2-M5?
- muricula 10mo agoI've played with something similar with my M1 using Apple's MLX framework. The problem is I'm compute bound. I've never managed to get my M1 Max's GPU to process more than ~7.8k tokens per second at bf16 precision, so to train a 112M parameter model on ~20 billion tokens I'd need to run the model training for ~30 days. One solution is to reduce the scope of the problem -- you can train on a smaller less diverse dataset such as TinyStories which is a collection of 1 billion tokens of chatGPT generated children's stories. After about 40 hours, less than one weekend, you'll have a model which can generate mostly grammatical children's stories. If you have a newer mac and/or an ultra chip you'll have more and faster GPU cores, and might be able to train on FineWeb or a similar, larger and more diverse dataset.
- gpjt 10mo agoOP here -- with a 112M model you should be able to get something worth playing with using 2.24B tokens. The Chinchilla heuristic is tokens = 20 x parameters. Obviously you cam get a better result by grinding through more tokens, but it will be very slow progress. It's worth noting that Andrej Karpathy is using the 20x thing for his nanochat project. I try to explain the Chinchilla paper in the post, but your favourite AI should be able to explain it well, and has the benefit that you can ask follow-up questions.
- goosers 10mo agoI’m experimenting with this, but using the CPU not the GPU. I’m finishing up writing the series now, but focused more on understanding the architecture than trying to build a useful model. Mine requires talking in the language of Shakespeare, and getting replies in the same, a proof of concept more than a useful tool. https://www.tag1.com/white-paper/part1-tokenization-building-an-llm-from-scratch-in-rust/ https://www.tag1.com/white-paper/part1-tokenization-building... I was interested in focusing on repeatability and using text sources anyone can legally obtain. It’s been fascinating, but after much experimentation it’s clear that working with more text and more diverse text would be extremely helpful.
- pwython 10mo agoFor those that have homebrewed a base model, does your output have the same AI-isms like overusing em dashes? If so/not, what dataset did you use?
- itissid 10mo agoDoes yours also use the oxford comma and generally more commas?
- miki123211 10mo agoAFAIK, those are mostly a consequence of posttraining.
- whimsicalism 10mo agothat is a post-training artifact
- pixigenie 10mo agothanks for sharing
- lacoolj 10mo agoMaybe I've been missing out, but can anyone give me a yay/nay on whether this is a worth-while 28-part-series to start from scratch and spend my time watching/reading? Is it along the same lines as https://github.com/karpathy/llm.c/discussions/677 https://github.com/karpathy/llm.c/discussions/677 ? He (karpathy) has a video series that also does something similar. I found it very informative and entertaining, even at the 1 hour + length it is (there are actually multiple videos, im not sure how long the others are).
- nfriedly 10mo agoThe full list of articles is at https://www.gilesthomas.com/llm-from-scratch https://www.gilesthomas.com/llm-from-scratch for anyone who's interested but wants to start at the beginning.
- fuddle 10mo agoThis is great to see, I'm also re-reading Sebastian Raschka's amazing book.
- noloman 10mo agoGreat article, thanks!
- noloman 10mo agoGreat article