24 ms·
NanoGPT
- bilsbie 4y agoReally cool. Can anyone answer these questions: Should I use this or minGPT? It says it needs 8XA100 40GB node. What is that and where do I acquire it? Could someone else train this and then send me the model? What would be required to run it as opposed to training it?
- vanpelt 4y agoA100’s are Nvidia GPU’s. You can rent them from providers like AWS or LamdaLabs. The readme has instructions for downloading the original GPT2 weights from OpenAI. You can also train a very simple version on a smaller dataset from your laptop as described in the README. If you just want to play with a similar but much better model goto https://chat.openai.com https://chat.openai.com
- theGnuMe 4y agoHow critical are training warmups and is an iteration here the same as an epoch?
- brossinthuon 4y agohttps://news.ycombinator.com/item?id=34336386 https://news.ycombinator.com/item?id=34336386
- aravindgp 4y agoThank you Andrej Karpathy for the work on ai and gpt models. It really helped me solve a problem as entrepreneur. I started making first few grand from ai.
- srge 4y agoMay i ask how? Consulting?
- rsiqueira 4y agoI could not find any sample (prompt and results). Can anyone provide samples of it's quality, even if it is in a narrow field of knowledge or specific use case? I tried GPT2, GPT-J 6B and GPT-NeoX 20B (implementation by Fabrice Bellard at textsynth.com/playground.html) but I could not find any production-quality scenario yet, only cherry-picked simple cases.
- boredemployee 4y agoThat's what I really miss to conclude if I should try it myself or not.
- visarga 4y agoAt this model size quality is not worth discussing. It is clearly another league from GPT-3.
- lossolo 4y agoIndeed, it is like comparing the speech of a 2-year-old child to that of a college professor.
- gpt-4 4y agoIs there a list of datasets like https://skylion007.github.io/OpenWebTextCorpus/ https://skylion007.github.io/OpenWebTextCorpus/ ?
- lukebechtel 4y agoI have taken several masters-level courses in Machine Learning -- and even with those credentials, I cannot recommend enough Andrej's youtube series, "Neural Networks: Zero to Hero". There, he teaches you, from scratch, how to build everything from the underlying automated gradient calculation system in pytorch, all the way up to the slower version of this model - `MinGPT`. [1] https://www.youtube.com/playlist?list=PLAqhIrjkxbuWI23v9cThsA9GvCAUhRvKZ https://www.youtube.com/playlist?list=PLAqhIrjkxbuWI23v9cThs... (edit: self-promo: I'm currently working on a Typescript follow-through of this same series of video lectures, if you want to follow along with stronger types for explanation: https://github.com/Marviel/lab-grad https://github.com/Marviel/lab-grad)
- randoglando 4y agoHow does it compare to fast.ai? As a engineer looking to learn, what should I start with?
- marviel 4y agoBoth are good for different things. Fast.AI is great, but it takes the top down, vs the bottom up, approach. It takes you from a production-level black box that you don't understand, down to the details. The benefit there is you get good high-level intuition of how it behaves at the "let me use this technology for a job" level. Separately, the fast.ai library is also highly recommendable -- it comes with some state-of-the-art image recognition models, and its training wrappers are really helpful particularly for image-recognition dataset training. Karpathy's "Neural Networks: Zero to Hero" video series starts at the level of individual neurons, and works you up to the final product. For some reason both this style, and Karpathy's conciseness appeal to me slightly more. I'm also super detail-oriented, though -- and any level of "hand waving" (even if further explanation comes later) always bothers me. He's also got some pretty high-profile industry experience which carries some weight with me. But I'll say that both are really high-quality. -- ultimately, my recommendation would be to follow whichever one speaks most to you personally after the first 1hr or so. EDIT: Per Jeremy's response below, if you want the bottom-up approach but like the fast.ai teaching style, you should check out "part 2" of the fast.ai set of tutorials, which is exactly that.
- marsven_422 4y ago[dead]
- 1986108 4y ago638c7215
- QuadrupleA 4y agoDoesn't huggingface have dozens of freely available pretrained models like this (including various sized implementations of GPT2) and isn't the source available on most if you wanted to train them yourself? All I see in the comments is praise for the author as a person, so just wondering what's unique about this that's not available elsewhere? 730 upvotes and counting, assuming I'm missing something...
- moneywoes 4y ago[flagged]
- minimaxir 4y agoAdditionally, in terms of the streamlining nanoGPT porports, HuggingFace's implementations play nice with optimization techniques such as ONNX/TensorRT, which will give you better performance than anything PyTorch-based even if minimal. That doesn't mean an ONNX-ed nanoGPT won't be better, but the field of optimized text generation isn't as new as people claim.
- isoprophlex 4y agoTrue, but the use cases arent the same. As he did before for other models, he has a knack for distilling the code down to beautiful, self-contained examples of high didactic value. It's an order of magnitude easier to grok the basics from this repo than from going through (admittedly more ergonomic or performant or production-ready) huggingface repos.
- visarga 4y agoThis is a didactic implementation. If you read the HuggingFace repo it is much more abstracted on account they implement many models in the same codebase. It's not fast or big, just easier to read and tweak.
- pms 4y agoIf so, then why the second line of its documentation says that "it is a rewrite of minGPT that prioritizes teeth over education"?
- surume 4y agoThank you so much for this! It is so impressive and I'm sure it took a lot of hard work! Is it able to re-write articles? And where could I find a guide on how to train it?
- karpathy 4y agoWow, fun to find this trending on HN this morning! I am currently also working on the associated video lecture (as the next episode of my video lecture series here https://karpathy.ai/zero-to-hero.html https://karpathy.ai/zero-to-hero.html ), where I will build nanoGPT from scratch and aspire to spell everything out, as with the earlier videos. Hoping to get it out in ~2 weeks or so.
- eternalban 4y agoThank you for sharing your knowledge. Anything that can be done to democratize machine learning is an invaluable social service. Hats off to you.
- katsucurry 4y agoI've found all of your code and lessons on youtube so incredibly useful. You're a wonderful teacher and I really appreciate all the work you've done with this!
- marviel 4y agoYour tutorials are effective and concise. Thank you for them! Accessible, from-scratch knowledge on these topics is essential at this time in history and you're really making a dent in that problem.
- de_nied 4y agoThank you for your constant contributions.
- gtoubassi 4y ago+1. I've benefited greatly from your content, e.g. your CNN lecture was incredibly accessible [0]. I still find transformers stubbornly elude my intuitions despite reading many descriptions. I would very much appreciate your video lecture on this topic. [0] I think https://www.youtube.com/watch?v=LxfUGhug-iQ https://www.youtube.com/watch?v=LxfUGhug-iQ
- cs702 4y agoAndrej: thank you! -- To the mod (dang): IMHO Andrej's comment should probably be at the top of the page, not my comment. UPDATE: Looks like that's done. Thank you :-)
- grogenaut 4y agoSomewhat off topic, does someone know how bing might integrate chat gpt into search. Is it to understand the prompt and filter results. Taking the question and summarizing it to search the index. Is it to summarize all the documents into an index and search that. Or to just be like chat gpt is now and use it to generate new results from it's knowledge base? I'm trying to connect the dots between a generative form like these are and how it would influence search in the future. Or is the lucene style index search on it's way out in a generative world?
- lossolo 4y ago> reproduces GPT-2 (124M) on OpenWebText, running on a single 8XA100 40GB node in 38 hours of training For comparison GPT-3 has more than 1000x more params (175B) and training time was around 2 months on ~1500 V100 GPUs which costs millions of dollars in cloud compute costs. Gopher with 280B params was trained on 4096 TPU-v3 chips, Microsoft Megatron-Turing NLG 530B trained on 2240 NVIDIA A100 cards (each card costs ~15k USD). And the most mind blowing is PaLM from Google with 540B params and trained on 6144 TPU v4, which costs around 10-30M USD in cloud compute to train.
- waiseristy 4y agoSo, are there any of these projects that aren't vendor locked to NVIDIA and are able to train large models with limited GPU RAM space? I don't mind letting my machine churn for 2-3 weeks. But I'm not looking to buy another 1000$ GPU just because CUDA is the only compute library researchers understand
- taylorius 4y agoWow, this is great. I can't wait for the video lecture, transformers are an aspect of modern machine learning that I'm not completely clear on. Andrej's lectures are brilliant - super detailed, and really answer the detailed questions I always have. Great stuff!
- jwithington 4y agoWhat would I google to figure out how to productionize the output of this? This repo trains a model--how would I prompt it and print the generated output?
- sexangel 4y ago[dead]
- Tepix 4y agoAnother cheap option is runpod According to https://www.runpod.io/gpu-instance/pricing https://www.runpod.io/gpu-instance/pricing renting out 4x A100 40GB costs $3.56 per hour. So that's $3.56 * 2 * 38hours = $270.56 then.
- awestroke 4y agoAre there any possible technologal or scientific leaps on the horizon that would reduce training time by an order of magnitude or more? GPT-3 took 355 years to train with incredibly expensive hardware, which means small players have no chance to push the state of the art
- make3 4y agosmall players will never have a chance to push the state of the art, as whatever optimization there is will also be applied at large scale with more money
- cypress66 4y agoA lot of SOTA comes from small players. It just isn't the case for LLMs.
- awestroke 4y agoGood point, but perhaps a leap could take small players into territories of language models that are large enough to be useful. GPT-3 crossed that threshold
- hankman86 4y agoTake a leaf from Seti@Home‘s book and try to come up with a distributed, volunteer-based approach to training an open source LLM. There is already an enormous amount of suitable ML hardware on end user devices.
- Der_Einzige 4y agoHuggingface actually recently did this, but I think it's for inference on their giant models like BLOOM
- noidiocyallowed 4y agoWork together and fuck up companies together. That's the way to go.
- legutierr 4y ago> The code itself is plain and readable: train.py is a ~300-line boilerplate training loop and model.py a ~300-line GPT model definition, which can optionally load the GPT-2 weights from OpenAI. That's it. What's the best source for these weights?
- benjamincburns 4y agoKaggle or HuggingFace
- iamflimflam1 4y agoI think the link should be: https://github.com/karpathy/nanoGPT https://github.com/karpathy/nanoGPT
- deleted 4y ago[deleted]
- mittermayr 4y agoAs someone who's been in software for almost 25 years now, I read through this in amazement of how much new stuff still keeps coming in. This industry never stops and that makes it such a fascinating (but arguably harsh) world to be in. Looking at this feels like seeing the source code of a 64k demo, learning about Mode 13h and trying to replicate it in Turbo Pascal. And, much like the old days of graphics programming, there's a good chance all of this knowledge will be mostly irrelevant soon, as the abstraction layers tend to come quicker and quicker and take care of the hard foundational work underneath. Then it'll be many of us here discussing whether or not it was good to have been with it from the start, to really get it, or whether playing with the highly-abstracted components is all that's needed to succeed with it. Either way, super cool to see the pace here and I loved the "I only have a macbook" section.
- eismcc 4y agoIt will be funny to look back from the future and think, wow, how did we get anything done with only 40GB RAM
- deleted 4y ago[deleted]
- henkdehenker 4y agoKarpathy is such a boss!
- yreg 4y agoIs there any trained model for text generation that you can run locally yet?
- throwaway743 4y agoPlenty. Huggingface alone has a ton
- deqwer 4y agoThere’s LAION working on open source[1] version of chatGPT [1] https://github.com/LAION-AI/Open-Assistant https://github.com/LAION-AI/Open-Assistant
- turmeric_root 4y agoThough their roadmap doc says they're looking into finetuning existing GPT-J/T5 models for this task. So you'll probably want a 3090 (24GB VRAM) and at least 16GB of CPU RAM to run inference if/when the project is complete.
- Metus 4y agoThis should be way higher up.
- wongarsu 4y agoGPT2 can be run locally (on a somewhat beefy consumer GPU)
- karmajuney 4y agoCan you add some info on what consumer GPU would be needed for this? Would a 3080 be able to handle this?
- wongarsu 4y agoAssuming you get the 12GB version of the 3080. A 2080TI is another option. Though you can reduce precision or use one of the smaller GPT2 versions to run on smaller cards as well.
- jamesfisher 4y agoFor casual readers like me: are there examples of what this can do once trained? E.g. it mentions training on Shakespeare, but gives no examples of fake Shakespearean.
- naasking 4y agoThe repo seems to imply that it matches GPT-2, so I imagine any analyses of GPT-2 will give you a good idea.
- kwerk 4y agoI’m not easily finding GPT-2 use cases. Any query guidance?
- fredoliveira 4y agoSomething that immediately comes to mind is text summarization. You'll by now be used to better results from GPT-3 or recent models, though.
- visarga 4y agoThe GPT family of models shines above 100B parameters. Almost nobody uses GPT2 today. It's too weak. If you want to go with <1B model, you use a BERT which is bidirectional or a T5 that is easier to fine-tune on other tasks.
- programmarchy 4y agoDoes anyone know the main differences between GPT-2 and GPT-3? Are there significant architectural changes, or is the advancement primarily from training?
- naasking 4y agoIf you google "GPT-2 vs GPT-3" you'll find lots of overviews and comparisons, like: * https://www.kdnuggets.com/2021/02/gpt2-gpt3-openai-showdown.html https://www.kdnuggets.com/2021/02/gpt2-gpt3-openai-showdown.... * https://bakztfuture.substack.com/p/the-chasm-between-gpt-2-and-gpt-3 https://bakztfuture.substack.com/p/the-chasm-between-gpt-2-a...
- cs702 4y agoAndrej doesn't need to do this. He's done it because he evidently loves it, and wants to share his hard-earned knowledge with the rest of the world. He may be a product of the ivory tower, but he's been in the trenches. He knows firsthand how f-ing hard it is to ship a product. And here he is, sharing useful personal code with everyone. This github repo now has collected ~4K stars and its antecessor (minGPT) has collected ~11K stars over the past couple of years. In my experience, the number of people who clone, copy, view or otherwise use code from a repo is one to two orders of magnitude larger than the number of people who star it, so we can safely say that Andrej has helped at least a few hundred thousand -- but likely more than a million -- individuals around the world learn how to build and tinker with GPT models. Remarkably, as I write this, no one else here has said thank you yet, so let me say it on everyone's behalf: THANK YOU ANDREJ. -- EDITS: I changed the wording in response to latexr's comments below.
- adam_arthur 4y agoI’m all for thanking open source contributors, but your excessively prostrating wording is a bit much for me.
- 1986108 4y ago[flagged]
- idiotsecant 4y agocan't tell if you're making some kind of clever quip or if this is some random spambot just entering a random reverse DNS lookup line.
- cs702 4y agoIf I overdid it, I'm sorry. I promise it wasn't intentional. My comment was spur-of-the-moment, motivated by sincere gratitude :-)
- mrg3_2013 4y agoThoughtful post! Everything so true! I am always amazed by individuals who truly are educators of the world.
- arturventura 4y agoThis is really good, and I was really excited by it but then I read: > running on a single 8XA100 40GB node in 38 hours of training This is a $40-80k machine. Not a diss, but I would love to see an advance that would allow anyone with a high end computer to be able to improve on this model. Before that happens this whole field is going to be owned by big corporations.
- windexh8er 4y agoBut how often do you need to run this? You can run 8xA1000 on LambdaLabs [0] (no affiliation) for $8.80/hr. So you should be able to run the entire data set for less than $350. [0] https://lambdalabs.com/service/gpu-cloud#pricing https://lambdalabs.com/service/gpu-cloud#pricing
- throwawaymaths 4y agoThey are acknowledged at the bottom for supporting andrej's research!!
- ProjectArcturis 4y agoThat's to train it from scratch, though, right? If you preload the GPT2 weights you don't need to do this. You can just give it additional training on your texts.
- anilshanbhag 4y agoIf GPT-2 / nanoGPT needs this setup, just imagine what GPT3 / chatGPT needs!
- Gigachad 4y agoSupposedly even running the trained model for ChatGPT is extremely expensive unlike the image generators which can largely be run on a consumer device.
- anigbrowl 4y agoWell, he does include instructions for running it on a personal computer, which looks like what I'm gonna be doing next week. Besides the rental options discussed below these nvidia boxen don't look too big so either used ones will be available for cheap relatively soon, or you could just locate and liberate one in Promethean fashion.
- albertTJames 4y agoCurious to know how close that training loop is to actual openai code.
- drjuice 4y ago[flagged]
- Terretta 4y ago14 hours ago: https://news.ycombinator.com/item?id=34331919 https://news.ycombinator.com/item?id=34331919 Curious why HN didn't merge the submission as it usually does. Is there a "no, submit this again" option?
- dang 4y agoHN lets reposts through if the story hasn't had significant attention yet. This is to give good submissions multiple chances at getting attention. Past explanations: https://hn.algolia.com/?dateRange=all&page=0&prefix=true&query=by%3Adang%20dupe%20detector%20porous&sort=byDate&type=comment https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que... I know it sucks when your submission was earlier and gets overlooked! We should eventually have some sort of karma-sharing to take care of this. In the meantime, it at least evens out in the long run, since the reason is randomness. https://hn.algolia.com/?dateRange=all&page=0&prefix=true&query=by%3Adang%20karma%20sharing&sort=byDate&type=comment https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que...
- eismcc 4y agoThe other post probably didn’t make it to the front page
- cpdomina 4y agoTo train small gpt-like models, there's also aitextgen: https://github.com/minimaxir/aitextgen https://github.com/minimaxir/aitextgen
- minimaxir 4y agoAs the creator of aitextgen, I'm mixed on continuing support since there doesn't seem to be as much demand as expected for small GPT models given the success and cost-effectiveness of GPT-3/ChatGPT, unfortunately. I still have a few ideas there (including another secret approach at better text generation) but it's hard to determine ROI.
- mboof 4y agoI think what you have created still has great demand. It give devs who do not have the budget or need for the gigantic models, something to train and use for their own specific language tasks. Not everyone is trying to replicate CHATGPT results for certain tasks.
- sebastianconcpt 4y agoWhat's the applicability? Can you give me some examples of what can be used this for?
- wongarsu 4y agoI imagine this might be interesting for domain-specific GPT models. Say training it on a mountain of technical documentation, or on every fanfiction published on the internet, or a sentiment analysis dataset. Of course fine-tuning GPT3 would give better results, but nanoGPT might allow you to make a much smaller model that's still good enough, to enable cheaper inference. Also the opportunity to play around with all the parameters fairly cheaply to find improvements. The todo section of the readme gives a small taste of that. Making bigger models works for OpenAI, but maybe the rest of us manage to make small models just perform better instead.
- justusthane 4y agoThis is a dumb question about language models in general, not necessarily specific to NanoGPT: why is all the focus on training? Can I download and run a pre-trained model locally? Surely the specs required to run a model are much, much lower than those required to train the model?
- nerdponx 4y agoIt's the equivalent of building from source versus downloading a compiled binary. Also you can perform "fine tuning" which means you start with a trained model and train it further on your own data, allowing you to customize the model for specific tasks.
- ausbah 4y agoinference can still be a bottleneck i think since you usually load the whole thing into memory which is 32-64GB+ usually?
- visarga 4y agoLanguage models range from 1 to 300+ GB when loaded. It depends on how you load them, if you load in int8 you get 4x reduction.
- anon291 4y agoIf you're only using pre-trained models, it's going to be harder to differentiate yourself. Training / specialization of models is where the moat-building is (due to access to different data sets / better ideas). By specializing / training, more of the token limit can be used for generation rather than prompting / better prompts can be made. The lower the cost of training, the more profitable any resultant business. You can even envision businesses that train the model regularly to bring in new knowledge. The cheaper this is, the more opportunities open up.
- code_runner 4y agoI believe the training is where the architecture of the model is most apparent. You can absolutely download plenty of pre-trained models. You will also probably need to fine tune for a specific use case, so a common approach is downloading a pre-trained model and fine tuning. I think including the “from scratch” tuning script is educational more than anything else.
- nprateem 4y agoIf I trained this on a 30,000 word document could it give me a summary? Or would there be no need to train it in that case, and I could just tell it "Summarise this: <insert 30,000 word document>"?
- londons_explore 4y ago30,000 words wouldn't be enough to train this from scratch - you'd ideally train from hundreds of millions of words at least. 30,000 words would be enough to finetune an existing model. If you did that, then the model would output text similar to the finetuning data. For example, if you finetuned it on shakespeare, then you might be able to use the model to make a new play, in shakespeare's style.
- ProjectArcturis 4y agoIf you finetuned it on the text of Shakespeare's plays, how would it link that text to the string "Shakespeare"?
- londons_explore 4y agoIt still has the knowledge from the main training on data from across the whole internet, so would still know the word Shakespeare... But you're right - the model finetuned on shakespeare would be good at writing a new play in the style of shakespeare, but would be bad at giving a critique of shakespeare's works.
- londons_explore 4y agoThe context window (block size) of this model is 1024 symbols. Symbols approximately map to words. So you can't ask it to summarize anything over 1024 words.
- nprateem 4y agoYeah that's the issue I was thinking of, how to get it to summarise large documents. Has anyone any ideas?
- imranq 4y agoI would love to see a minInstructGPT or a minRetro, or maybe something that combines instruction and retrieval into a readable codebase!
- siquick 4y agoExcuse my ignorance but what can a layman do with this?
- taneq 4y agoBecome less lay?
- deleted 4y ago[deleted]
- buzzdenver 4y agoFor an AI noob like me: can you use spot instances to train models? They are about 1/3rd the price on AWS compared to on demand ones, so it'd make a significant difference.
- satvikchoudhary 4y agoYes you should use them. They can be taken away from you with 2 min notice. (It doesn't happen a lot in practice though. I have been running a different instance for over a month. AWS doesn't force you if they don't have to) If you are going to run a long training job, ensure you are creating checkpoints. Be sure to use persistent storage, EBS and ensure that you check the option that it doesn't get deleted if the instance is stopped, so your checkpoint remain in the disk and you can easily restart. I haven't tried it but prices here are much cheaper. https://vast.ai/#pricing https://vast.ai/#pricing
- belter 4y agoYes you can. In Oregon you could eventually get this instance at $9. I say eventually, because of course Spot allocation is not guaranteed. ( And neither is On Demand ...but that is a story for another day) https://aws.amazon.com/ec2/spot/pricing/ https://aws.amazon.com/ec2/spot/pricing/
- yreg 4y agoWhy not? This is the exact use case of what Spot instances seem to be for. (Not hosting a service, but just calculating something for yourself.)
- swader999 4y agoWould it be possible to take all my user manuals and past customer Q&A and train on just on that to produce a customer helper chat bot?
- sharemywin 4y agoTo me this is the important quote: Unlike OpenWebText this will run in seconds. Finetuning takes very little time, e.g. on a single GPU just a few minutes. Run an example finetuning like:
- homarp 4y agosee also 'Cramming: Training a Language Model on a Single GPU in One Day' https://arxiv.org/abs/2212.14034 https://arxiv.org/abs/2212.14034 and https://github.com/JonasGeiping/cramming https://github.com/JonasGeiping/cramming
- jgalt212 4y agoSo is MSFT now extra grossly overpaying for ChatGPT?