12 ms·
I might be too new to this area -- but is this actually explaining how to create like a small version of the actual trained model -- not like "using the trained
by xt00 4y ago
I might be too new to this area -- but is this actually explaining how to create like a small version of the actual trained model -- not like "using the trained model for X"? like I can imagine in the future people won't start from pure scratch, there will be building blocks that everybody starts from, but mostly just wondering like how hard is it to actually replicate what openAI has done if you had the money to pay for the training?
- karpathy 4y agorough steps: 1. collect a very large dataset, see: https://www.lesswrong.com/posts/6Fpvch8RR29qLEWNH/chinchilla-s-wild-implications https://www.lesswrong.com/posts/6Fpvch8RR29qLEWNH/chinchilla... . scrape, de-duplicate, clean, wrangle. this is a lot of work regardless of $. 2. get on a call with the sales teams of major cloud providers to procure a few thousands GPUs and enter into too long contracts. 3. "pretrain" a GPT. one common way to do this atm is to create your own exotic fork of MegatronLM+DeepSpeed. go through training hell, learn all about every possible NCCL error message, see the OPT logbook as good reference: https://github.com/facebookresearch/metaseq/blob/main/projects/OPT/chronicles/OPT175B_Logbook.pdf https://github.com/facebookresearch/metaseq/blob/main/projec... 4. follow the 3-step recipe of https://openai.com/blog/chatgpt/ https://openai.com/blog/chatgpt/ to finetune the model to be an actual assistant instead of just "document completor", which otherwise happily e.g. responds to questions with more questions. Also e.g. see OPT-IML https://arxiv.org/abs/2212.12017 https://arxiv.org/abs/2212.12017 , or BLOOMZ https://arxiv.org/abs/2211.01786 https://arxiv.org/abs/2211.01786 to get a sense of the work involved here.
- mycall 4y agoGPUs are much less efficient than cores of the type Cerebras has made.
- vvrm 4y agoThanks for laying out the plan. I was trying to understand the cost of each of these steps below and started wondering about the following: > rough steps: > 1. collect a very large dataset, see: https://www.lesswrong.com/posts/6Fpvch8RR29qLEWNH/chinchilla https://www.lesswrong.com/posts/6Fpvch8RR29qLEWNH/chinchilla... . scrape, de-duplicate, clean, wrangle. this is a lot of work regardless of $. Pile seemed quite clean and manageable to me (I was able to preprocess it ~8 hours for a simple task on consumer grade hardware). Is Pile clean and rich enough for LLM training too ? > 2. get on a call with the sales teams of major cloud providers to procure a few thousands GPUs and enter into too long contracts. It seems like the standard instructGPT model itself is based on a 1 billion param GPT model. Wouldn't that fit on a 24GB RTX 3090 ? Might take longer, maybe not enough opportunity for hyper-parameter search, but still possible right ? Or is hyper-parameter search on a thousand machines in parallel the real magic sauce here ? > 3. "pretrain" a GPT. one common way to do this atm is to create your own exotic fork of MegatronLM+DeepSpeed. go through training hell, learn all about every possible NCCL error message, see the OPT logbook as good reference: https://github.com/facebookresearch/metaseq/blob/main/projec https://github.com/facebookresearch/metaseq/blob/main/projec... Sounds like a good opportunity to learn. No pain, no gain :-) > 4. follow the 3-step recipe of https://openai.com/blog/chatgpt/ https://openai.com/blog/chatgpt/ to finetune the model to be an actual assistant instead of just "document completor", which otherwise happily e.g. responds to questions with more questions. Also e.g. see OPT-IML https://arxiv.org/abs/2212.12017 https://arxiv.org/abs/2212.12017 , or BLOOMZ https://arxiv.org/abs/2211.01786 https://arxiv.org/abs/2211.01786 to get a sense of the work involved here. Maybe somebody would open source the equivalent datasets for this soon ? Otherwise the data collection seems prohibitively expensive for somebody trying to do this for fun: contract expert annotators, train them, annotate/reannotate for months ?
- NicoleJO 4y agoStep number 1 is already the first problem.
- tysam_and 4y agoToday to me is the equivalent of the phone phreaking days when people are just doing as much as they can and getting away with as much as they can for as long as they can, until the regulations come. It will be an interesting time in the next few years as I think StabilityAI's guerilla marketing tactics have inadvertently by proxy also placed the ML dataset debate right in the laps of the larger consumer market.
- anigbrowl 4y agoExtremely interested in your take on where language/reasoning competency ends and knowledge retrieval begins. OpenAI stuff has succeeded in part because it can synthesize good bullshit* on a huge variety of topics. For many purposes this makes it as good as asking someone in the same room to look something up for you on Wikipedia. But while vast general and somewhat special knowledge is very impressive, comprehension and reasoning ability can exist without it. We know from our own human experience that general knowledge is useful to have, but not the same thing as intelligence or wisdom. It seems rational to think that the size of model needed to get ChatGPT's adequate level coherence and rationality is much less than that required to also encode sufficient general knowledge to be informative on just about any topic, most of which are not language specific. * in the Frankfurtian sense of 'information provided without regard to its correctness'
- somenameforme 4y agoThis is why it has always seemed to me that the 'chat bot' -> AI pathway has felt quite analogous to the 'chess bot' -> AI pathway. We're constantly trying to replicate things that look like demonstrations of intelligence, but never really bothering with what intelligence is. What I mean is that a man lifting 400kg is a demonstration of exceptional athleticism. A 400kg man sitting on a balance and having 400kg go up on the other side is not, even though if we only observe the output (400kg goes up) then it is absolutely identical. This isn't just a 'only humans can be intelligent' type argument, but emphasizing that what we want and what we're pursuing seem to be quite different. Newton deriving the inverse square law of gravitational attraction by observing things fall on Earth and watching the celestial bodies in the sky - that is an application of the sort of intelligence that we want. Asking a student to memorize and later recite that the gravitational force is proportional to m1*m2/r^2 is the sort of intelligence that we're building. And it's not like the latter leads to the former, of course it's the exact opposite!
- tysam_and 4y agohence the artificial in the name
- 4y ago
- ingenieroariel 4y agoOn number 2, even if you are John Carmack you may have trouble getting the right people on the phone. https://twitter.com/id_aa_carmack/status/1305967411749892098?lang=en https://twitter.com/id_aa_carmack/status/1305967411749892098... Anyone at Google Cloud out there? It seems I can't get my GPU quota raised to 40 x V100 as an independent researcher. I was told that setting up a website would help, but I would rather not. I can pay the bills...
- ultrons 4y ago@ingenieroariel, sorry for this trouble. I am product manager for Cloud TPU, I would be happy to connect you with my GPU colleagues and also explore if Cloud TPU can help with your research as well. What's the best way to connect with you?
- eyegor 4y agoIf you're actually an independent researcher, sometimes you can find professors at universities or national labs that are willing to help out in exchange for credits on the paper. I've had success at [redacted] labs in the New Mexico region as well as folks from my previous university. The trick is asking people who do research that's sort of adjacent to your field.
- mattnewton 4y agoI gave up and moved to lambdalabs, then they ran out of quota across the board, and now I use a combination of Vast, and Coreweave today.
- Centigonal 4y agob-but you work for- why wouldn't they... you know what, never mind.
- tysam_and 4y agoIt is sometimes far more convenient to avoid the paper trail and bureaucracy of having to provision things internally, similarly I'm sure to how renting GPUs online for short periods of time avoids the issue of having to pay for the maintenance and time costs of maintaining them onsite.
- tysam_and 4y agoI was so confused by the saltiness until I saw the username. I'm sure you've earned it. I got into deep learning because of your char-rnn posts a while ago -- it inspired me to do an undergrad thesis on the topic. I read arxiv papers after that and implemented things from the ground up until a startup liked my work and hired me in a neural network engineer position. Fast forward a few years and I was enamoured with minGPT and it stuck with me. I wanted a CIFAR10 experimentation toolbench so I took my hand at my best swing at applying the minGPT treatment on the current best single-GPU Dawnbench entry, added a few tweaks and got https://github.com/tysam-code/hlb-CIFAR10 https://github.com/tysam-code/hlb-CIFAR10. It currently (AFAIK) holds the world record for training to the 94% mark by a fair bit. It's about 600 lines in a monolithic file, only requiring torch and torchvision, but it's my first project like this and I'd like to learn how to better minify codebases like this. It seems like the hardest part is knowing how to structure inheritance and abstraction, but I don't know if you had any good outside references/resources that you used or would recommend. If you have any feedback or help, I am open to receiving it, as I am very much a newbie at this particular art/science. It is quite a fun one, however (especially as it is a useful tool for my day-to-day work). I'm also hoping to apply the same treatment to a small language model at some point by taking the Dawnbench approach -- picking a good target validation loss value or some reasonable metric, then optimize around that obsessively to build a good tiny reference model. I don't know if you'd know anyone that's interested in that kind of thing, but I feel like that would be a fun next step for me.
- sethammons 4y ago> I was so confused by the saltiness I was confused by what you thought was salty. I don't see it remotely.
- raffraffraff 4y agoSame. I think it just lacks wide-eyed optimism.
- rockwotj 4y agoYou can skip to step 4 using something like GPT-J as far as I understand: https://github.com/kingoflolz/mesh-transformer-jax#links https://github.com/kingoflolz/mesh-transformer-jax#links The pretrained model is already available.
- tysam_and 4y agoGPT-J I think hasn't gone beyond 20B parameters, and while it is not the most obvious I think the original question is asking about the full 180B parameter+ kind of model. :) :thumbsup:
- srajabi 4y agoIt's cool you're on here and I'm sure I speak for many people in saying I really appreciate your video series and comments like the above! Thank you!
- eddsh1994 4y agoI would love to know what your thoughts are on how software engineering (and jobs in general) will change over the next 10 years and what we lowly developers can do to keep up & maybe even be involved in that change
- mrisse 4y agoAndrej wrote an interesting post some years back titled Software 2.0 about the direction he saw software engineering going. It's more about changes in software than the changes in the job market, but I suspect you'd still find it interesting. https://karpathy.medium.com/software-2-0-a64152b37c35 https://karpathy.medium.com/software-2-0-a64152b37c35
- eddsh1994 4y agoThanks!
- osigurdson 4y ago>> de-duplicate, clean, wrangle. this is a lot of work regardless of $. This sounds like a great job for a specialized GPT!
- superflit 4y ago[dead]
- zachur 4y agoYeah, he's explaining how you would create the base model, which is actually one of the more straightforward parts given that they've published their architecture (though I'm sure they've withheld a bit of their special sauce). In reality, putting aside the millions of $$ needed to pay for the GPUs to train the model, the complexity actually lies in the training data acquisition/cleaning and the infrastructure needed to harness the 1000s of GPUs to train it in a remotely reasonable timeframe. That being said there are a number of companies (Google, AI21, Cohere, and probably others) who have successfully created large language models like GPT3, so it's definitely not impossible when you have the resources.
- sva_ 4y agoHe's building the model from scratch, as the title suggests. He only trains a small model with 10M parameters on it, something that is feasible with a single GPU. In comparison, GPT-3 has 175B parameters. > wondering like how hard is it to actually replicate what openAI has done if you had the money to pay for the training? It would most certainly be possible for another company to build something very similar (models of similar size have even be released publicly). I'm honestly unsure why Microsoft would rather pay $10B to acquire less than half of OpenAI, as they have the hardware to do it (OpenAI uses MS cloud products.) Must be some business reasons I don't understand. OpenAI definitely has some very talented people working for it, though.
- poulpy123 4y agoBy single GPU, is a normal one would suffice?
- pcthrowaway 4y agoDoes the time to train the model increase linearly with the number of parameters, or exponentially? In other words, GPT-3 is 17,500X the number of parameters but does that mean you can train it in 17,500X the amount of time it takes to train the 10M param model?
- tysam_and 4y agoI am not from the LLM world, but I believe it's mostly constrained by the standard multiprocessing limits -- communication and synchronization of multiple workers, some of whom operate over an exceedingly slow Ethernet interface.
- bitL 4y agoIn theory it should be linear, however, the parallelization is not perfect and some overlapping parts of gradients are computed on multiple GPUs at the same time so expect some constant factor slowdown on average.
- sebzim4500 4y agoOn top of what other people have said about parallelism overheads, you normally need more data to train a bigger network and the training time is roughly proportional to network size * training data. IIRC OpenAI used a million times more data to train GPT3 than karpathy used in this video, so a naive estimate would be that it would take about 20 billion times more compute. This is could be a significant overestimate since Karpathy probably used each bit of the training set more times than openAI did.
- hluska 4y agoI use this model to help me reason through it. Realistically, it takes five steps. At a glance, the five steps are simple. But when you dig in, you realize that people have 15+ years of experience in each of the individual steps. At a glance, their insight into the individual problems will seem way too big to actually use, but as you dig deeper, you’ll find edge case after edge case that uses those insights. So, it’s really five steps, but very smart people have devoted their lives to figuring out each one of those. We’re lucky to live in a time where we can stand on those shoulders, but it can take quite the leap to get up there in the first place. And of course, at each point you’ll get some nice rewards. It feels good. It’s exciting. And my inner twelve year old feels like this is why we fell in love with tech in the first place. So, it is quite hard but the fact it exists means it’s obtainable. If you get into it, I have a lot of respect for you and hope you have a ridiculous amount of fun. On the other hand, if you get into it and just don’t enjoy it, who cares? There are many other interesting fields!!
- gmadsen 4y agoAs someone who hasn't really done much deep learning, I've always wondered if the work itself is fullfilling or if it is just the fact that there is absurbly cool outcomes? The math isn't super complex, it seems the majority of the effort is data cleaning and tuning. Is it just a massive labor of love? I also worry that the labor itself doesn't build on itself and becomes obsolete knowledge like a web framework.
- c7b 4y agoYeah, that would be really interesting, open source models with similar quality to GPT that even smaller players could use to train models tailored for their application. Kind of the language equivalent to what the OS community already achieved for image models (eg replicating Dreambooth for StableDiffusion). I think one problem standing in the way of that would be that the computational requirements for language models just seem to be a lot higher.