3 ms·
This makes sense but I feel there might be an opportunity for people who knows how to train compact versions of these LLMs which run on local machines and solve
by subidit 3y ago
This makes sense but I feel there might be an opportunity for people who knows how to train compact versions of these LLMs which run on local machines and solve a very specific use case for that business.
What could be a bootcamp like curriculum for someone who wants to learn only how to train a LLM (upto a sellable standard)?
- PaulHoule 3y agoI want to say “go back in time and start a PhD program 7 years ago” in that in a PhD program, no matter what the subject, you learn to solve problems on your own. Because it is such a new thing I don’t think there is any curriculum available, what you really can do is “learning through doing” and also reading papers both to get some idea of what people are doing (I like the papers where run of the mill researchers solve run of the mill problems because that gives me an idea of what to expect with my run of the mill problem) and also to pick out results to try to replicate. My best advice right now is that it is important to scale down to something that lets you do a lot of experiments. For instance I have a classifier that takes about 30 seconds to train and I can rapidly iterate on it because I can run it 1000s of times a day. I have another that takes 30 minutes and I am loathe to put effort into it because it so easily becomes a tarpit. Thus start with something that gets meaningful results in a minimal time, get really comfortable with it, then scale it up deliberately until it is large enough. Feel free to send me an email if you want to talk more.