4 ms·
I would have expected fine-tuning to be good at imparting knowledge. Pre-training is often done for a single epoch only and models soak up the knowledge like cr
by WanderPanda 2y ago
I would have expected fine-tuning to be good at imparting knowledge. Pre-training is often done for a single epoch only and models soak up the knowledge like crazy without multiple passes so why would fine-tuning be any different?
- eightysixfour 2y agoBecause the learning rate vs. pre-training is completely different. This is not accurate, but my mental model is that the LLM's initial training establishes the "space of concepts and ideas" while tuning (like RLHF and fine-tuning) changes how it expresses those concepts and ideas. It works well for me in deciding my approach.
- leobg 2y agoThe problem is that it’s almost impossible to teach knowledge to an LLM without teaching a specific form of expression at the same time. When you have a question/answer tuple in your training data, you are also teaching the model that every other way of answering the question is wrong. So while the LLM would probably be capable of generating maybe 100 answers to the question that would be equally useful (just using different phrasing, different choice of words, etc.) come on, you are forcing it to update its parameters to suppress all of these, except the one specific form that you selected. So you’re not really adding knowledge. Instead, you’re chiseling knowledge away.
- eightysixfour 2y agoI'm not sure this is entirely true, but I guess we could test it by generating a large enough dataset around a specific concept, and see if we can add that concept to a model. Or change one that exists to something else entirely. For example, create an "idea", generate thousands of Q&A pairs about that idea from different angles, and generate conversations about that idea, then train the model on it. This is essentially the Phi process, but with a single concept instead of "everything." My guess is that fine-tuning cannot add the concept to the model without suffering from catastrophic loss everywhere else. However if we added that same data to the pre-training dataset and retrain the model, the model would express the idea correctly.