3 ms·
> which I assume would be retraining ALL the weights from scratch I have been curious about this as well, and I am not in the field, so I ask ChatGPT about how
by EMM_386 4y ago
> which I assume would be retraining ALL the weights from scratch
I have been curious about this as well, and I am not in the field, so I ask ChatGPT about how it can best learn new information without having to fully retrain the model and all the compute involved in that.
It suggest incremental learning, where its parameters are updated or an additional layer is added to the underlying neural network. Adding a layer involves "adding a new set of nodes or neurons to the existing architecture of the model".
This is far less computationally expensive than retraining the full model, because "adding new layers allows the model to incorporate new information without modifying the existing layers, which can preserve the previously learned knowledge".
You can make of that what you will, after all it was generated by ChatGPT itself. But it seems there are techniques to have these models learn new information without having to do another very costly training run.
- beiller 4y agoI tried to touch on that in my comment. I've used training methods such as LoRA: Low-Rank Adaptation of Large Language Models. I feel they work well to train it on a specific subject, but it's at the cost of over-riding other weights in the model, and it tends to distort things that are unrelated because what a ML model considered to be related is likely not at all what we as humans believe to be related. I've tried LoRA in the context of stable diffusion only. You can see images using LoRAs can be tricky, and even harder to combine 2 LoRAs etc. Adding too many LoRA modifiers can straight up give distorted messed up images.