5 ms·
https://twitter.com/Tim_Dettmers/status/1689375417189412864 https://twitter.com/Tim_Dettmers/status/1689375417189412864
by perplexitywiz 3y ago
https://twitter.com/Tim_Dettmers/status/1689375417189412864 https://twitter.com/Tim_Dettmers/status/1689375417189412864
- DebtDeflation 3y agoI'm not sure to whom he is responding, since no one is claiming LoRA performs as well as traditional fine tuning. If you click through to the original Tweet he shared, it says "when you have a lot of data and limited compute go for LoRA, while with limited data and ample compute go for full finetuning" which I think is absolutely correct and few would disagree. As these models get bigger and bigger though, fewer and fewer people are going to have the "ample compute" required for full fine tuning.
- scv119 3y agoThe tweet is referring to a paper that fine tunes Chinese dataset on english base model. I'm not surprised with LoRA's poor result in this setup.
- yousif_123123 3y agoI'm not sure less data should require full fine-tuning. If I had 5 pages of text, I don't see why I need to train billions of parameters that are already trained pretty well on general internet knowledge, and already know how to chat.. From a practical perspective, unless cost is really immaterial, I think most will end up starting with Lora, especially for 13b or 70b models.. you could do 10 fine-tuning runs for the cost of a few full fine-tunings. But it's still all witchcraft to me to some degree, and I'd probably try full and Lora.