3 ms·
I'm pretty excited about LoRA MoEs, but for the sake of conversation I'll point out a reply someone made to me when I commented about them: https://news.ycombin
by Me1000 3y ago
I'm pretty excited about LoRA MoEs, but for the sake of conversation I'll point out a reply someone made to me when I commented about them: https://news.ycombinator.com/item?id=37007795 https://news.ycombinator.com/item?id=37007795
Any LoRA approach is obviously going to be perform a little worse that a fully tuned model, but I guess the jury is still out on whether this approach will actually work well.
Exciting times!
- spmurrayzzz 3y agoYea its definitely a tradeoff. My intuition here is that, much like the resistance you get to catastrophic forgetting when using LoRAs, adapter-based approaches will be useful in scenarios where your "experts" largely need to maintain the base capabilities of the model. So maybe the experts in this case are just style experts, rather than knowledge (this is pure conjecture, we will see as we eval all these approaches).