3 ms·
Do you mean leaving most of the model in its initial, randomised state and only training a LoRA?
by sp332 2y ago
Do you mean leaving most of the model in its initial, randomised state and only training a LoRA?
- buildbot 2y agoI’ve tested specifically this (on my personal time) :) It will train but I found the loss is proportional to the number of trainable parameters. So roughly to hit the performance of a standard 70m param model, you need to train ~70m lora params anyway.
- cheald 2y agoIt's worse than that, because lora requires two matrices per layer. At full rank, you have an additional NxN parameters to learn versus full finetuning, where N is min(input_features, output_features). For example, tuning a layer of 128 in x 256 out is 32k params. Learning a full-rank lora for that layer would be two matrices of 128x128 and 128x256 = 48k params.
- buildbot 2y agoYeah, exactly. Though the 48k param lora might be as good as a 48k param layer of higher rank, I haven't looked into that case really.