4 ms·
>In practice this means you can fine tune a 30B parameter model on a consumer GPU in a couple of hours. Consumer GPU, yes, but in practice LoRA doesn't actuall
by arugulum 4y ago
>In practice this means you can fine tune a 30B parameter model on a consumer GPU in a couple of hours.
Consumer GPU, yes, but in practice LoRA doesn't actually reduce training time. What it mainly reduces is memory requirements. In fact LoRA training can often require more training steps than full fine-tuning and therefore be slower (you can imagine why this is the case: the optimization is trying to modify the mode's behavior a smaller number of parameters, and so has a harder job)
- MacsHeadroom 4y agoModern peft methods with LoRA actually do reduce training time by orders of magnitude. Here's an example of 20 seconds per epoch on a single consumer GPU: https://github.com/johnsmith0031/alpaca_lora_4bit/issues/7#issue-1635228539 https://github.com/johnsmith0031/alpaca_lora_4bit/issues/7#i...