3 ms·
It’s actually completely different. What you linked is about zero shot learning by adjusting the prompt, vs Lora which is about actually fine tuning the weights
by stu2b50 4y ago
It’s actually completely different. What you linked is about zero shot learning by adjusting the prompt, vs Lora which is about actually fine tuning the weights of the model.
- timmg 4y agoIn that case, you can think of the prompt as being one vector of the model that is being tuned while the rest is frozen. Not exactly the same, to be sure. But fulfills a similar need: more efficient "fine tuning" of a large model.
- stu2b50 4y agoI suppose that is true. You can even train the prompt with gradient descent. But in practice, it ends up being fairly different.
- eternalban 4y agoThey address prompt tuning's issues in the paper: "The other direction, as exemplified by prefix tuning (Li & Liang, 2021), faces a different challenge. We observe that prefix tuning is difficult to optimize and that its performance changes non-monotonically in trainable parameters, confirming similar observations in the original paper. More fundamentally, reserving a part of the sequence length for adaptation necessarily reduces the sequence length available to process a downstream task, which we suspect makes tuning the prompt less performant compared to other methods." https://ar5iv.labs.arxiv.org/html/2106.09685 https://ar5iv.labs.arxiv.org/html/2106.09685 This is key imo: "More fundamentally, reserving a part of the sequence length for adaptation necessarily reduces the sequence length available to process a downstream task".
- arugulum 4y agoLoRA conversely has different downsides. LoRA can be used in two ways: merged or unmerged. Unmerged (which is how it's trained) incurs a non-trivial computation cost. Merged means you are modifying the model weights, which means you are stuck with that one model on that device (though, this usually applies for most implementations for the unmerged versions too). The benefit of prompt and prefix tuning (note: these are two separate methods) is that you can serve different soft-prompts and soft-prefixes efficiently with a single shared set of model weights.
- eternalban 4y agohttps://ar5iv.labs.arxiv.org/html/2106.09685/assets/x1.png https://ar5iv.labs.arxiv.org/html/2106.09685/assets/x1.png > incurs a non-trivial computation cost The hit seems to be in energy/cpu not time since the W0 computation is in parallel with the BAx. (My assumption based on the latency claims in paper.) So an issue in edge deployments (battery life, etc.). > you are stuck with that one model on that device Upfront I have 0 clue on the actual numbers, but from a purely software architecture pov [in unmerged setup], having that W0 forward process once with n distinct BAx paths (for distinct fine tunings!) would address that, no? [p.s. say an application that takes as input A/V+Txt, runs that through an Ensemble LoRA (ELoRA™ /g) which each participant contributing its own BAx finetuing processing, sharing the single pre-trained W0.]
- arugulum 4y ago> My assumption based on the latency claims in paper. The latency claims are based on the merged version, where the modifications are merged into the model weights. Hence there is no latency cost, since the final model has the same shape as the original. > having that W0 forward process once with n distinct BAx paths (for distinct fine tunings!) would address that, no? The tl;dr is that that works, but is more expensive. Not ridiculously more expensive, but certainly more expensive that processing a few additional tokens with prefix/prompt tuning.
- edwardjhu 4y ago> Merged means you are modifying the model weights, which means you are stuck with that one model on that device (though, this usually applies for most implementations for the unmerged versions too). If one is careful with floating point issues, it's straightforward to unmerge the weights. W_0 = W_1 - BA Yes, prompt-based methods don't involve swapping weights.
- arugulum 4y agoRight, it's mathematically easy (again, up to floating point issues) to recover the weights as needed, but in terms of distribution/serving I'm guessing the plan is to have the original weights and carry around the LoRA weights and merge as necessary. (Also, I'm assuming you're the first author of LoRA.)
- arugulum 4y agoBoth LoRA and prompt tuning are parameter-efficient tuning methods. Both of them inject new weights into the model and tune them. Prompt tuning does so by injecting addition prefix tokens in the input to the model. LoRA does so by injecting low-rank matrices that are additive modifications to a set of linear layers in the model. They both do something slightly different, but are very much in the same class of methods.