3 ms·
This reminds me of bias tuning, a LoRA competitor. One can get decent adapters by only finetuning a vector added to each linear layer activations. I think I saw
by benob 3y ago
This reminds me of bias tuning, a LoRA competitor. One can get decent adapters by only finetuning a vector added to each linear layer activations. I think I saw it first while reading [1] but there are other instances.
[1] https://arxiv.org/pdf/2304.15010.pdf https://arxiv.org/pdf/2304.15010.pdf
- elcomet 3y agoPlease try to share abstract links instead of pdf links, for mobile or low connection readers.
- aspenmayer 3y agoA fine suggestion. For you and others: https://arxiv.org/abs/2304.15010 https://arxiv.org/abs/2304.15010 also available at: https://doi.org/10.48550/arXiv.2304.15010 https://doi.org/10.48550/arXiv.2304.15010