4 ms·
Standard LoRA (W_delta = B@A with standard inits) generally underperforms FT, primarily because of "intruder dimensions" (new high-ranking singular vectors whic
by cheald 1y ago
Standard LoRA (W_delta = B@A with standard inits) generally underperforms FT, primarily because of "intruder dimensions" (new high-ranking singular vectors which misalign with the singular vectors of the underlying weights) as outlined in the paper.
There are techniques like PiCa and SVFT which can mitigate much of the loss, though.
- tangjurine 1y agopica came out two days ago, how did you find out about it?
- cheald 1y agoThe one I was referring to was from this paper, first published in May: https://arxiv.org/abs/2505.20211v1 https://arxiv.org/abs/2505.20211v1 I don't recall how I found out about it, but it was either paperswithcode or an LLM research session working through the intruder dimensions problem. In my Stable Diffusion tests, it substantially improves LoRA training speed and fidelity, though I've got some experiments that seem to even further substantially improve on it by adding learnable rotations of the singular vectors.