5 ms·
I'd like to see a writeup of LORAs for something that is readily tractable without LORAs. E.g., a pretrained ResNet34 ImageNet model that gets a LORA instead of
by carbocation 3y ago
I'd like to see a writeup of LORAs for something that is readily tractable without LORAs. E.g., a pretrained ResNet34 ImageNet model that gets a LORA instead of being fine-tuned or fully re-trained. The pedagogical value is that it can be compared to the alternatives which are tractable in this setting (and which are not in an LLM setting).
- mufasachan 3y agoDisclaimer: This is just my intuition, I do not have knowledge about LoRA on small models. It's possible that does not work. LoRA (for Low Rank) benefits from the "small changes" introduced during finetuning of a model. The update of the weights has a low rank. If you take a smaller model, it might induce that the rank is not so low, resulting in degradation in metrics by LoRA compression. I would be interested to see if LoRA still has a benefit in this configuration.
- dragonwriter 3y ago> LoRA (for Low Rank) Pedantic, but it is actually for Low Rank Adaptation.
- knorker 3y agoIf you write it "LORA" then people won't know if you mean "LoRA" or "LoRa". Microsoft already created this huge confusion, so I'd recommend you help not make it worse by using the relevant capitalization.
- carbocation 3y agoBoth this article and Microsoft are referring to the same thing (low-rank adaptation, LoRA, which I have lazily miscapitalized). What is the confusion that you think will be addressed by my fixing that?
- wlesieutre 3y agoLoRa is a LOng RAnge radio system https://lora-alliance.org/ https://lora-alliance.org/ LoRA is LOw Rank Adaptation, the LLM tuning technique
- carbocation 3y agoSo you are saying that in the context of an article about low-rank adaptation, where I was talking about applying this to ResNets, there was actually confusion that I was referring to long range radio?
- xd1936 3y agoSample size of one, but I clicked the article thinking that someone was using an LLM to "tune" the range performance of LoRa radio somehow. :(
- wlesieutre 3y agoI'm not the one who originally commented to complain about your capitalization, but sure. Even if not in this specific comment thread, maybe having correct capitalization here means someone else knows how to write the one they mean to talk about in another conversation later, and then it will have been helpful. If everyone went around calling both of them LORA all the time absolutely that's going to sometimes have people talking past each other about different technologies. Even with them capitalized correctly it's still going to, but not as much. Too bad Microsoft didn't call theirs LoRaA and it would've been more obvious.
- nl 3y agoThe original LoRA paper does exactly this using GPT2 and BERT type language models. They are explicit about it being designed for LLMs, but since it seems to work really well for stable diffusion it would be interesting to see metrics in the image domain.