8 ms·
LoRA Learns Less and Forgets Less
- hybridtupel 2y agoThis is about Low-rank adaptation. Not to be confused with LoRa the long range proprietary radio communication technique, which hopefully doesn't learn at all.
- martinky24 2y ago"Why the hell is LoRa learning" was indeed my first thought...
- HeatrayEnjoyer 2y agoThis is how the subs were knocked offline in Terminator III
- gregmac 2y agoThis is "Low-Rank Adaptation", "a widely-used parameter-efficient finetuning method for large language models." Not to be confused with LoRa ("long range") [1], an Internet of Things radio technology. [1] https://en.wikipedia.org/wiki/LoRa https://en.wikipedia.org/wiki/LoRa
- chaos_emergent 2y agoIsn’t this fairly obvious after a two second glance at the abstract
- SubiculumCode 2y agoThis is Low-rank adaptation. Not to be confused with Lake of the Ozarks Recreation Area.
- 0cf8612b2e1e 2y agoApparently constructed in 1929. You think those wireless people would have been more careful when they reappropriated the name.
- chaos_emergent 2y agoThe findings are that the best fine-tune performance comes from fine-tuning all weights, followed my MLPs, followed by attention heads, using LoRA. Authors assert that the performance difference is based on the target module of the NN. Isn’t an equally valid argument that MLPs tend to constitute a greater number of weights in transformer networks than attention heads, and the performance difference can be traced to a greater number of weights having freedom to change? I’d be curious to know if randomly choosing a subset of matrices to train, regardless of where they are in the network, would provide analogous performance to LoRA on a specific module with comparable learnable weights.
- chaos_emergent 2y agoas a follow up curiosity, has anyone tried using LoRA on the entire model for pretraining to compare regular training model performance to LoRA?
- cabidaher 2y agoThis paper [1] does atempt that and reports similar performance compared to conventional pre-training. However, they do start off by doing a normal full-rank training and claim that it is needed to 'warm start' the training process. [1] https://arxiv.org/abs/2307.05695 https://arxiv.org/abs/2307.05695
- danielhanchen 2y agoOh yes this paper! The main issue is the scaling of the A and B LoRA matrices. Some papers show scaling the B matrix with larger learning rates (LoRA+) could be beneficial. DoRA for eg learns an auto scaling vector of numbers which tries to alleviate these issues. Galore might be more equivalent to full pretraining with the gradients being low rank.
- sp332 2y agoDo you mean leaving most of the model in its initial, randomised state and only training a LoRA?
- thepasswordis 2y agoI really wish people would be more careful about choosing names for these things. LoRa has been a popular wireless protocol for like 10 years.
- sva_ 2y agoYes, but this is LoRA, clearly not LoRa.
- squarefoot 2y agoPP has a point though. I entered "LoRA" on Google, DuckDuckGo, Startpage and Bing, and all returned results in all first pages were about the communication protocol (1). They could have inferred my interests from previous searches, but I never used Bing in the last year or so, so it seems to me someone didn't care about name clashes. (1) well, except Google which -surprise- returned about mid page an ad of a local quite expensive chandeliers brand called "LORA".
- sva_ 2y agoI usually just add a term like 'ml' or 'nn' after my search to give the machine context and it is sufficient in most cases.
- deleted 2y ago[deleted]
- refulgentis 2y agoWait until you find out its a name too
- johnisgood 2y agoMost of the time you have to add a keyword as to what it is related. We cannot expect everything to have unique names, unless we are perfectly fine with random pronounceable strings as names.
- atherton33 2y ago
- ssl-3 2y agoWhat can we learn about Low Rank Acronyms today?
- chriskanan 2y agoThis study is great and addresses a question I've had about LoRA for a while. In a continual learning paper from last year, I found LoRA was extremely effective for faster fine-tuning and not forgetting the original dataset: https://arxiv.org/abs/2306.01904 https://arxiv.org/abs/2306.01904
- MarcoZavala 2y ago[dead]
- rzzzt 2y agoThis paper has 12 authors, which fascinates me to no end for some unexplainable reason. How does it work? Is it a common occurrence to have this many people working on a submission? Did each of them get at least a paragraph in edgewise?
- repsak 2y agoI raise you the Gemini paper https://arxiv.org/abs/2312.11805 https://arxiv.org/abs/2312.11805
- SubiculumCode 2y agoFor a serious answer, this is how it works in my field A researcher gets a grant with 3-7 co-investigators. This generates a bunch of data and other resources that will support 10 or more papers. Coinvestigators and PIs will ask their postdocs and grad students to write up a paper. PIs and co-Is go on every paper...because it's a paper from their grant. Then the 1 to 4 grad students and post-docs go on the paper, depending on their specific material contributions to the work, be it analysis, conception, or execution, or writing. The numbers can stack up.
- PeterisP 2y agoThe general criteria for authorship require including the people who worked on the experiments and data for the paper, which can be more important contribution than most of the text in that paper. In other experimental fields, there are papers with dozens or even hundreds of authors, because it can take many people to get to a measurement of a single number in the paper.
- iudexgundyr 2y agoI feel like this is a trivial conclusion. Keeping the rank low in the optimization is a common regularization technique.
- Saris 2y agoWhat does LoRa have to do with LLMs? Whoever named this thing screwed up big time.
- yinser 2y agoThis was a poor study, https://x.com/danielhanchen/status/1791900967472140583?s=46&t=gIaMUOLxHiEBBS2XqapYcg https://x.com/danielhanchen/status/1791900967472140583?s=46&...