10 ms·
LoRA from scratch: implementation for LLM finetuning
- ignoramous 3y agoI've been keeping track of the techniques through Maxime Labonne's LLMs 101: https://github.com/mlabonne/llm-course#4-supervised-fine-tuning https://github.com/mlabonne/llm-course#4-supervised-fine-tun...
- pama 3y agoThanks for the resource. It seems useful enough to warrant its own thread here.
- dymk 3y agoNot to be confused with LoRa ("long range"), a radio communication protocol. At first I thought this could be about using LLMs to find optimal protocol parameters, but alas.
- cpfohl 3y agoI had the exact same confusion
- OJFord 3y agoIt's the first thing that comes to my mind too, but this is mentioned in every thread (and there are far more of them for LoRA than LoRa atm), and in this case there's unlikely to be much confusion because it starts by spelling out the acronym: 'LoRA, which stands for Low Rank Adaptation, [...]'.
- rasbt 3y agoHah, yeah that's LoRA as in Low-Rank Adaptation :P
- thelastparadise 3y agoThis caught me off-guard as well. I really wish they could have used abother acronym.
- deleted 3y ago[deleted]
- deleted 3y ago[deleted]
- the__alchemist 3y agoConcur; or at least don't use a mix of lower and upper-case, like the radio. I think there would be less mis-assumptions if they had called it "LORA", "Lora", "lora" etc. "LoRA" is asking for trouble.
- andy99 3y ago"From scratch" seems to be a matter of opinion. "Pure pytorch" maybe, except it uses HF transformers. So it's LoRA on top of common frameworks...
- rasbt 3y agoYeah, the LoRA part is from scratch. The LLM backbone in this example is not, this is to provide a concrete example. But you could apply the exact same LoRA from scratch code to a pure PyTorch model if you wanted to: E.g. class MultilayerPerceptron(nn.Module): def __init__(self, num_features, num_hidden_1, num_hidden_2, num_classes): super().__init__() self.layers = nn.Sequential( nn.Linear(num_features, num_hidden_1), nn.ReLU(), nn.Linear(num_hidden_1, num_hidden_2), nn.ReLU(), nn.Linear(num_hidden_2, num_classes) ) def forward(self, x): x = self.layers(x) return x model = MultilayerPerceptron( num_features=num_features, num_hidden_1=num_hidden_1, num_hidden_2=num_hidden_2, num_classes=num_classes ) model.layers[0] = LinearWithLoRA(model.layers[0], rank=4, alpha=1) model.layers[2] = LinearWithLoRA(model.layers[2], rank=4, alpha=1) model.layers[4] = LinearWithLoRA(model.layers[4], rank=4, alpha=1)
- 2024throwaway 3y agoThis apple pie recipe claims to be from scratch, but they cooked it in an off the shelf oven. So it's from scratch on top of the universe...
- deleted 3y ago[deleted]
- _bifc 3y agoIf anyone is interested in a more 'pure' or 'scratch' implementation, check out https://github.com/michaelnny/QLoRA-LLM https://github.com/michaelnny/QLoRA-LLM. (author here) It also supports 4-bit quantized LoRA, using only PyTorch and bitsandbytes, without any other tools.
- HumblyTossed 3y ago[flagged]
- deleted 3y ago[deleted]
- huqedato 3y agoExcellent and practical example! I'm curious if there's a comparable one using Julia or JavaScript.
- ijhuygft776 3y agoI wish the wireless LoRa protocol would be open source...
- ijhuygft776 3y agohttps://www.epfl.ch/labs/tcl/wp-content/uploads/2020/02/Reverse_Eng_Report.pdf https://www.epfl.ch/labs/tcl/wp-content/uploads/2020/02/Reve...
- broabprobe 3y agowow definitely thought this was about LoRa at first.
- gourabmi 3y agoSomeone somewhere is already working on naming their project Lehsun.. /s
- rsweeney21 3y agoIt's still strange to me to work in a field of computer science where we say things like "we're not exactly sure how these numbers (hyper parameters) affect the result, so just try a bunch of different values and see which one works best."
- manojlds 3y agoDivine benevolence
- r3trohack3r 3y agoI feel like it's the difference between something that has been engineered and something that has been discovered. I feel like most of our industry up until now has been engineered. LLMs were discovered.
- arketyp 3y agoI understand your distinction, I think, but I would say it is more engineering than ever. It's like the early days of the steam engine or firearms development. It's not a hard science, not formal analysis, it's engineering: tinkering, testing, experimenting, iterating.
- peddling-brink 3y ago> tinkering, testing, experimenting, iterating But that describes science. http://imgur.com/1h3K2TT/ http://imgur.com/1h3K2TT/
- amelius 3y agoAI requires a lot of engineering. However, the engineering is not what makes working in AI interesting. It's the plumbing, basically.
- justanotheratom 3y agoand finally, this justifies the "science" in Computer Science.
- 3y ago
- kakayako 3y ago[dead]
- chenxi9649 3y agoIt's still not too clear to me when we should fine tune versus RAG. In the past, I used to believe that finetuning is mostly for model behavioral change, but recently it seems that certain companies are also using fine-tuning for knowledge addition. What are the main use cases for fine tuning?
- rasbt 3y agoI think the main use case remains behavior changes: instruction finetuning, finetuning for classification, etc. Knowledge addition to the weights is best done via pretraining. Or, if you have an external database or documentation that you want to query during the generation, RAG as you mention. PS: All winners of the NeurIPS 2023 LLM Efficiency Challenge (finetuning the "best" LLM in 24h on 1 GPU) used LoRA or QLoRA (quantized LoRA).
- ignoramous 3y agoFrom what I gather, fine-tuning is unreasonably effective [0] because in-context learning really depends on how powerful the underlying model is and just how you do RAG (process queries, retrieve embeddings, rank outcomes, etc [1]). Per this paper I read, fine-tuning may add new domain knowledge (but as another commenter pointed out, knowledge is better represented from data of the pre-training stage) or boost specific knowledge; while RAG is limited to boosting only; nevertheless, both techniques turn out to be similarly capable with different trade-offs [2]. -- [0] Fast.ai: Can Models learn from one sample, https://www.fast.ai/posts/2023-09-04-learning-jumps/ https://www.fast.ai/posts/2023-09-04-learning-jumps/ / https://archive.is/eJMPR https://archive.is/eJMPR [1] LlamaIndex: Advanced RAG, https://blog.llamaindex.ai/a-cheat-sheet-and-some-recipes-for-building-advanced-rag-803a9d94c41b https://blog.llamaindex.ai/a-cheat-sheet-and-some-recipes-fo... / https://archive.is/qtBXX https://archive.is/qtBXX [2] Microsoft: RAG vs Fine-tuning: Pipelines, Tradeoffs, and a Case Study, https://arxiv.org/html/2401.08406v2#S6 https://arxiv.org/html/2401.08406v2#S6 / https://archive.is/UQ8Sa#S6 https://archive.is/UQ8Sa#S6
- CuriouslyC 3y agoFine tuning is better than RAG when the additional data isn't concise, or requires context. This is because too much context (or "unfocused" context) can dilute prompt following behavior, and RAG doesn't help the model with higher order token associations so you have to get lucky and pull what you need from the augmentation material, at which point it's not much better than a fancy search engine. Of course this is mostly an issue when you're dealing with a specialized corpus with its own micro-dialect that isn't well represented in public data sets, such as with government/big corporation internal documents.
- yatugalu 3y ago[dead]
- jamesblonde 3y agoI prefer the not from scratch, but from configuration approach by Axolotl. Aolotl supports fine-tuning mistral, llama-2, with lots of the latest techniques - sample packing, flash attention, xformers. I concentrate on collecting and curating the fine-tuning data, do "data-centric" fine-tuning - not learning LoRA from scratch.
- wfalcon 3y agothis is also what our (Lightning AI) lit-gpt library does. https://github.com/Lightning-AI/lit-gpt https://github.com/Lightning-AI/lit-gpt
- jamesblonde 3y agoThanks, hadn't seen this.
- denysvitali 3y agoLoRA != LoRa. I keep on getting confused and hate that they chose to reuse an existing acronym
- sschueller 3y agoIt's unfortunate that those two so far unrelated technologies have the same acronym.
- daemonologist 3y agoLikewise. My day job is machine learning and I still, or maybe consequently, do a double-take every time I see the acronym with minimal context (like on the HN front page, where either usage would be normal).
- travisgriggs 3y agoAnd my day job involves a lot of LoRa. I always do a double take on these. I'm grateful that at least the caps is now being done differently.
- sbrother 3y agoWait, what is the meaning other than "Low-Rank Adaptation"? It's hard to google the difference.
- boolemancer 3y ago
- facu17y 3y agoWhat's the performance penalty of LoRA?
- rasbt 3y agoDuring training, it's more efficient than full finetuning because you only update a fraction of the parameters via backprop. During inference, it can ... 1) ... be theoretically a tad slower if you add the LoRA values dynamically during the forward pass (however, this is also an advantage if you want to keep a separate small weight set per customer, for example; you run only one large base model and can apply the different LoRA weights per customer on the fly) 2) ... have the exact same performance as the base model if you merge the LoRA weights back with the base model.
- yandrypozo 3y agogotta say naming is hard I thought this was about LoRa (from "long range") or LoRaWAN, the IoT sensors communication.
- somethingsome 3y agoNice article, I'm not in this field, however, my understanding of the original paper was that the LoRA was applied only on the last dense layer, and not to all independently (maybe I misread it originally). Digging a bit in why the implementation is like this in the link, I found that in QLoRA they used this and it seems to have some interesting effects, maybe adding a note on the QLoRA decision would be nice :) I'm not sure I understand why it works though, my neophyte view was that applying LoRA to the last layer made sense, but, I do not wrap my mind on the rationale of applying it repeadly to each linear layer. Can someone explain their intuition?
- icyfox 3y agoLike most things in ML, the answer of which layers to use come down to empirical evidence more than theory. In a typical Lora training pipeline, you freeze the contents of the base model and just adjust the Lora layers. The more layers you convert to lora layers the more degrees of freedom you have for the optimization. There are some finetuning regimens that only recommend finetuning the last layer since this is theorized to have the "highest order" representation of the inputs. Other training regimens will finetune all layers. It's largely data and problem dependent. Lora just mirrors this convention.
- somethingsome 3y agoYeah, but if I remember correctly the paper, LoRA followed the logic that only the last layers on a llm changed drastically during finetuning, and the layers above remained almost unchanged, so it made sense to alterate only the last ones, breaking this by adding a LoRA at each linear layer doesn't seem to follow the logic of why LoRA was created and why it works.
- icyfox 3y agoWell, Lora works just because it's a low rank approximation of full updates - much in the same way that SVD works, and regular gradient updating works. It delivers good results by both acting as a regularizer and by allowing larger models to be updated with smaller memory footprints. My point is that the original Lora paper choosing the last layer is one choice. And it is likely the most common one because of its higher symbolic nature typically being all that's needed for good performance on downstream tasks. Depending on the size of your finetuning job I've personally seen updating more layers (or updating some only on a certain learning rate schedule) to be more effective. Lora is just the mathematical technique of updating, it doesn't really have a hypothesis on the ideal training regimen.
- helloericsf 3y agoHN friends, What are the most popular libraries for fine-tuning? (Not from scratch)
- jasonjmcghee 3y agohttps://github.com/OpenAccess-AI-Collective/axolotl https://github.com/OpenAccess-AI-Collective/axolotl
- mintrain 3y ago[dead]
- Rudeg 3y agonice, looks very cool and useful! I'll definitely try it!
- fnordfnordfnord 3y agoI thought this was going to be some neat software defined radio stuff. Still quite interesting though.
- z3ugma 3y agoit's all about whether the 'A' is capitalized or not. LoRa - radio LoRA - machine learning
- tussa 3y agoIt's cheap and sleazy to steal a name from another project to ride it's fame.