7 ms·
Takeaways from hundreds of LLM finetuning experiments with LoRA
- naillo 3y agoGood article but my god why is there like a 2 second delay before every user interaction on this page. Scrolling, selecting text, everything has some kind of 2 second 'bootup' time after which things work normally but if you stop interacting with it for a bit it goes back to some idle mode. Really weird.
- mike-cardwell 3y agoSuggest upgrading your browsing experience by switching to Firefox with uBlock Origin.
- Tiberium 3y agoMaybe this is why :) https://i.imgur.com/odYgMpd.png https://i.imgur.com/odYgMpd.png
- swatcoder 3y agoAs others noted, the latency is from a whole slew of trackers that privacy extension would block. Maybe this is a good nudge for you to go install one. It's a problem that's already widespread and will continually get worse. But on top of that, this is also one of the many fun examples of where a static, linear document that's so simple it could have been written in markdown requires megabytes of download across dozens of files. Nobody cares about user experience right now -- it's all about developer and content management experience, so runtime bulk and responsiveness are ignored in favor of... whatever these kinds of developers think they're gaining on the production end. I often have a hard time guessing what that is without assuming the responsible devs are not-so-great.
- carbocation 3y agoI'd like to see a writeup of LORAs for something that is readily tractable without LORAs. E.g., a pretrained ResNet34 ImageNet model that gets a LORA instead of being fine-tuned or fully re-trained. The pedagogical value is that it can be compared to the alternatives which are tractable in this setting (and which are not in an LLM setting).
- mufasachan 3y agoDisclaimer: This is just my intuition, I do not have knowledge about LoRA on small models. It's possible that does not work. LoRA (for Low Rank) benefits from the "small changes" introduced during finetuning of a model. The update of the weights has a low rank. If you take a smaller model, it might induce that the rank is not so low, resulting in degradation in metrics by LoRA compression. I would be interested to see if LoRA still has a benefit in this configuration.
- dragonwriter 3y ago> LoRA (for Low Rank) Pedantic, but it is actually for Low Rank Adaptation.
- knorker 3y agoIf you write it "LORA" then people won't know if you mean "LoRA" or "LoRa". Microsoft already created this huge confusion, so I'd recommend you help not make it worse by using the relevant capitalization.
- carbocation 3y agoBoth this article and Microsoft are referring to the same thing (low-rank adaptation, LoRA, which I have lazily miscapitalized). What is the confusion that you think will be addressed by my fixing that?
- wlesieutre 3y agoLoRa is a LOng RAnge radio system https://lora-alliance.org/ https://lora-alliance.org/ LoRA is LOw Rank Adaptation, the LLM tuning technique
- carbocation 3y agoSo you are saying that in the context of an article about low-rank adaptation, where I was talking about applying this to ResNets, there was actually confusion that I was referring to long range radio?
- munro 3y agoThis is amazing, thank you!! > My hypothesis is that the Alpaca dataset does not contain any related arithmetic tasks, and the model actively unlearns basic arithmetic when it focuses more on other tasks. I'm surprised this wasn't verified, it's a major benchmark stat. My eyes keep getting drawn to it, because it seems to have the most variance. Does anyone know? Also throwing it out, I would love to see a Neptune/Wnb of the hyperparameter tuning :)
- pbhjpbhj 3y ago>the model actively unlearns Aka "catastrophic forgetting" (CF).
- moffkalast 3y agoThere's also been a recent discovery that the MMLU benchmark contains a significant percentage of answers with completely wrong ground truth, where models answering correctly lots points. The most popular benchmarks and datasets are mostly haphazardly cobbled together with hardly any oversight or verification, sometimes even synthetically generated from gpt 3.5 et al. without checking the output at all. Frankly it's amazing that any of it even works when people blindly train and test with what's essentially self contradicting garbage.
- munro 3y agoOh wow, yea I see. A web search brings up a lot of examples. >> As a result of an accident, Abdul lost sight in his right eye. To judge the distance of vehicles when he is driving, Abdul is able to rely on cues of >> - A. I only >> - B. II only >> - C. III only >> - D. I and II only > You didn’t read that wrong. The question never explains what I, II or III are. This appears to have been improperly copied from crackap.com. Somehow Platypus 2 still gets the right answer with high confidence. Is this a sign it has merely memorized the answers? I checked the second best ranked model upstage/LLama-2–70b-instruct-v2 and it also somehow got the answer right (the third best Open LLM also gets this question right so I don’t know what is happening). https://derenrich.medium.com/errors-in-the-mmlu-the-deep-learning-benchmark-is-wrong-surprisingly-often-7258bb045859 https://derenrich.medium.com/errors-in-the-mmlu-the-deep-lea...
- lappa 3y agoGreat work, lots of useful information here. The only thing I wish you did different was explored alpha > 2 * r. In this blog post, the author found that alpha of 4 * r (where r=64) outperformed all smaller alphas in terms of loss when finetuning Llama-7b on databricks-dolly-15k. https://medium.com/@drishtisharma96505/comparative-analysis-of-lora-parameters-on-llama-2-with-flash-attention-574b913295d4 https://medium.com/@drishtisharma96505/comparative-analysis-... Additionally you identify (alpha = 2*r) r=16 as inferior to r=256, however aside from arithmetic, r=16 actually outperforms all others. And the base model outperforms any finetuning for both arithmetic metrics.
- Der_Einzige 3y agoLoRAs are, like most fine-tuning, a spectrum. LoRAs can be nearly the same size as the original model, with nearly the same representation capacity/trainability, or they can be a tiny tiny fraction of the original model size, with correspondingly fewer learnable parameters. As such, they are suitable for nearly all tasks. We should be asking if they are better than regular fine-tuning, or soft-prompts (aka textual inversion), or slapping new trainable layers on the end (aka hypernetworks). The stable diffusion community seems to think that they are.
- stan_kirdey 3y agoclaude2 tldr: Here are a few key points I gathered from the article: The article explores optimizing LoRA hyperparameter settings for finetuning large language models. The goal is to maximize performance while minimizing memory usage and training time. The base model used is Llama-2 7B. Experiments compare default LoRA, QLoRA (4-bit quantized), AdamW, SGD optimizers, and different choices for rank r and alpha hyperparameters. Key findings: QLoRA provides substantial memory savings (6GB less than default LoRA) at the cost of slower training. Performance impact is minor. AdamW vs SGD makes little difference in memory or performance. Increasing training iterations from 50k to 100k hurts performance, likely because the Alpaca dataset lacks diversity. Tuning rank r and alpha is most impactful. Good rule of thumb is to set alpha=2*r. Best model uses r=256, alpha=512. Improves over base model on most tasks, except arithmetic. The optimized LoRA model was submitted to the NeurIPS efficiency challenge and showed improvements on several benchmarks compared to the base Llama-2 model. Takeaways are practical tips for tuning LoRA hyperparameters and trading off memory, compute, and model performance.
- snitty 3y agoLoRA, not LoRa. I was VERY confused for a minute.
- wkipling 3y agoI was excited to see long range radio to meet machine learning
- QuantumG 3y agoSeemed to take terrible performing models and marginally improve them. None of them are fit for purpose at the end of the training, so what was the point?
- nl 3y agoScience.
- nl 3y agoI'm outside my edit window but I'm surprised at the negative reaction to this. This experimental process on small NNs us exactly how science is done. You develop and test theories of what will work using small networks because you can do them quickly. Then you test on large, production size networks based on the lessons from the small ones.
- grandmczeb 3y agoProbably because one word replies come across as low effort. IMO your second reply is great.
- cypress66 3y agoIt had the right amount of effort that the parent question deserved.
- IOT_Apprentice 3y agoHow would anyone know if they did not posit an idea or theory and then implement it as a test? Publish results and it may spur others to tweak it or try other approaches.