6 ms·
For those wondering why this is interesting: This technique is being used to reproduce[0] the Alpaca results from Stanford[1] with a few hours of training on co
by numlocked 4y ago
For those wondering why this is interesting: This technique is being used to reproduce[0] the Alpaca results from Stanford[1] with a few hours of training on consumer-grade hardware.
I believe there will soon be a cottage industry of providing application-specific fine-tuned models like this, that can run in e.g. AWS very inexpensively. The barrier today seems to be that the base model (here, Meta's LLaMA) is encumbered and can't be used commercially. Someone will soon, I'm confident, release e.g. an MIT-licensed equivalent and we'll all be off to the races.
[0] https://github.com/tloen/alpaca-lora https://github.com/tloen/alpaca-lora
[1] https://crfm.stanford.edu/2023/03/13/alpaca.html https://crfm.stanford.edu/2023/03/13/alpaca.html
- outside1234 4y agoOr, more importantly than in AWS, locally in disconnected or poorly connected scenarios like in-vehicle or in-home.
- romanzubenko 4y agoToday Databricks announced [0] 6b parameter model from EleutherAI finetuned on Alpaca dataset. According to their CEO[1], training took 3 hours, and costed $30. They didn't release any details on how it was trained, but likely with LoRa. [0] https://www.databricks.com/blog/2023/03/24/hello-dolly-democratizing-magic-chatgpt-open-models.html https://www.databricks.com/blog/2023/03/24/hello-dolly-democ... [1] https://twitter.com/alighodsi/status/1639251347777388544 https://twitter.com/alighodsi/status/1639251347777388544
- numlocked 4y agoInteresting. I wonder what the training cost was for: https://huggingface.co/EleutherAI/gpt-neox-20b https://huggingface.co/EleutherAI/gpt-neox-20b Perhaps it’s in the paper…
- michaelhartm 4y agoThey used the 6b GPT4-J, not 20B. That's what's interesting, it's a smallish large language model :).
- dragonwriter 4y agoGPT-J, not GPT4-J.
- m3affan 4y agoLet the revolutionbbegin
- int_19h 4y agoThere are also some LLaMA LoRAs that are trained on the Anthropic dataset specifically for chat: https://huggingface.co/serpdotai https://huggingface.co/serpdotai I haven't done any formal tests on this yet, but with llama-13b, the overall structure of its responses definitely becomes much more ChatGPT-like. It would be very interesting to see how the 65B model performs.
- polyterative 4y agoThanks! Hard to follow this stuff sometimes with all the news
- GaggiX 4y agoIn addition, for the past 1/2 month this technique has been used to fine-tune Stable Diffusion models.
- terafo 4y agoCloser to 4 months. It is much better than having a bunch of 2-4gb models laying around.
- GaggiX 4y ago4 months? I don't think so, people really start using LoRA when it was added to the diffusers library less than 2 months ago, this library is used by the training plugin of automatic webui, I guess time seems to flow more slowly when many things happen.
- dragonwriter 4y agoThe 1/2 month seems to match Lycoris/LoCon, which as I understand (haven’t dug into the details on this) is a newer refinement of LoRa. LoRa has been used for longer, correct.
- GaggiX 4y agoThe LyCORIS/LoCon repo started committing 1 month ago and almost no one is using it except for a few experiments (not even the automatic webui supports it without a plugin).
- dragonwriter 4y agoJudging from activity on Civitai, I think “almost no one is using it except for a few experiments” is very wrong. Sure, A1111 needs a plugin for it; it needs a plugin for ControlNET, too, but that is also quite popular.
- Agentlien 4y agoControlNet is built in as of maybe two weeks ago and no longer requires an extension. I started using it when the built-in support arrived and have had a lot of fun with it since.
- smaddox 4y agoThere's already RWKV, if you want a decent performing pre-trained model that's Apache 2.0 licensed: https://twitter.com/BlinkDL_AI/status/1638555109373378560?s=20 https://twitter.com/BlinkDL_AI/status/1638555109373378560?s=...
- pffft8888 4y agohttps://news.ycombinator.com/item?id=35281026 https://news.ycombinator.com/item?id=35281026
- arugulum 4y ago> This technique is being used to reproduce[0] the Alpaca results from Stanford[1] Reproduced is a strong statement, without any rigorous justification other than a few cherry-picked examples. Alpaca-LoRA is simply LLaMA with LoRA-tuning on the Alpaca data. There are no metrics, no measurements, no evaluations to show that the Alpaca-LoRA performs similarly to Alpaca, when it is well-known in the field that parameter-efficient fine-tuning always pays a cost in terms of performance relative to full fine-tuning (which is what Alpaca does). (This has been a huge nit for me because of the recent flood of Alpaca-replications, or even claims that Alpaca comparable to ChatGPT, rushing to market themselves, but with nothing to justify their claims.)
- numlocked 4y agoI agree - my comment originally had a parenthetical about this fact, but I thought it was probably confusing to people who just wanted to understand what this was about. Perhaps I shouldn't have edited it out. It also bothers me that a lot of LoRA claims read like "You won't believe how little it costs to train these models!", when of course 99%+ of the complexity and cost is in the LLaMA (or whatever) model that underpins it. Folks are talking about it in a loose way that implies some kind of miraculous overall training cost breakthrough.
- GaggiX 4y ago>when it is well-known in the field that parameter-efficient fine-tuning always pays a cost in terms of performance relative to full fine-tuning The LoRA paper clearly states the performance of the method "LoRA performs on-par or better than fine-tuning in model quality on RoBERTa, DeBERTa, GPT-2, and GPT-3, despite having fewer trainable parameters, a higher training throughput, and, unlike adapters, no additional inference latency. ": https://arxiv.org/abs/2106.09685 https://arxiv.org/abs/2106.09685
- arugulum 4y agoI don't want to get into the weeds of the subtleties of evaluation, hyperparameter-tuning and model comparisons, but let's just say that subsequent studies have shown that LoRA (consistent with most parameter-efficient tuning methods) underperform full fine-tuning: https://arxiv.org/abs/2203.06904 https://arxiv.org/abs/2203.06904 As simple way to think about it is this: if LoRA really gives full fine-tuning performance, why would anyone ever fully fine-tune a model?