12 ms·
Having spent quite a bit of time playing around with llama.cpp, alpaca.cpp, loras, and the many other llama-based weights lately, here is my impression: The bi
by 2bitencryption 4y ago
Having spent quite a bit of time playing around with llama.cpp, alpaca.cpp, loras, and the many other llama-based weights lately, here is my impression:
The biggest deal with this isn't the published lora adapter (which seems limited to llama 7b), but the cleaned training data, which is likely better than the previous data sets used to train the alpaca-inspired loras that have been publicly released so far. [0]
If you're really limited to running "just" llama 7b, this is great for you. But the biggest value will be when people inevitably release lora adapters for the 13b, 30b, and 65b, based on this training data (assuming it really is better than the previously released adapters).
[0] admittedly, this is based off anecdotes and github issues, and not real measurements. but smarter people than I have claimed the currently most popular loras were trained on messy data, and have started an effort to clean that data and retrain. So if the training data in this repo is high quality like the authors claim, it will benefit models of all sizes.
- m3kw9 4y agoHow does training on just 800k pieces of data need 7b parameters?
- sp332 4y agoLlama 7B was trained on a trillion tokens. The Lora is a small fraction of extra neurons that get integrated into the structure, and those are what get trained on the new data. It's like fine-tuning but takes less RAM and compute than retraining the whole model.
- comex 4y agoBecause it’s fine-tuning an existing 7B-parameter language model, not training from scratch.
- KRAKRISMOTT 4y agoThe chinchilla formula demands a 20:1 ratio
- airstrike 4y ago"demands" is great
- KRAKRISMOTT 4y agoHey I didn't invent Chinchilla, blame google for that.
- adt 4y agoDeepMind (apparently there's some friction between those two!)
- alchemist1e9 4y agoI’ll ask a dumb question. On another of the numerous LLM related posts I was asking if any of the self host-able open model can do code summaries at close to the quality of GPT 3.5 turbo. I was basically told nowhere close yet. Can this potentially do that? Ideally I’d like to have it generate descriptions of large amounts of code but would rather not burn tokens and lose privacy via OpenAI api. But I’d gladly keep a high end GPU burning on such task , even if that was actually slightly more expensive. Edit: To clarify I do this partially now on batches of code via openAI api currently. It’s around 1-3 cents for a typical 400-800 line source code file. And I don’t mean feeding the full code base in as a single input.
- 2bitencryption 4y agoHere's my experience, having used llama+lora 7b, 13b, and 30b, on both cpu and gpu: On gpu, processing the input prompt, even for huge prompts, is almost instant. Meaning, even if your prompt is huge, it will start generating new tokens after your prompt very quickly. On a rented A6000 gpu, using llama+lora 30b, you can use huge prompts and it will start giving a new output right away. On cpu (i.e. the project llama.cpu), it takes a very, very long time to process the input prompt, before it begins to generate new tokens. Meaning, if you provide a huge copy/paste of code, it will take a long time to ingest all that input, before it begins outputting new tokens. Once it finally starts outputting new tokens, the rate is surprisingly fast, not much slower than gpu. I wish I knew the reason for this, but I'm not an expert :) I've just seen this in practice.
- alchemist1e9 4y agoThat sounds promising but how about the quality of the output. I’ve been using OpenAI API with chatblade and been giving it code in various languages and it’s quite surprising how well it describes the purpose and code implementation in english. The english description would be useful and relevant for developers trying to quickly familiarize themselves with a code base. For a typical 400-600 line file it looks like it would cost around 1-3 cents per file. However loss of privacy isn’t so great. How do you find output quality of llama + lora for such a task? Code is research code I’d run it on.
- tmountain 4y agoThis sentence defies lay people: The biggest deal with this isn't the published lora adapter (which seems limited to llama 7b), but the cleaned training data, which is likely better than the previous data sets used to train the alpaca-inspired loras that have been publicly released so far.
- refulgentis 4y agoA Lora is a layer on top of a model, the big deal isn’t that this exists (it’s a Lora for the weakest llama), but the fact they shared their dataset. The stronger llamas trained with this data will produce even better Lora’s and better results.
- axlee 4y agoWhat is a lora or llama? Google gives me nothing.
- bestcoder69 4y agollama: gpt-3 alternative that you can download and run on a toaster lora: efficient way of fine-tuning a model like llama, where instead of recreating an entire model, you're keeping the base model and generating a fine-tunings file to apply on top of it. toaster: any machine with like 4GB of RAM available to fit the model
- h11h 4y agoLLaMA is Facebook's LLM (large language model, comparable with GPT). It's publicly available (anyone can download the weights and run it themselves), so it's popular here. LoRA, or Low-Rank Adaptation of Large Language Models, lets people fine tune a LLM (making it perform better for a particular application) using vastly less resources. Paper: https://arxiv.org/pdf/2106.09685.pdf https://arxiv.org/pdf/2106.09685.pdf
- dpiers 4y agoLoRA: https://arxiv.org/pdf/2106.09685.pdf https://arxiv.org/pdf/2106.09685.pdf LLaMA: https://ai.facebook.com/blog/large-language-model-llama-meta-ai/ https://ai.facebook.com/blog/large-language-model-llama-meta... Both have been the subjects of numerous HN posts in the last month.
- thelittleone 4y agoCould envision some dystopian future where we pay for access to AI with varying tiers of training data.
- bioemerl 4y agoThat's not dystopian, that's already happened and is happening
- pmoriarty 4y agoIt is dystopian when the way humans are using AI is causing inequality to deepen.
- teamspirit 4y agoI'm sorry, this medical ai model only has a small training set and runs on limited resources. It may only provide a treatment with a 50% chance of survival. If you'd like, you can apply for a loan for our MedAI 3000 that creates custom drugs to target your child's cancer. We're headed to Elysium and that's dystopian.
- saurik 4y agoTo verify, the reason that this is not dystopian is because you are assuming we aren't already in the dystopia?
- Tepix 4y agoYes, i haven't seen any fine-tuned LLaMA-65B model so far unfortunately. I guess the cost is a bit high. Perhaps with LoRa someone will do it.
- Rzor 4y agoKeep an eye on these guys: https://github.com/ZrrSkywalker/LLaMA-Adapter/issues/2 https://github.com/ZrrSkywalker/LLaMA-Adapter/issues/2
- saurik 4y agomaybe https://huggingface.co/chavinlo/Alpaca-65B/tree/main https://huggingface.co/chavinlo/Alpaca-65B/tree/main via https://github.com/antimatter15/alpaca.cpp/issues/124#issuecomment-1485075763 https://github.com/antimatter15/alpaca.cpp/issues/124#issuec... ? (I have not tried this yet.)
- SilentM68 4y agoMight want to check out these guys as well: https://cocktailpeanut.github.io/dalai/#/ https://cocktailpeanut.github.io/dalai/#/
- Tepix 4y agoThat's for CPU (and they only provide 7B and 13B).
- pmoriarty 4y agoIs it possible to use AI's to clean training data?
- bitshiftfaced 4y agoThe Alpaca folks used GPT to generate training data. Yeah you can also use it to find issues. It's not perfect, though. What's interesting is the idea of training an LLM, using it to improve the training data, train a better LLM with that, and repeat.