17 ms·
Here's my experience, having used llama+lora 7b, 13b, and 30b, on both cpu and gpu: On gpu, processing the input prompt, even for huge prompts, is almost insta
by 2bitencryption 4y ago
Here's my experience, having used llama+lora 7b, 13b, and 30b, on both cpu and gpu:
On gpu, processing the input prompt, even for huge prompts, is almost instant. Meaning, even if your prompt is huge, it will start generating new tokens after your prompt very quickly. On a rented A6000 gpu, using llama+lora 30b, you can use huge prompts and it will start giving a new output right away.
On cpu (i.e. the project llama.cpu), it takes a very, very long time to process the input prompt, before it begins to generate new tokens. Meaning, if you provide a huge copy/paste of code, it will take a long time to ingest all that input, before it begins outputting new tokens.
Once it finally starts outputting new tokens, the rate is surprisingly fast, not much slower than gpu.
I wish I knew the reason for this, but I'm not an expert :) I've just seen this in practice.
- alchemist1e9 4y agoThat sounds promising but how about the quality of the output. I’ve been using OpenAI API with chatblade and been giving it code in various languages and it’s quite surprising how well it describes the purpose and code implementation in english. The english description would be useful and relevant for developers trying to quickly familiarize themselves with a code base. For a typical 400-600 line file it looks like it would cost around 1-3 cents per file. However loss of privacy isn’t so great. How do you find output quality of llama + lora for such a task? Code is research code I’d run it on.
- ineedasername 4y agoHow long is very very long? Am I going to get coffee while it works, going to lunch, doing it right before I go to bed, or hoping it finishes in time to come up with the most perfect epitaph on my tombstone? ;)
- londons_explore 4y agoThat is likely a bug. There is nothing in the maths that should make prompt tokens slower on CPU.
- loufe 4y agoIt is a bug, it's been clearly discussed under the issues of LLaMA.cpp. The wait time is far from atrocious, with a 5600x and the 13B I would wait 15 seconds before input starts for a relatively complex prompt.
- kolinko 4y agoAny tips on setting up llama+lora on 30b? There are so many resources that I can't figure out which models to use, and which projects to use to set everything up.
- nickthegreek 4y agotext-generation-webui (https://github.com/oobabooga/text-generation-webui https://github.com/oobabooga/text-generation-webui)
- kolinko 4y agoI tried it, but I can't find the right model for llama+lora 30B . Google, nor Bing ;) is not helpful.
- nickthegreek 4y agoI think this is the one you want: https://huggingface.co/elinas/alpaca-30b-lora-int4 https://huggingface.co/elinas/alpaca-30b-lora-int4
- stolsvik 4y agoAs I understand these models, that makes no sense whatsoever. Is it the tokenisation that takes this crazy amount of time? Because each additional token should take exactly the same amount of time as the first.