5 ms·
LLaMA 30B or 60B can be very impressive when correctly prompted. Deploying the 60B version is a challenge though and you might need to apply 4-bit quantization
by juliensalinas 3y ago
LLaMA 30B or 60B can be very impressive when correctly prompted.
Deploying the 60B version is a challenge though and you might need to apply 4-bit quantization with something like https://github.com/PanQiWei/AutoGPTQ https://github.com/PanQiWei/AutoGPTQ or https://github.com/qwopqwop200/GPTQ-for-LLaMa https://github.com/qwopqwop200/GPTQ-for-LLaMa . Then you can improve the inference speed by using https://github.com/turboderp/exllama https://github.com/turboderp/exllama .
If you prefer to use an "instruct" model à la ChatGPT (i.e. that does not need few-shot learning to output good results) you can use something like this: https://huggingface.co/TheBloke/Wizard-Vicuna-30B-Uncensored-fp16 https://huggingface.co/TheBloke/Wizard-Vicuna-30B-Uncensored...
The interesting thing with these Uncensored models is that they don't constantly answer that they cannot help you (which is what ChatGPT and GPT-4 are doing more and more).
- Llamamoe 3y agoHow low can you get the memory and computational power requirements that way?
- Vecr 3y agoYou can run that model (Wizard-30) on a computer with 64 gigabytes of RAM (or smaller, I don't know how tight you can cut it). You obviously want fast RAM and a good CPU, but you don't need a GPU.
- ed_mercer 3y agoAFAIK you can get away with a swapfile, no need for large amounts of RAM.
- dacryn 3y agowont that nearly kill your ssd if you do it for extended periods of time?
- ZiiS 3y agoMost of the ram is for storing the model once it is loaded it is read only so will not harm an SSD.
- Incipient 3y agoIt only reads from memory,not swap directly. If it needs to read something from swap, it'll write out something from memory to swap, then read the swap into memory. Reading 1gb of swap, will essentially write 1gb to the ssd too. (rough numbers) Correct me if I misunderstand swap?
- mihaic 3y agoIf the underlying data hasn't changed, the page isn't written to disk. CPUs do keep track of writes that mark a memory page as "dirty".
- simcop2387 3y agoThat's basically right. I'm not sure if Linux or windows will keep track of the pages it read out of swap to know if they're still there and valid, but there's a better way for this that I think at least ggml supports where it stores a copy of the model unpacked and ready on disk as a cache for doing the work rather than relying on the OS virtual memory to handle it. This should be faster than the OS VMM (though probably not by much) but since it'll know which pieces it needs to leave on disk and where they are it should be much safer as far as writes go since it will know enough to not write multiple times like that.
- lossolo 3y agoYou can also travel on a bike from NY to LA.
- nobody9999 3y ago>You can also travel on a bike from NY to LA. You can. In fact, my brother did so a bunch of years ago. He found it to be a wonderful experience that made his life better. He's also flown on a commercial airplane from NY to LA (as have I, as well as millions of others) and while it got him to Los Angeles, it didn't provide the levels of sensory input, personal interactions and experience that riding his bicycle did. That's not to say everyone should ride bicycles across the US every time they need/want to make such a trip, but doing so at least once can be a more positive experience than sitting next to some strangers for five hours. The satisfaction of doing so, or the experiences in interacting with people and the landscape during such a trip aren't quantifiable, but reducing the value of doing so (if I'm missing your point here, my apologies) to the time required to make such a trip is reductive in the extreme IMHO. Edit: Clarified my prose.
- Llamamoe 3y agoLove this comment so much. Your brother sounds cool :)
- deleted 3y ago[deleted]
- logicchains 3y agoJust imagine if Boeing lobbied the government to pass a law banning you biking from NY to LA.
- snowycat 3y agoI am running 30b llama models (4 bit quantized using llama.cpp) on 32 gb of ram and no GPU. I get around 2 tokens/second.
- koheripbal 3y ago4-bit quantization removes a lot of the model's sophistication, and 60B parameters is still smaller than what GPT4 is using.
- johnnyworker 3y agoThe point is that it's infinitely better in not being there "just to take your jobs and make a few VCs richer". Nobody even claimed it's more performant. It's like the difference getting nothing, but keeping your land, and getting glass pearls, losing your land. You have to completely ignore the meat of the argument to even pretend there is a contest. And this is without considering what happened if we stopped feeding hostile actors and supported ourselves, instead of keeping to do the reverse. Not just here and there, but consistently for decades.
- motoxpro 3y agoThere is no way I am going to spin up my own worse LLM so a few people will make less money. Even if it was 1-5% better. It's just not worth the time.
- johnnyworker 3y agoIt's not "a few people making less money", it's a few gigantic monoliths carving up the future, like blind watch-destroying gods -- or at least wanting to, no matter how nicely they dress it up. And it's not about utility or chance of success for everyone, either, but rather trying to do something in an ethical or more clean way just because that's more fun for them. But I have to admit to being an idealist, and while I disagree with you because of that, I don't think you should be downvoted for basically just bringing up the majority position. It's easy to complain over people not being starry-eyed idealists that make great personal sacrifices to bring along an utopia for people in 10 generations, or whatever. It's way harder to find and teach the joy of doing something for the sake of doing it, and at the same time come up with medium and long-term ideas that are realistic enough to make working towards them fulfilling, but also genuinely beautiful and true. The whole "rather than teaching to build a ship (we can't even agree on!), teach people how to long for the ocean" thing. It's a really hard problem.
- kyledrake 3y agoHow do you correctly prompt it? A lot of people are not familiar with how to do this. I think this would improve how many people are using the non-OpenAI models.
- lolinder 3y agoJust a reminder that LLaMA is not open—in order to use it legally you have to agree to Meta's terms, which currently means research use only. The versions circulating on torrents are essential pirated, and while I don't have an ethical problem with that at all you can't use it safely in a business. The open replacements for LLaMA have yet to reach 30B, let alone 65B.
- immibis 3y agoIf anyone has a copyright claim to an LLM, the creators of the input data have more of a copyright claim than the company that trained it. There's a good chance they are not copyrightable at all. I'd bet there's a lot of people willing to take on that risk. However, they might still fall under trade secret law.
- fallingknife 3y agoWhy would an LLM be any less copyrightable than any other piece of software?
- meithecatte 3y agoThe software is the matrix multiplication and gradient descent. We are talking about the numbers in the matrices. They are the output of a training algorithm, so we can only talk about the copyright on the training algorithm, and on its input data.
- abtinf 3y agoFor the same reason that phone books cannot have copyright.
- KingMob 3y agoThe model weights could be seen as a derived work, for which they didn't get the permission of the original copyright holders. Alternatively, it can be argued that the LLMs are no different than a fanfic writer trying to imitate the style of their favor author. It's not obvious which way it will go, but I can see the point of those arguing that LLM data are ill-gotten gains.
- stOneskull 3y ago> The interesting thing with these Uncensored models is that they don't constantly answer that they cannot help you (which is what ChatGPT and GPT-4 are doing more and more) that's great to hear. that political correctness in gpt is annoying.