7 ms·
People have been running large language models locally for a while now. For now the general consensus is that llama is not fundamentally better than local model
by Spiwux 4y ago
People have been running large language models locally for a while now. For now the general consensus is that llama is not fundamentally better than local models with similar resource requirements, and in all the comparisons it falls short of an instruction-tuned model like Chat GPT
- simonw 4y agoMy argument here is that this represents a tipping point. Prior to LLaMA + llama.cpp you could maybe run a large language model locally... if you had the right GPU rig, and if you really knew what you were doing, and were willing to put in a lot of effort to find and figure out how to run a model. My hunch is that the ability to run on a M1/M2 MacBook is going to open this up to a lot more people. (I'm exposing my bias here as a M2 Mac owner.) I think the race is now on to be the first organization to release a good instruction-tuned model that can run on personal hardware.
- stonerri 4y agoAs someone who just got the 7B running on a base MacBook M1/8GB, I strongly agree. The rate of tool development & prompt generation should see the same increase that Stable Diffusion did a few months (weeks?) ago. And given how early the cpp port is, there is likely plenty of performance headroom with more m1/m2-specific optimization.
- Vetch 4y agoIt's not just that it's accessible, it's also significantly higher in quality than previous local runnable causal LMs. I suspect people saying it's not good are prompting it like ChatGPT, not realizing how much trickier a raw model is to prompt. Getting the hyperparameters for good sampling is another stumbling block. The models are very good if you do everything properly.
- reasonabl_human 4y agoInteresting, where can I learn more about prompting, and tuning a raw model?
- simonw 4y agoThere are a few initial tips here in the LLaMAA FAQ: https://github.com/facebookresearch/llama/blob/main/FAQ.md#2-generations-are-bad https://github.com/facebookresearch/llama/blob/main/FAQ.md#2...
- version_five 4y agoBut llama is the most performant model with weights available in the wild. Personally I hope we quickly get to the stage that there's a real open llm like SD is to DALL-E. It sucks to have to bother with Facebook's core model, and give it more attention than it deserves, just because it's out there. If facebook had actually released it as an open model, I would have said that all the credit should go to them. But instead people are doing great open source work on top of their un-free model just because it's available, and in the popular conception they're going to get credit that they shouldn't
- loufe 4y agoI've been following LLaMa closely since release and I'm surprised to see the claim that it's "general consensus" that's it isn't superior. I've seen machine and anecdotal evidence to the contrary. I'm not suggesting you're lying, but I am curious, can you point me to something you're reading?
- zone411 4y agoYeah, there is no such "general consensus." I don't know where the OP got this idea.
- bestcoder69 4y agoWhat instruction tuned LLM is better?
- yunyu 4y agoFLAN-UL2
- MacsHeadroom 4y agoNot per bit or per watt. LLaMA-30B is 16GB and draws 40 watts from a 4090 GPU.
- yunyu 4y agoLLaMa isn't instruction tuned
- nl 4y agoI've run a lot of language models locally. Llama 7B is much better that something like GPT-Neo at text generation.
- swyx 4y agowhat is needed to be done to instruction tune Llama.cpp? like is all that is needed just a few thousand labeled rows of data?