3 ms·
Definitely don't quantize to 4-bit at this param count, if that's what you're doing. The redpajama.cpp repo is a bit crazy for making that the default. I'd say
by bestcoder69 3y ago
Definitely don't quantize to 4-bit at this param count, if that's what you're doing. The redpajama.cpp repo is a bit crazy for making that the default.
I'd say at this size don't count on accuracy. Think: fancy auto-complete, simple NLP task-doer, or a base for fine-tuning a task specific model. IME RedPJ-3B was too small to do anything at all other than ramble. The novel contribution of this model here being you can use this at work, and it has a GGML implementation unlike Falcon and MPT (last I checked) - meaning CPU-based inference on the cloud VM you use at work, maybe.
- redox99 3y agoPeople that run 7B 4bit is probably because they don't have the memory for more. And there are tests showing that quantizing a bigger model is always better than running a smaller model full precision (i.e LlaMa 13B 4bit > 7B fp16)
- anon373839 3y ago> a base for fine-tuning a task specific model Yes. The 7B size can perform very well when tuned on a specific menu of tasks: 1. Goat: Fine-tuned LLaMA Outperforms GPT-4 on Arithmetic Tasks https://huggingface.co/papers/2305.14201 https://huggingface.co/papers/2305.14201 2. Gorilla: Large Language Model Connected with Massive APIs https://arxiv.org/abs/2305.15334 https://arxiv.org/abs/2305.15334