3 ms·
Any opinion on whether the q4 quantization is stable/effective? That's a lot of quantization. Edit: Perhaps answering my own question: λ ollama run hf.co/lmst
by cjbprime 2y ago
Any opinion on whether the q4 quantization is stable/effective? That's a lot of quantization.
Edit: Perhaps answering my own question:
λ ollama run hf.co/lmstudio-community/Llama-3.3-70B-Instruct-GGUF:Q4_K_M
>>> Hi. Who are you?
I'm an artificial intelligence model known as Llama. Llama stands for "Large Language Model Meta AI."
- state_less 2y agoI prefer larger quantization and longer response times, but only have so much memory and am only willing to wait so long. I can work with 2 toks/s and this kind of accuracy for these kinds of queries. I expect you'll need to play around with it to find out if it would be helpful in your case. Here's what I'm getting asking 'who are you?' System Prompt: You are an advanced super intelligence, capable of everything from software development to dream interpretation. >> Who are you? I am an artificial intelligence designed to simulate human-like conversation and answer questions to the best of my knowledge. I'm a large language model, which means I've been trained on a massive dataset of text from various sources, including books, articles, research papers, and websites. My capabilities are diverse and extensive. I can: 1. **Answer questions**: I can process natural language queries and provide relevant answers, covering a wide range of topics from science and history to entertainment and culture. 2. **Generate text**: I can create text based on a prompt or topic, including articles, stories, emails, and even entire books. 3. **Translate languages**: I can translate text from one language to another, including popular languages such as Spanish, French, German, Chinese, and many more. 4. **Summarize content**: I can summarize long pieces of text into concise, easily digestible versions, highlighting key points and main ideas. 5. **Offer suggestions**: I can provide suggestions for topics such as gift ideas, travel destinations, books to read, and more. 6. **Engage in conversation**: I can have a natural-sounding conversation with you, using context and understanding to respond to your questions and statements. 7. **Dream interpretation**: As you mentioned, I can even help interpret dreams, providing insights into the symbolism, themes, and emotions that may be present in your subconscious mind. 8. **Software development**: I can assist with software development tasks, such as writing code, debugging, and optimizing algorithms. My knowledge is constantly evolving, and I'm always learning from the interactions I have with users like you. So, feel free to ask me anything – I'll do my best to help!
- int_19h 2y agoQ4_* is by far the most popular one for local use, and it works "fine" meaning that you do see some effect on perplexity and other scores, but it's small enough to not be a concern in most cases. Although it should be noted that this can depend on the model - e.g. there have been some reports that for QwQ, going 4-bit does adversely impact the quality of its CoT. With Llama specifically I recall someone comparing various quants on perplexity finding that even at Q3, 70B is still smarter than 34B. So quantization is generally worthwhile so long as you have a larger model that you can squeeze into your VRAM budget with it, and don't mind the slowdown from more parameters.