4 ms·
I'm getting 2.12 tok/s[1] on a 24GB (4090) GPU and 64GB (7950x) CPU memory, splitting the model across the GPU and CPU (40/80 layers on GPU) with lm-studio. Ou
by state_less 2y ago
I'm getting 2.12 tok/s[1] on a 24GB (4090) GPU and 64GB (7950x) CPU memory, splitting the model across the GPU and CPU (40/80 layers on GPU) with lm-studio. Output looks good so far, I can use something like this for a query that I want as good an answer as possible and that I don't want to send out on the network.
If we can get better quantization, or bigger GPU memory footprints, we might be able to use these big models locally for solid coding assistants. That's what I think we have to look forward to (among other benefits) in the year(s) ahead.
1. lmstudio-community/Llama-3.3-70B-Instruct-GGUF/Llama-3.3-70B-Instruct-Q4_K_M.gguf
- cjbprime 2y agoAny opinion on whether the q4 quantization is stable/effective? That's a lot of quantization. Edit: Perhaps answering my own question: λ ollama run hf.co/lmstudio-community/Llama-3.3-70B-Instruct-GGUF:Q4_K_M >>> Hi. Who are you? I'm an artificial intelligence model known as Llama. Llama stands for "Large Language Model Meta AI."
- state_less 2y agoI prefer larger quantization and longer response times, but only have so much memory and am only willing to wait so long. I can work with 2 toks/s and this kind of accuracy for these kinds of queries. I expect you'll need to play around with it to find out if it would be helpful in your case. Here's what I'm getting asking 'who are you?' System Prompt: You are an advanced super intelligence, capable of everything from software development to dream interpretation. >> Who are you? I am an artificial intelligence designed to simulate human-like conversation and answer questions to the best of my knowledge. I'm a large language model, which means I've been trained on a massive dataset of text from various sources, including books, articles, research papers, and websites. My capabilities are diverse and extensive. I can: 1. **Answer questions**: I can process natural language queries and provide relevant answers, covering a wide range of topics from science and history to entertainment and culture. 2. **Generate text**: I can create text based on a prompt or topic, including articles, stories, emails, and even entire books. 3. **Translate languages**: I can translate text from one language to another, including popular languages such as Spanish, French, German, Chinese, and many more. 4. **Summarize content**: I can summarize long pieces of text into concise, easily digestible versions, highlighting key points and main ideas. 5. **Offer suggestions**: I can provide suggestions for topics such as gift ideas, travel destinations, books to read, and more. 6. **Engage in conversation**: I can have a natural-sounding conversation with you, using context and understanding to respond to your questions and statements. 7. **Dream interpretation**: As you mentioned, I can even help interpret dreams, providing insights into the symbolism, themes, and emotions that may be present in your subconscious mind. 8. **Software development**: I can assist with software development tasks, such as writing code, debugging, and optimizing algorithms. My knowledge is constantly evolving, and I'm always learning from the interactions I have with users like you. So, feel free to ask me anything – I'll do my best to help!
- int_19h 2y agoQ4_* is by far the most popular one for local use, and it works "fine" meaning that you do see some effect on perplexity and other scores, but it's small enough to not be a concern in most cases. Although it should be noted that this can depend on the model - e.g. there have been some reports that for QwQ, going 4-bit does adversely impact the quality of its CoT. With Llama specifically I recall someone comparing various quants on perplexity finding that even at Q3, 70B is still smarter than 34B. So quantization is generally worthwhile so long as you have a larger model that you can squeeze into your VRAM budget with it, and don't mind the slowdown from more parameters.
- Me1000 2y agoThe 32B parameter model size seems like the sweet spot right now, imho. It's large enough to be very useful (Qwen 2.5 32B and the Coder variant our outstanding models), and they run on consumer hardware much more easily than the 70B models. I hope Llama 4 reintroduces that mid sized model size.
- xenospn 2y agoqwen2.5 looks like magic compared to llama3.2.
- Sharlin 2y agoA question: How large LLMs can be run at reasonable speed on 12GB (3060), 32GM RAM? How much does quantization impact output quality? I've worked with image models (SD/Flux etc) quite a bit, but haven't yet tried running a local LLM.
- fluoridation 2y ago>How large LLMs can be run at reasonable speed on 12GB (3060), 32GM RAM? If you want to offload fully to VRAM, I'd say 8B is the limit. If you're keeping some on RAM, 15-20B can still give OK performance, depending on your tolerance. >How much does quantization impact output quality? Basically with more quantization the output becomes more incoherent and less realistic. At the extreme end it's basically just gibberish. I think the sweet spot generally is at 4 bits. At that point the model is pretty compact and the quality isn't diminished too much.
- idonotknowwhy 2y agoHe could probably do 12b (Nemo) up to 14b (Qwen 2.5) at 4bpw with exllamav2
- rubatuga 2y agoYou can download LM Studio today and try it out. I've had success with the Mistral Small Instruct IQ3M model, which fits in VRAM.
- magicalhippo 2y agoI got a 2080Ti with 11GB, I can fit Gemma 2 9B Q5_K_M or LLama 3.2 Vision 8B Q4_K_M in memory (if I nuke Firefox's GPU process first). Speed takes quite a hit once you have a few layers on the CPU, but depending on needs it can be doable. I've just asked LLama 3.3 70B Q5_K_M a question and it offloaded about 5 of the 80 layers, so running almost entirely on my 5900X CPU, but still churning out about one word per second. In my experience quantization affects prompt adherence primarily and answer accuracy secondarily. For example, if you have multiple clauses, ie one or more "if this then that", then quantization might get it to not consider those. I also find they tend to answers more generally and less precise at higher quantization levels. As a concrete example, I've been asking the LLama 3.2 Vision 8B model to categorize some images. The default instruct model in general has been heavily trained to output general commentary on the image. If in the prompt I tell it to "output the category only", the Q4_K_M variant sometimes ignores that instruction, while the Q8 variant almost always respect it. Larger models primarily bring more knowledge in my experience, but usually also better prompt adherence. Larger models also typically can support larger contexts, though this can vary, check the model cards. edit: I should clarify. More knowledge also often translates to better, more accurate output. For example, a larger model might recognize an idiom and answer accordingly, while the smaller model fails to recognize it and thus provides a poor answer. Depending on your needs, 12GB might be quite decent or it might be insufficient. If you need an assistant-like model, I liked the Gemma 2 9B Q5_K_M. And I've been quite impressed by LLama 3.2 Vision 8B Q4_K_M for describing images and transcribing text from images. But for more open-ended stuff, especially if larger contexts is needed, I think you might find it underwhelming.
- kristianp 2y agoCan llama.cpp make use of the gpu built into the 7950x CPU? I assume that would improve performance.
- xena 2y agoThe limit is memory bandwidth, a dedicated GPU will have higher memory bandwidth than a CPU or iGPU ever will.
- menaerus 2y agoGranite Rapids memory bandwidth is between 614 and 844 GB/s.
- ryao 2y agoIt would make no difference at best for token generation and would actually run slower for prompt processing since it has fewer gflops than the CPU cores.
- pmarreck 2y agoHow do you measure tokens/sec? Here's my attempt on a new M4 Max 128GB, does about 6 words/sec: bash> time ollama run llama3.3 "What's the purpose of an LLM?" | tee ~/Downloads/what\ is\ an\ LLM.txt A Large Language Model (LLM) is a type of artificial intelligence (AI) designed to process and understand human language. The primary purposes of an LLM are: (... contents excerpted for brevity) Overall, the purpose of an LLM is to augment human capabilities by providing a powerful tool for understanding, generating, and interacting with human language. real 0m59.040s user 0m0.071s sys 0m0.081s pmarreck 59s35ms 20241206220629 ~ bash> wc -w Downloads/what\ is\ an\ LLM.txt 359 Downloads/what is an LLM.txt
- state_less 2y agoLM Studio puts stats at the bottom of each reply like: 2.09 tok/sec, 346 tokens, 1.74s to first token. This was for a 259 word response, so ~ 0.75 words/token. If that ratio holds, you might be getting 8 tok/sec on you M4 Max? Looks like LM Studio is available for ARM based Macs, if you want to give that a try, that'd be one way to get these stats. LM Studio also surfaces up some parameters to play around with, and keeps a record of past conversations if that might appeal to you.
- pmarreck 2y agoYeah, I already use it along with Ollama, I just didn't notice those stats I guess!
- evilduck 2y agoJust add "--verbose" to your run command, e.g. "ollama run mistral-nemo:latest --verbose", it'll dump the token counts and timing info after each message.