3 ms·
I haven’t ran QWQ yet, but it’s a 32B. So about 20GB RAM with Q4 quant. Closer to 25GB for the 4_K_M one. You can wait for a day or so for the quantized GGUFs t
by syntaxing 2y ago
I haven’t ran QWQ yet, but it’s a 32B. So about 20GB RAM with Q4 quant. Closer to 25GB for the 4_K_M one. You can wait for a day or so for the quantized GGUFs to show up (we should see the Q4 in the next hour or so). I personally use Ollama on an MacBook Pro. It usually takes a day or two for it to show up. Any M series MacBook with 32GB+ of RAM will run this.
- aledalgrande 2y agohttps://ollama.com/library/qwq https://ollama.com/library/qwq
- int_19h 2y agohttps://huggingface.co/lmstudio-community/QwQ-32B-Preview-GGUF https://huggingface.co/lmstudio-community/QwQ-32B-Preview-GG...
- Terretta 2y agoOn Macbooks with Apple Silicon consider MLX models from MLX community: https://huggingface.co/collections/mlx-community/qwq-32b-preview-67485e0db78b18207ba61359 https://huggingface.co/collections/mlx-community/qwq-32b-pre... For a GUI, LM Studio 0.3.x is iterating MLX support: https://lmstudio.ai/beta-releases https://lmstudio.ai/beta-releases When searching in LM Studio, you can narrow search to the mlx-community.
- coconut08 2y agoon macos with lm-studio is it better to use the mlx-community releases over the one that lm-studio releases? also I didn't install a beta and mine says i'm using 3.5 which is what the beta also says. is there a difference right now between the beta and the release version?
- Terretta 2y agoYou're right, looks like 0.3.5 is now on the home page.
- swyx 2y ago> 20GB RAM with Q4 quant. Closer to 25GB for the 4_K_M one how does this math work? are there rules of thumb that you guys know that the rest of us dont?
- mekpro 2y agoAs a quick estimation, the size of q4 quantized model usually be around 60-70% of the model's parameter. You can preciselly check the quantized model size from .gguf files hosted in huggingface.