3 ms·
I'm not much interested in vibe coding (for those who aren't aware that LLMs have other uses). The specific model I've been using with Ollama is hf.co/unsloth/Q
by bachmeier 5mo ago
I'm not much interested in vibe coding (for those who aren't aware that LLMs have other uses). The specific model I've been using with Ollama is hf.co/unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF:UD-Q4_K_XL and it's amazing how fast it is on 64 GB of RAM and i5-13400 CPU. No GPU on this computer. Gemma 4 E4B will think for a couple of minutes vs 3-5 seconds for Qwen. It's hard to believe how much you can do with such limited hardware using their models.
- metalliqaz 5mo agoI have a much more powerful PC and I would not call Qwen3-Coder-30B-A3B "fast" on my machine by any stretch of the word. How are you running it?
- pancsta 5mo agoProbably in the chat / no sys prompt mode for docs only. Try to refac a codebase using a CPU only…
- maille 5mo agoWhat are your use cases?
- throwa356262 5mo agoSince you are using unsloth models from HF, why not use Unsloth Studio instead of Ollama? It is supposed to be faster + they will update the new models multiple times during the first month to correct bugs and performance issues https://unsloth.ai/docs/new/studio#quickstart https://unsloth.ai/docs/new/studio#quickstart