4 ms·
I think not much can run without a dedicated GPU
by aphroz 2mo ago
I think not much can run without a dedicated GPU
- noduerme 2mo agoWhat's the story with Mac laptops? Worth a try?
- rawland 2mo agoYes. mtplx runs it at 25 tok/sec on a M4 Max with 48GB RAM.
- selcuka 2mo agoThe author tested in on an M5 laptop too: > It feels pretty slow on both the M5 Mac and the DGX Spark.
- cyberrock 2mo agoDense ones like this are more bandwidth-hungry, so you want to try MoE ones like Qwen3.6-35B-A3B (35 Billion params but only 3 Billion Active) or Gemma 4. Unfortunately it seems like we might not be getting a 3.8 MoE.
- madduci 2mo agoTill now I was using successfully Qwen 3.5 and Gemma 4 at a reasonable speed
- mobelkh 2mo agowere you running the MoE models? those perform better speed wise
- pyrale 2mo agoThere is no way you would run a dense 27b model on that spec. I ran 3.6 27b on a 64gb ram, 24 gb vram, and it felt like the lower limit for this model with a decent context window. If you want a better experience, maybe wait for either a moe model (like 3.6 35b A3) or a model with less parameters (like 9b). Qwen has been releasing those in the past, so maybe we’ll have them for 3.8 too.
- DanielHB 2mo agoFrom my experience if it doesn't fit on vram it is rarely worth to bother except for a few narrow tasks. For example make an essay about something where you don't actively engage with the LLM after the initial prompt. So mostly one-shot prompts.
- madduci 1mo agoThe issue was caused by Ollama. I've tried again with llama.cpp and mounting the igpu device correctly, I get 18-23 tk/s on the Iris Xe card under Debian, MUCH better than before.
- kzrdude 2mo agoI think (maybe I missed something) that identical size and quant versions of Qwen 3.5 and 3.8 should run at the same speed. It’s the exact same architecture.
- madduci 2mo agoTried the 4 Bit versions. It loads, bit the <thinking> output isn't even coming out.