3 ms·
I am very much a beginner to local LLM stuff and I find it incredibly hard to figure out how to run models optimally with the correct settings for my hardware.
by DanielHB 2mo ago
I am very much a beginner to local LLM stuff and I find it incredibly hard to figure out how to run models optimally with the correct settings for my hardware. The number of different variations of the same model and how each quant work is super confusing as well.
When I tried to run llama.cpp directly I was getting max 9tk/s on qwen3.5-9B, then I tried LM Studio with the same model and got 77tk/s. I haven't figured out yet how to get MTP working properly in either.
- car 2mo agoIf you are on Mac, have a look at the Llama-macOS app. They claim sensible settings for the linked model downloads. I'd expect the authors of Llama.cpp and the Huggingface folks to know this stuff. https://news.ycombinator.com/item?id=49328008 https://news.ycombinator.com/item?id=49328008
- tiahura 2mo agoIf your package manager / configurator isn’t claude code or codex, you’re wasting time.
- dofm 2mo agoUnless, of course, your goal is to actually understand what is going on, regardless of whether that is difficult.
- blactuary 2mo agoSome of us prefer to avoid Anthropic/OpenAI
- throw1234567891 2mo agoYour funny.
- roosterIllusi0n 2mo agoI have had luck telling the free chatgpt my graphics card brand/vram and asking it to recommend latest qwen3.8 or gemma4 model variants. Then I pasted in my server command and ask it to optimize it. I also pasted in the token per second logs to get further tweaks. If you paste the token per second info log info back to chatgpt, you can iterate with the free chatgpt to get better settings.