3 ms·
Is there a resource somewhere that tells me, "Given that you have X Gb RAM on your laptop, these are the current local models you should consider trying and her
by Myrmornis 1mo ago
Is there a resource somewhere that tells me, "Given that you have X Gb RAM on your laptop, these are the current local models you should consider trying and here's how to configure them to leave enough memory for other applications on your machine."?
- blueferret 1mo agoSeveral exist actually. Try whichllm.app or fitmyllm.com.
- rcktmrtn 1mo ago> https://www.whichllm.app/ https://www.whichllm.app/ - Linux, general use case, balance - 16 GB RAM - 10 GB VRAM Recommendation: Kimi-K3 This checks out.
- sandos 1mo agoIt says Kimi-K3 for very set of parameters I put in! 64 or 128GB RAM, 6 or 64GB of VRAM.... Not sus at all.
- tencentshill 1mo agoRoleplay, windows, 16gb/4gb Kimi-K3 I think the slop site is broken or compromised...
- rurban 1mo agowhichllm sucks. It only recommends gguf models, and llama.cpp is not supported. fitmyllm is much better
- crossroadsguy 1mo agoIt gives astronomically optimistic results for my 16GB M1 Pro :) By the way own exploration sort of led me to qwen3.5:9b for the best case scenario balanced model considering almost 10-11GB of RAM is almost always gone anyway. Even with aggressive app quitting.
- chrishare 1mo agoNot that I know of, but https://www.canirun.ai/ https://www.canirun.ai/ might be of use
- bhelkey 1mo agoThis calls the top coding model for the Apple M1 Pro: qwen2.5-coder-7b, a model released September 18, 2024 [1]. I question this choice. Coding models have improved significantly in the past 2 years. [1] https://qwen.ai/blog?id=qwen2.5-coder https://qwen.ai/blog?id=qwen2.5-coder
- sandos 1mo agoI just asked my current LLM for that advice. Funny that they dont block it, I guess they are not very threatened. For my anemic 6GB built-on 14GB Qwen seems to be the best bet, not great reviews but from my limited testing its pretty impressive.
- bhelkey 1mo agoAre you familiar with Ollama [1]? It is a particularly easy to use tool to download and run local models. They sort models recent popularity and specify size for the various quantization levels. I would try using ~1/2 your available ram and iterate from there. If you have 32GB of RAM, I would give Qwen 3.8 a try. All you would have to do is run "ollama pull qwen3.8:27b" then "ollama run qwen3.8:27b". If you have 16GB of RAM, I would try Gemma 4. [1] ollama.com