3 ms·
even 30B model is too large to large on local device (low end). meta should provide free hosted model api to use it.
by heysagnik 2mo ago
even 30B model is too large to large on local device (low end). meta should provide free hosted model api to use it.
- Schlagbohrer 2mo agoMeanwhile those of us with 128GB RAM plus some VRAM don't have any good modern (last 8 months) open weights models to make use of all that. I don't care if it would run 5 tok/s, I want a smarter model than Qwen3.6 which avoids loops and can handle more context than 80k before crashing.
- heysagnik 2mo agowhy don't you use the quantized version of kimi-k3
- Schlagbohrer 2mo agoI do not see a quantized version of kimi-k3 on huggingface that can fit into 128GB. The smallest Unsloth q version is 594GB https://huggingface.co/unsloth/Kimi-K3-GGUF https://huggingface.co/unsloth/Kimi-K3-GGUF
- Godsend69 2mo ago[dead]