4 ms·
It's in Unsloth Desktop already. Looks like it's 73GB, so 128GB Mac or Strix Halo etc will work. Exciting!
by pram 1mo ago
It's in Unsloth Desktop already. Looks like it's 73GB, so 128GB Mac or Strix Halo etc will work. Exciting!
- cwizou 1mo agoDownload is available, but likely need to wait for an update, I get this which is understandable with the architectural change : Original error: llama.cpp does not support this GGUF's model architecture ('qwen4exp') Edit : Saw the pull request, should arrive soon enough https://github.com/ggml-org/llama.cpp/pull/27742 https://github.com/ggml-org/llama.cpp/pull/27742
- andy99 1mo agoI only see a 1-bit quant posted on unsloth HF and it’s 72.5 GB. Is that what you mean? That’s much bigger than I expected. If you can’t run a 4 bit quant in on Strix Halo it becomes a lot less interesting. https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF
- agile-gift0262 1mo agoIn their page they say it will need at least 112GB[0], so including context, that would be a tight fit. I'm also hoping I can make a q4 fit on my 128GB strix halo [0]: https://unsloth.ai/docs/models/qwen3.8-next#qwen3.8-flash-next-requirements https://unsloth.ai/docs/models/qwen3.8-next#qwen3.8-flash-ne...
- walrus01 1mo agoin llama-server PR 27742 it fits fine in 128GB RAM on a CPU only system , this is with --load-mode mlock to stuff the whole thing persistently into memory at llama-server launch time, no mmap 0.01.033.250 I common_memory_breakdown_print: | memory breakdown [MiB] | total free self model context compute unaccounted | 0.01.033.261 I common_memory_breakdown_print: | - Host | 118186 = 106166 + 8898 + 3122 | 0.01.092.684 I common_params_fit_impl: projected to use 118186 MiB of host memory vs. 128855 MiB of total host memory
- Wheen 1mo agoJust a hunch, but it might be because of the 51B parameter n-gram embedding. At 125B, you'd expect ~16gigs for a 1-bit quant. Add 51gigs for the n-grams and you're not far off the actual size. If that's true, it'd scale linearly with number of bits in the quant with an offset of about 51gigs. So Q4 should be a bit bigger than 82gigs, I'd guess in the 90s (as opposed to a ~280gig q4 if the whole 70gigs of the 1-bit quant scaled linearly).
- dist-epoch 1mo ago73GB for the 1 bit model...
- naasking 1mo agoThat probably includes the 51b ngrams too. It's possible that those could be streamed from NVMe on-demand. The Engram paper that developed this technique streamed from RAM to VRAM at only ~1% performance degradation, but these strix halo boxes and the spark have much slower memory, so it's possible moving down another rung on the memory hierarchy wouldn't affect their performance too much. This will almost certainly require changes to llama.cpp or vllm to do it right.
- naasking 1mo agoThis guy claims 6% throughout hit for this approach: https://x.com/0xBakeer/status/2092694905978237224?s=20 https://x.com/0xBakeer/status/2092694905978237224?s=20 Crazy how fast things move these days.
- petu 1mo agoIt's not 1 bit. It's ~4bit for n-gram and ~2.8bit for the model. Not idea why it's called Q1, but likely it's preliminary quant just for PR testing / very likely to be remade after llama.cpp support is merged.