3 ms·
When the Unsloth quant of the flash model does appear, it should show up as unsloth/... on this page: https://huggingface.co/models?other=base_model:quantized:
by a_e_k 9mo ago
When the Unsloth quant of the flash model does appear, it should show up as unsloth/... on this page:
https://huggingface.co/models?other=base_model:quantized:zai-org/GLM-4.7-Flash https://huggingface.co/models?other=base_model:quantized:zai...
Probably as:
https://huggingface.co/unsloth/GLM-4.7-Flash-GGUF https://huggingface.co/unsloth/GLM-4.7-Flash-GGUF
- homarp 9mo agoit'a a new architecture. Not yet implemented in llama.cpp issue to follow: https://github.com/ggml-org/llama.cpp/issues/18931 https://github.com/ggml-org/llama.cpp/issues/18931
- dumbmrblah 9mo agoOne thing to consider is that this version is a new architecture, so it’ll take time for Llama CPP to get updated. Similar to how it was with Qwen Next.
- cristoperb 9mo agoApparently it is the same as the DeepseekV3 architecture and already supported by llama.cpp once the new name is added. Here's the PR: https://github.com/ggml-org/llama.cpp/pull/18936 https://github.com/ggml-org/llama.cpp/pull/18936
- khimaros 9mo agohas been merged