5 ms·
>But I would suggest using UD-IQ3_XXS for 10.9GB for 16GB machines or Q2_K_XL For those of us with a 16GB GPU, how do they compare with ExllamaV4 at 4-bit (4.0
by idonotknowwhy 1mo ago
>But I would suggest using UD-IQ3_XXS for 10.9GB for 16GB machines or Q2_K_XL
For those of us with a 16GB GPU, how do they compare with ExllamaV4 at 4-bit (4.0bpw)?
It looks like that fits in 12.5GB of VRAM since embedding are left in DRAM, Unsloth Studio and other llama.cpp derivatives have to load these weights in VRAM for tied embedding models like Qwen3.8.
ExllamaV3 4.0bpw fits in 12.5G of VRAM and beats IQ4_XS according to the measurements here: [turboderp/Qwen3.8-27B-exl3](https://huggingface.co/turboderp/Qwen3.8-27B-exl3 https://huggingface.co/turboderp/Qwen3.8-27B-exl3)
But those were compared against UD2.0 I guess. Also plans to support these (SOTA) quants in Unsloth Studio?