4 ms·
Thank you, looking forward to it. I just saw this simple patch to enable MTP (potentially 2x performance) on older GPUs (Kepler etc), so maybe it will work for
by russianGuy83829 3mo ago
Thank you, looking forward to it.
I just saw this simple patch to enable MTP (potentially 2x performance) on older GPUs (Kepler etc), so maybe it will work for you
https://github.com/ggml-org/llama.cpp/pull/25680 https://github.com/ggml-org/llama.cpp/pull/25680
Also, for Qwen, the 4 bit _XL quantization seems to have a good balance of performance to size.