3 ms·
ExLlama is GPU only right? This speedup is for GPU + CPU split use cases.
by nulld3v 3y ago
ExLlama is GPU only right? This speedup is for GPU + CPU split use cases.
- modeless 3y agoOh I see, they are running a 40B model unquantized, whereas exllamav2 would have to use 4-bit quantization to fit. Given the quality of 4-bit quantization these days and the speed boost it provides I question the utility of running unquantized for serving purposes. I see they have a 4-bit benchmark lower down in the page. That's where they ought to compare against exllamav2.