7 ms·
Seems interesting! https://github.com/turboderp/exllama https://github.com/turboderp/exllama "A more memory-efficient rewrite of the HF transformers implement
by helloericsf 2y ago
Seems interesting!
https://github.com/turboderp/exllama https://github.com/turboderp/exllama
"A more memory-efficient rewrite of the HF transformers implementation of Llama for use with quantized weights."