4 ms·
I have written a llm compression quantization library called glq because of the high ram prices. Glq uses qtip Trellis quantization which are efficient at low b
by acd 2mo ago
I have written a llm compression quantization library called glq because of the high ram prices. Glq uses qtip Trellis quantization which are efficient at low bpw 2-4 bits. As a gamer and ai developer I want to squeeze more out of the same hardware. Glq runs with vLLM.
Open source
https://github.com/cnygaard/glq https://github.com/cnygaard/glq
Pypi glq
https://pypi.org/project/glq/ https://pypi.org/project/glq/