2 ms·
I quantize my models with llama.cpp and it's usually one command. Some of their quants are fine-tuned by architecture but it's only to squeeze out every little
by kittikitti 2mo ago
I quantize my models with llama.cpp and it's usually one command. Some of their quants are fine-tuned by architecture but it's only to squeeze out every little performance benefit.
- suprjami 2mo agoUnsloth imatrix data puts their quants at lower KLD than almost all others. It's true they make architecture-specific changes like keeping certain layers at F16 but it's also more than that.