3 ms·
That isn't unquantized, it's de-quantized. They went from Q5 to fp16 for use in Pytorch instead of the GGUF ecosystem.
by MallocVoidstar 3y ago
That isn't unquantized, it's de-quantized. They went from Q5 to fp16 for use in Pytorch instead of the GGUF ecosystem.
- Taek 3y agoI never thought people would be upscaling models by increasing quantization precision. The rationale makes sense bit its also a goofy outcome.
- nullc 3y agoYou should be able to upscale and fine tune to recover performance, I suppose! Clearly we should train a diffusion model to denoise the weights of LLM transformer models. Yo dawg.
- throwaway9274 3y agoYes, that’s correct. Good correction.