3 ms·
I’m also a bit confused by the quantization thing. Why exactly is everybody running the same program on the same file? Why not just include the quantized weight
by braingenious 4y ago
I’m also a bit confused by the quantization thing. Why exactly is everybody running the same program on the same file? Why not just include the quantized weights?
It seems like if somebody figured out the “correct” way to quantize the 7b weights it would make way more sense to just torrent the output rather than distribute a fixed program.
- homarp 4y agoquantization takes lots of RAM: https://github.com/qwopqwop200/GPTQ-for-LLaMa/blob/main/README.md https://github.com/qwopqwop200/GPTQ-for-LLaMa/blob/main/READ... says llama-13B takes 42GB and 33B takes more than 64GB...
- disgruntledphd2 4y agoI quantised all of the models(though the 65Bn one came out weird) and I only have 64Gb of RAM, so dunno.
- sebzim4500 4y agoThere is no reason not to do that, except i) Distributing large files through torrents is slightly annoying if you don't already happen to have a seedbox ii) People are still messing around with quantization settings, they might think that they are a few days away from a much better version iii) No one wants to be sued by Meta. I think the risk is pretty small but not zero.