3 ms·
When Meta releases the quantized 70B it will give another > 2X speedup with similar accuracy: https://ai.meta.com/blog/meta-llama-quantized-lightweight-models/
by GavCo 2y ago
When Meta releases the quantized 70B it will give another > 2X speedup with similar accuracy: https://ai.meta.com/blog/meta-llama-quantized-lightweight-models/ https://ai.meta.com/blog/meta-llama-quantized-lightweight-mo...
- YetAnotherNick 2y agoYou don't need quantization aware training on larger models. 4 bit 70b and 405b models exhibit close to zero degradation in output with post training quantization[1][2]. [1]: https://arxiv.org/pdf/2409.11055v1 https://arxiv.org/pdf/2409.11055v1 [2]: https://lmarena.ai/ https://lmarena.ai/
- WanderPanda 2y agoI wonder why that is? because they are trained with dropout?
- david-gpu 2y agoProbably because of how bloody large they are. The quantization errors likely cancel each other out over the sum of so many terms. Same reason why you can get a pretty good reconstruction when you add random noise to an image and then apply a binary threshold function to it. The more pixels there are, the more recognizable will be the B&W reconstruction.
- ipsum2 2y agoProbably not. Cerebras chip only has 16bit and 32bit operators.