3 ms·
For us less technical folks (in this field), what’s the big take away here / why does this matter / why should we be excited?
by cyrillite 3y ago
For us less technical folks (in this field), what’s the big take away here / why does this matter / why should we be excited?
- buildbot 3y agoTypically, you need to use some tricks for pre-training in lower precision (finetuning seems to work at low precision), with FP16 you need loss scaling for example. With MX, you can train in 6 bits of precision without any tricks, and hit the same loss as FP32.
- imjonse 3y agoFuture hardware implementations for these <8bit data types will result in much larger (number of parameters) models fitting in the same memory. Unless they are standardized, each vendor and software framework will have their own slightly different approach.
- deleted 3y ago[deleted]
- andy99 3y agoOn most hardware, handwritten math is required for all the nonstandard formats, e.g. for quantized int-8 https://github.com/karpathy/llama2.c/blob/master/runq.c#L317 https://github.com/karpathy/llama2.c/blob/master/runq.c#L317 Integer quantization doesn't typically just round, it has scaling and other factors in blocks so it's not just a question of manipulating int8's. And the FP16/FP8 are not supported by most processors so need their own custom routines as well. It would be great if you could just write code that operates with intrinsics on the quanitzed types.