3 ms·
AFAIK they used half-precision (Float16)
by skdotdan 6y ago
AFAIK they used half-precision (Float16)
- cs702 6y agoThanks. I should have written "if using Float32," which is what I meant -- instead of "with Float32," which in hindsight reads a bit ambiguous. But regardless of which floating-point representation is used, the number of weights is still in the hundreds of billions... which is insane.