3 ms·
No, you can't. The exponent in a float16 is too small. You'd rather convert back to a 32-bit float, do your operations, and then throw away the surplus precisio
by warpspin 3y ago
No, you can't. The exponent in a float16 is too small. You'd rather convert back to a 32-bit float, do your operations, and then throw away the surplus precision and convert back to bfloat16.