3 ms·
I recently tried to implement a 16bit 64k point fft in C using FP32. It thought it would work easily. Turns out it was very difficult, what worked fine in doubl
by tails4e 3y ago
I recently tried to implement a 16bit 64k point fft in C using FP32. It thought it would work easily. Turns out it was very difficult, what worked fine in double gave huge inaccuracies in float. It was due to the large small issue, if you subtract a small FP32 from a large tje result cannot be represented so it's as if the subtraction did not happen. These kinds of errors accumulated and affected the result. I found a workaround for my particular algorithm, but it was eye opening to me that FP32 was not enough to just work in this case.
- hpcjoe 3y agoCatastrophic loss of precision is, as the name implies, catastrophic in terms of the calculation context. For scientific/engineering codes, and things requiring a preservation of resolution for proper functioning, FP32 is rarely sufficient. FP64 is usually better. For ML/AI apps, resolution isn't nearly as important.
- tails4e 3y agoAbsolutely. If anything ML is migrating to lower precision types, like fp16, mx9, mx6, int8, etc.