3 ms·
> Why use floats when you can basically have doubles for the same cost? Floats costing the same as doubles is a myth stemming from x87 arithmetic, which is obs
by LPDWORD 7y ago
> Why use floats when you can basically have doubles for the same cost?
Floats costing the same as doubles is a myth stemming from x87 arithmetic, which is obsolete. On an x64 CPU running 64-bit code, your compiler can often pack four floats into one SSE register. Even when that doesn't happen, CPU microcode can likely do more with floats than with doubles.
Lastly, memory bandwidth usage and cache occupancy doubles, which is true even with x87.
- vardump 7y agoThe difference is meaningless in scalar code. Not everything is or can be vectorized. Pretty much no difference between one double and one float in a SSE XMM register. > Even when that doesn't happen, CPU microcode can likely do more with floats than with doubles. I have no idea what that means. As far as I know, there's no CPU microcode dealing with floating point numbers. > Lastly, memory bandwidth usage and cache occupancy doubles, which is true even with x87. So use floats when you have a lot of data.
- LPDWORD 7y agoIf you already knew there was a cost difference, why didn't you just say so? > The difference is meaningless in scalar code. Not everything is or can be vectorized. Multiple scalars can be packed into one SSE register even in scalar code. This is not the same as vectorization. > Pretty much no difference between one double and one float in a SSE XMM register. Even then, some arithmetic operations have higher throughput with floats than with doubles. > I have no idea what that means. As far as I know, there's no CPU microcode dealing with floating point numbers. Your CPU doesn't execute the ISA directly in hardware, it first converts it into architecture specific micro-ops. That includes FP operations. > So use floats when you have a lot of data. That's not the point. A lot of the time, the cost difference will indeed be irrelevant. The point is that there is a cost difference.
- vardump 7y ago> Multiple scalars can be packed into one SSE register even in scalar code. This is not the same as vectorization. They can, but can you honestly call it a common case for scalar code? > Even then, some arithmetic operations have higher throughput with floats than with doubles. By far the most (90-99%) of FP computation is additions and multiplications (or fused multiply adds). For scalar case (1 double or float in SSE register), they take precisely as long on modern x86 hardware. Sure, float div executes in 11 instead of 13-14 clocks for doubles and I'm sure transcendentals are even worse, but they're rarely needed. Even then, if the dependency chain allows, the cost is often OoO scheduled away in integer dominated code. > Your CPU doesn't execute the ISA directly in hardware, it first converts it into architecture specific micro-ops. That includes FP operations. Except that SSE instructions pretty much are micro-ops as-is. Despite similarly sounding term, microcode has nothing to do with micro-ops. > That's not the point. A lot of the time, the cost difference will indeed be irrelevant. The point is that there is a cost difference. Well, I've written a lot of SSE, AVX etc. SIMD code. There sure is a big difference when you're processing large amounts of data. But... I've seen floats introducing silly precision related bugs [0] and a ton of useless float -> double -> float conversion chains. Most of the time most programmers should default to double. [0]: Example: https://randomascii.wordpress.com/2012/02/13/dont-store-that-in-a-float/ https://randomascii.wordpress.com/2012/02/13/dont-store-that...