3 ms·
On the PIC32, fixed add is ~2 cycles and floating point is is ~60. And about a 2x increase for multiplies. On architectures with floating point hardware some of
by vha3 5y ago
On the PIC32, fixed add is ~2 cycles and floating point is is ~60. And about a 2x increase for multiplies. On architectures with floating point hardware some of the advantage is lost, but it's fantastic for DSP on microcontrollers.
- remram 5y agoDo you have numbers for x86? A web search didn't give me a clear answer, especially for non-vectorized code.
- vha3 5y agoI don't unfortunately. I came up with those numbers by using the timers on the PIC32 to measure execution time, in case that's helpful. Or alternatively (and often more simply) by setting and clearing a GPIO pin before/after the operation being timed, and then using the oscilloscope to measure execution time
- klodolph 5y agoThere are half a million x86 microarchitectures. You can look up a particular microarchitecture in Agner Fog’s instruction tables (link below). There will be both throughput and latency values for different instructions (or something else, depending on microarchitecture). Keep in mind that x86 has both x87 and SSE floating-point, it sounds like what you want is the single precision, non-vectorized SSE instructions (like ADDSS). Also note that most microarchitectures are superscalar and may have multiple units capable of executing the instruction you are looking at. https://www.agner.org/optimize/instruction_tables.pdf https://www.agner.org/optimize/instruction_tables.pdf
- nwallin 5y agoIn general, recent consumer x86 CPUs will be faster at floating point code than fixed point code. Floating point multiply is already faster than integer multiply; fixed point has the additional cost of the bitshift. (which is very small, but non-zero) Addition/subtraction are similar. Mid-range smartphones will generally have acceptable floating point units, and you should stick to floating point math. If you're targeting low end smartphones, you never know what you're going to get. Fixed point is generally only still useful on embedded. If you're building software for fixed point, it's generally because you know exactly what CPU your software is going to run on, and you know it doesn't have a FPU.
- Const-me 5y ago> Fixed point is generally only still useful on embedded. Also multimedia. Modern CPUs compute FP32 floats equally fast as integers. However, when you only need 8 or 16 bits of precision, RAM bandwidth often dominates computations. Most image and video codecs are still using 8 bits per channel. Profit from fixed point can be quite large for these use cases. That’s assuming the implementation is good, with SSE2, AVX2 or NEON SIMD.