3 ms·
While CORDIC is great for fixed point, it has limitations for floating point. The original 8087 fsin and fcos instructions used CORDIC, but later versions of th
by GeertB 8y ago
While CORDIC is great for fixed point, it has limitations for floating point. The original 8087 fsin and fcos instructions used CORDIC, but later versions of the architecture switched to polynomial approximations, see https://software.intel.com/sites/default/files/managed/f8/9c/x87TrigonometricInstructionsVsMathFunctions.pdf https://software.intel.com/sites/default/files/managed/f8/9c.... Today it's possible to develop implementations of these elementary functions on x86 CPUs that are more precise and more performant using regular multiply/addition/fused multiply add than even the current improved post-CORDIC fsin and fcos functions.
The main issue is that having an instruction executing a fixed-function block with a given (high) latency and little if any pipelining tends to be far worse than having many more fully pipelined multiply/add instructions. The other issue is that argument reduction and approximation over the reduced domain are not independent. For some parts of the domain, such as computing the sine of a number very close to a multiple of pi, you may need to spend more cycles reducing the argument accurately to counter cancelation effects. However, as the reduced argument is then very close to zero, a simple polynomial suffices.
So, for most modern systems, I'd put the effort in efficient pipelined fused-multiply-add and use that for all elementary functions. Fixed-function hardware for elementary functions has generally been proved sub-optimal.
- seedless-sensat 8y agoThe author's method is trivially pipeline-able into up to 32 stages.
- im3w1l 8y ago> For some parts of the domain, such as computing the sine of a number very close to a multiple of pi, you may need to spend more cycles reducing the argument accurately to counter cancelation effects. I'm curious about this. How would you compute FLOAT32_32 (or FLOAT64_MAX) mod PI with correct rounding?
- stephencanon 8y agoThe most approachable explanation I know of is K C Ng’s: https://www.scribd.com/document/64949170/Ng-Argument-Reduction-for-Huge-Arguments-Good-to-the-Last-Bit https://www.scribd.com/document/64949170/Ng-Argument-Reducti...