3 ms·
Since OP didn't answer, it's probably multiplying terms by zero to mask them out.
by sounds 4y ago
Since OP didn't answer, it's probably multiplying terms by zero to mask them out.
- maximilianburke 4y agoYou don't need to multiply by zero; there are comparison operations which will generate a mask value which can then be used to select the appropriate value (`vsel` for altivec, `_mm_blend_ps` for sse4, or a combination of `_mm_and_ps` / `_mm_andnot_ps` / `_mm_or_ps` for ssse3 and earlier) At least on Xbox 360 / Cell, high performance number crunching code would often compute more than needed (ie: both the if-case and else-case) and then use a branchless select to pick the right result. It may seem like a waste but it was faster than the combination of branch misprediction as well as shuffling values from one register file to another to be able to do scalar comparisons.
- sounds 4y agoThen they would have said "masking" instead of "folding" Thanks for mentioning masking though, in case anyone thought we didn't already know about that. ;-)
- Const-me 4y ago> `_mm_blend_ps` for sse4 That instruction uses a mask provided at compile time; it encodes the mask as a part of the instruction. SSE 4 set has a similar one, `_mm_blendv_ps`, which does take the mask from another vector register.
- maximilianburke 4y agoGood catch! Thanks, I'll correct it. (edit: oops, can't fix it, it's been too long since I first wrote the comment)