4 ms·
That should be vmulq_lane_f32(), but yeah, lane broadcast is free on a number of NEON operations. Many operations also have built-in narrowing, widening, satura
by ack_complete 3y ago
That should be vmulq_lane_f32(), but yeah, lane broadcast is free on a number of NEON operations. Many operations also have built-in narrowing, widening, saturation, and rounding. One of the more ridiculous ones is vqrdmlah_lane_s16(), which translates to: signed saturating rounding doubling multiply accumulate returning high half (with a lane broadcast).
The downside is that the latencies can be a bit high sometimes compared to other CPUs. 128-bit vector integer adds, for instance, have 2c latency even on an Apple M1.
Another thing to watch out for is that some NEON guides are outdated and only tell you about ARMv7 features, missing some goodies added in ARMv8 like horizontal operations (vaddv) and rounding on conversions other than truncate.