3 ms·
Surely SIMD combined with multiple streams would beat both approaches. (This would be separate streams in each SIMD lane and separate streams in different SIMD
by thrtythreeforty 3mo ago
Surely SIMD combined with multiple streams would beat both approaches. (This would be separate streams in each SIMD lane and separate streams in different SIMD variables.) There are multiple SIMD execution units, just like the 6 scalar units you mention. The latency of SIMD ops will be similar to scalar, except in cases you mention like shifts.
- manucorporat 3mo agogonna try it! i also suspect that that cranelift's vector lowering is not ideal, so the reasons i described could be wrong