4 ms·
I'd argue most SIMD hardware instruction sets do support that. x86 has had pshufb since SSSE3 (most desktop+laptop, and ~all mobile x86), providing an arbitrar
by mtklein 11y ago
I'd argue most SIMD hardware instruction sets do support that. x86 has had pshufb since SSSE3 (most desktop+laptop, and ~all mobile x86), providing an arbitrary byte shuffle across 16 bytes, and NEON has vtbl, pretty much the same but limited to 8 bytes of output per instruction.
Now, I will admit that those instructions are not always a good idea (particularly on mobile x86, where pshufb is often several cycles), and they're essentially never a good idea when a specialized instruction (e.g. punpcklbw, vtrn) can do the job.
- ndesaulniers 11y ago> x86 has had pshufb since SSSE3 SIMD.js uses SSE2 as the baseline, due to NEON compatibility, though our implementer and TC39 champion, sunfish, can probably answer more in depth.
- sunfish 11y agoIt's not just SSE2; SIMD.js also includes things like Int32x4.mul, which is a little tricky without SSE4.1's pmulld. It's kind of a balancing act between several concerns. Also, it's a base, and we definitely plan to iterate and add more features on top of it.
- sunfish 11y agoFor SIMD types with elements other than Int8, in addition to the pshufb, there's also the cost of computing the pshufb byte indices. For additional context, the current version of SIMD.js is aimed at covering the basics that strike some balance of being fast on most hardware and being useful. It's also just the beginning, and we expect it'll evolve to add many more features, quite likely including pshufb-like functionality.