3 ms·
Though, it does seem like anyTrue() and allTrue() are pretty non trivial on ARM with NEON. Those functions seem pretty geared towards x86 movemask/ptest instru
by mtklein 11y ago
Though, it does seem like anyTrue() and allTrue() are pretty non trivial on ARM with NEON. Those functions seem pretty geared towards x86 movemask/ptest instructions. I'd think those would have to be compound, if not serial, on ARM.
- sunfish 11y agoYou're right. There's enough diversity in SIMD instruction sets that no single rule seems sufficient for deciding what to include. Operations which are "fast" on all popular hardware are obviously great, but SIMD.js also includes some operations needed by popular use cases, such as allTrue() and anyTrue(). Also, while these operations are "fast" on x86 as you say, on NEON they're at least no worse than what applications would do otherwise.
- Marat_Dukhan 11y agoI think SIMD.js need more operations, which may not directly map to hardware, but are "at least no worse than what applications would do otherwise": 1. The most obvious operation is multiply-accumulate, which would map to either multiply + add instructions, multiply-add instruction (with intermediate rounding), or FMA instruction. For a linear algebra (BLAS) library, it would be a one-minute fix to make use of this operation, and it would double performance on modern CPUs. 2. Another kind of operations I would like to see is load-and-deinterleave/store-and-interleave, e.g. operation which loads 12 floats of interleaved RGB data and returns 3 Float32x4 vectors with red, green, blue components. These operations map to a single instruction on ARM, and map to a nice sequence of instruction on x86 with SSSE3, and to a much longer sequence of instructions on x86 with SSE2. The best a developer can do with the current SIMD.js API is the analog of very suboptimal SSE2 code. 3. Operations which extend the type, e.g. load 4 16-bit signed integers and extend them to 4x32 vector. The optimal sequence is: - PMOVSXWD xmm, [mem] on x86 with SSE 4.1 - MOVQ xmm, [mem] + PXOR xmm2, xmm2 + PCMPGTB xmm2, xmm + PUNPCKLBW xmm, xmm2 on x86 with SSE2 - VLD1.16 dTemp, [rAddr] + VMOVL.S16 qOut, dTemp on ARMv7. The best the developer can do with SIMD.js now is SSE2 approach. If you are interested in these ideas, I have more operations in mind that would be useful.
- sunfish 11y agoThanks for the suggestions! FMA is definitely something we want to add. The question is what to do when hardware doesn't have FMA, as it's quite expensive to compute manually. Some applications would want to fall back to discrete multiply and add, but of course that rounds differently and other applications don't want implicit rounding differences between platforms. I believe the solution is to give applications a say in what happens, though the details are still being discussed. The load-and-deinterleave/store-and-interleave operations are good ideas too. They're a little more complex than we could accommodate the initial version of SIMD.js, but they're definitely things we should consider. Working with 8-bit and 16-bit data is an area where the current version of SIMD.js is fairly limited overall. Extending loads are a good idea, and in general I'm hoping a future version of SIMD.js will provide a much more complete set of operations. I'm interested in any other ideas you have as well. SIMD.js is being developed at https://github.com/tc39/ecmascript_simd https://github.com/tc39/ecmascript_simd and you're welcome to file issues to send us your ideas. Thanks!
- dbaupp 11y agoYeah, they're compound but don't have to be serial, due to the VPMIN/VPMAX instructions, but definitely much nicer on x86. For instance: int all(uint32x4_t v) { uint32x2_t x = vpmin_u32(vget_low_u32(v), vget_high_u32(v)); uint32x2_t y = vpmin_u32(x, x); return y[0] != 0; } (You can get away with a single UMINV/UMAXV on AArch64.)