3 ms·
Not being able to inline and having to branch on every call to a simd function can sometimes make it slower than the basic scalar version.
by elabajaba 3y ago
Not being able to inline and having to branch on every call to a simd function can sometimes make it slower than the basic scalar version.
- Pet_Ant 3y agoJust a thought, but would it be possible to hot patch at the time of loading the binary? I realise it might require updates to the binary format, but it might be very well justified.
- anonymoushn 3y agoThis sounds similar to shipping all these routines in dynamic libraries and loading the right one at runtime. I'm not sure if the memory cost of having all these duplicate simd kernels is a big deal though.
- dzaima 3y agoYou sould branch at the level where inlining doesn't make sense, which would usually be some function wrapping the big loop, which should be rather free. Which is the same situation as on x86-64 if you want to target pre-AVX2/post-AVX2/AVX-512.