2 ms·
> How many pieces of code are out there where an easy x4 speedup could be achieved > today if there were written with batched operations from the start? (It a
by yaantc 6y ago
> How many pieces of code are out there where an easy x4 speedup could be achieved
> today if there were written with batched operations from the start? (It also shows
> how limited are the compiler in autovectorizing)
Even before auto-vectorization, I'd love a functional automated "loop unrolling with interleave" that works on large functions. There is a pragma for this in Clang, but when I checked it in clang-9 it didn't work. I'll have to try again as v9 is a bit old now. When this is well supported it will avoid easily "filling the pipe" on multiple issues cores when doing batched operations, without having to manually unroll the loops as is done in VPP for example:
https://gerrit.fd.io/r/gitweb?p=vpp.git;a=blob;f=src/vnet/ip/ip4_forward.c;h=bb70805b4e66f821616946ad895c309e1f08d96f;hb=35ef865678d82b5a6fd3936716de8afb2fd49e60#l1250 https://gerrit.fd.io/r/gitweb?p=vpp.git;a=blob;f=src/vnet/ip...
Manual unrolling works, but getting the same effect with a simple pragma on top of the loop looks so much more attractive ;)