3 ms·
Are there efforts to include the neccessary context in compilers to autovectorize?
by benob 1y ago
Are there efforts to include the neccessary context in compilers to autovectorize?
- yvdriess 1y agoWhat do you mean with necessary context? Modern compilers all autovectorize really well. Usually writing plain canonical loops with plain C arrays is a good way to write portable optimal SIMD code. The usual workflow I use is to translate the vector notation (RIP Cilk+ array syntax) in my paper notes to plain C loops. The compiler's optimization report (-qopt-report for icx, gcc has -fopt-info-vec and -fopt-info-vec-missed) gives feedback on what optimizations it considered and why it did not apply them. In more complex scenarios it can be helpful to add `#pragma omp simd` pragmas or similar to overrule the C semantics.
- Sesse__ 1y ago> Modern compilers all autovectorize really well. Usually writing plain canonical loops with plain C arrays is a good way to write portable optimal SIMD code. I don't think I've seen this happen once at work (always using recent Clang) the last… three years? And I've written a fair bunch of SIMD code. Autovectorization works fairly well when all of your code is very vertical (no interference between the elements) _and_ the compiler knows that your N is divisible by some pretty number (or else, that code bloat is OK enough that it can unroll and deal with the tail). And in the occasional very standard horizontal case (e.g., sum a bunch of bytes, N % 16 == 0). Otherwise, it's just a sad, sad story IME. For most of the algorithms in the presentation at hand, the compiler will have no chance at autovectorizing them.
- camel-cdr 1y agoI think you can coax most of them into autovectorizing via openmp and a bit of restructuring. Can you give 1-3 examples of these type of problems where you think autovectorization is lacking? I want to give it a try.
- Sesse__ 1y agoLet me give you the last couple of routines that I was involved in vectorizing, no cherry-picking: - Skip to the first unmatched } (i.e., if there's two { in there, you are looking for the third }), ignoring anything inside "", and that any character may be escaped using a backslash. (I've simplified away a bunch of complicating cases, or it would just sound unreasonable. The optimal path depends on whether you have e.g. PCLMUL.) - Skip to the next <, \r, \0 or & (note that the optimal path on SSSE3+ uses pshufb, or similarly VTBL on NEON) and return its position. A typical case is that it's ~10 bytes away, but it could be kilobytes. - Parse the decimal part of a double, given that only accuracy up to 1e-7 is needed. (The optimal solution differs somewhat between NEON and SSE2.) The second one is the only one where I believe that autovectorization could have reasonably done a fair job, since it's so standard (but it didn't in practice).
- camel-cdr 1y agoSo I think I managed to do all of the above with varying degrees of success: https://godbolt.org/z/WY99vxs76 https://godbolt.org/z/WY99vxs76 * parse_fract: I got that one down to 23 instructions on icx, although gcc took 81 and clang 91 Since both of the other two return an index, I decided to keep it simple and use a 128 iteration inner loop that accumulates the index into a 8-bit integer, so I don't have to widen. 128 instead of 256, because I needed a sentinel value. * find_unmatched: obviously the compiler couldn't figure out the clmul trick. icx: 0.86 instr/byte, gcc: 0.625 instr/byte, clang at failed to vectorize the +-scan. * find_special: The LUT didn't end up working that well, so I'm doing the four comparisons separately. icx: 0.45 instr/byte, gcc: 0.30 instr/byte, clang: 0.25 instr/byte (I used znver5 as the target for gcc and clang, but znver4 for icx) These were more painful to do than need be, somebody should try it with ISPC and see how that compares. I didn't know about the inclusive scan support in OpenMP before writing this. It's almost good, but the implementations are slightly buggy, and it seems to be designed with threading, not SIMD, in mind. In the sense that you have to write the scan into an array, while in SIMD you don't need that, in multi-threading you need the buffer to do a scan-tree-reduction. The other problem is early exit loops, which should totally be permissible. icc also had support for early_exit, but icx doesn't support it anymore. Wouldn't you "just" need to do an or reduction on the condition mask and break if one bit was set? Thanks for the suggestions. Sounds like you are working on some kind of parser?
- deleted 1y ago[deleted]