6 ms·
Optimizing your programs for Arm platforms
- astrange 2y agoThis isn't a good article. I would say that if you're trying to rely on `restrict` and autovectorization you're doomed and should write it yourself. Even if it works on one compiler version, it won't work on all of them. (It could possibly work in a language that isn't C and is designed for it; Fortran or shader programs are easier to autovectorize, and something like ISPC starts out "vectorized" and gets "autoscalarized".) This is why ffmpeg writes SIMD in assembly and is more successful than all the people constantly replying "um actually you never need to write anything in assembly" to them.
- ColonelPhantom 2y agoAren't shader programs more like ISPC (or OpenCL/CUDA), in that the programming model is based around 'pretend each SIMD lane is thread'?
- astrange 2y agoDepends on the target architecture. Some GPUs have used vectors in the past, and some people try to run shaders on CPU. (like for OpenCL or for emulation)
- ColonelPhantom 2y ago> Some GPUs have used vectors in the past Which ones? I'm aware of AMD/ATi Terascale using VLIW, but I'm pretty sure that architecture also used SIMD (requiring the use of both VLIW and large 'waves' to achieve maximum occupancy). And running shaders on CPU is, in essence, a similar programming model to ISPC. When you run OpenCL on a CPU I am quite certain that the runtime pretends each lane is a program instance, same as ISPC.
- SubjectToChange 2y agoWriting good assembly is a niche skill, especially SIMD assembly. Projects like ffmpeg are able to do it because they're pulling from a massive pool of contributors. In general writing raw assembly should be avoided unless you're genuinely in a position of knowing better. ...people constantly replying "um actually you never need to write anything in assembly" to them. Honestly, who is saying that?
- astrange 2y ago> Projects like ffmpeg are able to do it because they're pulling from a massive pool of contributors. It has the opposite problem; it's drawing from a small pool of skilled contributors, because not enough people have learned it, because so much other incorrect advice thinks it's fine to use autovectorization that doesn't work. > Honestly, who is saying that? The recent article here about ffmpeg's use of assembly exclusively these comments, or people thinking it was a joke, even though everyone replying who'd actually used it explained why it was good. https://news.ycombinator.com/item?id=39813724 https://news.ycombinator.com/item?id=39813724 (note asm vs intrinsics is a different tradeoff - it doesn't use intrinsics because they aren't actually easier to work with; they are not faster, not more portable, and on Intel not even more readable because of Hungarian notation. Although they are easier to debug.)
- SubjectToChange 2y agoIt has the opposite problem; it's drawing from a small pool of skilled contributors,.. The project has 2000+ direct contributors and even more indirect contributors on its mailing lists. ...because so much other incorrect advice thinks it's fine to use autovectorization that doesn't work. There are few high performance programmers who genuinely believe that autovectorization can compete with hand written assembly. The recent article here about ffmpeg's use of assembly exclusively these comments, or people thinking it was a joke, even though everyone replying who'd actually used it explained why it was good. I don't see anyone thinking it was a "joke". Comments range from std::simd to SIMD support in Java/C#. A few others quibble over the problems of hand written assembly, but only one or two users genuinely push back against the assembly. This is hardly persecution. That said, I don't exactly understand your gripe with those people. Should they be showering ffmpeg et al. with praise or something? Like, it's great that the ffmpeg developers can afford to duplicate the same routines across different architectures and SIMD instruction sets, but hardly anyone else can justify doing that. For everyone else the best they can hope for are custom languages and/or better optimizing compilers.
- dzaima 2y agoAny modern compiler that bothers should be able to autovectorize most practical vectorizable things without issue, even without restrict. Of course there'll be some small inefficiencies or failed autovectorization sometimes, but small and big missed optimizations are in no way at all a problem unique to vectorization, so it's a moot point here.
- kimixa 2y ago"restrict" is getting around language issues where it's easy to "accidently" block things like vectorization, or having to reload values multiple times just in case pointers aliased. There's no compiler in the world that can work around this without more guarantees on the expected behavior from the programmer, as it would be incorrect according to the spec. I've seen "obvious" big wins missed without it.
- dzaima 2y agoBoth gcc and clang check for aliasing at runtime if not provable statically for autovectorization (granted, that can fail if you have reverse/strided/gather addresses, but those are less common; and yeah it does lead to some constant overhead, though likely not significant often). Of minor note is that you can add "#pragma clang loop vectorize(assume_safety)" or "#pragma GCC ivdep" on the respective compilers to a loop to allow them to vectorize anything, which in my experience is much more functional than restrict. And even the vast majority of benefit I've gotten from it was just removing the alias check overhead (though it did catch a case of a reversing loop failing to vectorize due to a 32-bit index variable or something)
- astrange 2y ago> Both gcc and clang check for aliasing at runtime if not provable statically for autovectorization This is often enough to make it unworkable, because it means you're inserting checks into hot loops. Also, if you partially vectorize something yourself you have to write similar setup code, which might involve scalar versions of the loop, but then autovectorization can come by and vectorize those, so now you have duplicate setup code making it worse than nothing.
- BoingBoomTschak 2y agoThe big problem is that gcc/clang don't seem to have a concept of optimization notices, like SBCL does. Nobody is more appropriate than the compiler to warn you that it couldn't optimize something costly and why.
- dzaima 2y agoClang does have "-Rpass-missed=vectorize" among others.
- BoingBoomTschak 2y agoHuh, the more you know.
- Phyx 2y agoGCC has among others -fopt-info-vec-missed
- rarepostinlurkr 2y agoIt’s good so much attention is being given to arm! Apple also recently released more details on optimization https://developer.apple.com/documentation/apple-silicon/cpu-optimization-guide https://developer.apple.com/documentation/apple-silicon/cpu-...
- electricshampo1 2y agoThanks for this link; did not realize that they did this.