4 ms·
Another example is writing hand optimized matrix and vector operation routines tailored to specific hardware for BLAS libraries [0]. [0] https://en.m.wikipedia
by steppi 3y ago
Another example is writing hand optimized matrix and vector operation routines tailored to specific hardware for BLAS libraries [0].
[0] https://en.m.wikipedia.org/wiki/Basic_Linear_Algebra_Subprograms https://en.m.wikipedia.org/wiki/Basic_Linear_Algebra_Subprog...
- KeplerBoy 3y agoIs this really still a thing? Do people go further than using instrinsics for let's say AVX?
- retrac 3y agoSure. You'll see it very often in codec implementations. From rav1e, a fast AV1 encoder mostly written in Rust: https://github.com/xiph/rav1e/tree/master/src/x86 https://github.com/xiph/rav1e/tree/master/src/x86 Portions of the algorithm have been translated into assembly for ARM and x86. Shaving even a couple percent off something like motion compensation search will add up to meaningful gains. See also the current reference implementation of JPEG: https://github.com/libjpeg-turbo/libjpeg-turbo/tree/main/simd/x86_64 https://github.com/libjpeg-turbo/libjpeg-turbo/tree/main/sim...
- steppi 3y agoYeah. I'm going to be helping to work on expanding CI for OpenBLAS and have been diving into this stuff lately. See the discussion in this closed OpenBLAS issue gh-1968 [0] for instance. OpenBLAS’s Skylake kernels do rely heavily on intrinsics [1] for compilers that support them, but there's a wide range of architectures to support, and when hand-tuned assembly kernels work better, that's what are used. For example, [2]. [0] https://github.com/xianyi/OpenBLAS/issues/1968 https://github.com/xianyi/OpenBLAS/issues/1968 [1] https://github.com/xianyi/OpenBLAS/blob/develop/kernel/x86_64/cgemm_kernel_8x2_skylakex.c https://github.com/xianyi/OpenBLAS/blob/develop/kernel/x86_6... [2] https://github.com/xianyi/OpenBLAS/blob/23693f09a26ffd8b60ebdfbfd79ffa5b649d1289/kernel/arm64/cgemm_kernel_8x4_cortexa53.c#L76-L290 https://github.com/xianyi/OpenBLAS/blob/23693f09a26ffd8b60eb...
- KeplerBoy 3y agointeresting stuff. thanks for the links
- mikebenfield 3y agoFWIW I've found that compilers' code generation around intrinsics is often suboptimal in pretty obvious ways, moving data around needlessly, so I resort to assembly. For me this has just been for hobby side projects, but I'm sure people doing it for stuff that matters run into the same issue.
- riceart 3y ago> Is this really still a thing? Why wouldn’t it be? Compilers haven’t advanced tremendously in the past two decades in terms of optimizations and don’t have much new to add to high performance SIMD numeric kernels.