5 ms·
It's cool and everything but I don't go along with the whole why-don't-people-just tone of the whole thing. I think it's fairly understandable to want to limit
by ris 3y ago
It's cool and everything but I don't go along with the whole why-don't-people-just tone of the whole thing. I think it's fairly understandable to want to limit the number of micro-optimized implementations of however many hot routines your library has & maintain them for an ever expanding array of SIMD extensions and variants, each of which you either have to ship with your resulting ever-expanding binaries (if you're doing runtime capability detection) or will likely never see the light of day for 99.9% of users who get shipped a baseline binary for compatibility (if you're doing build-time capability detection).
I'm sure libpng is in a better place with his improvement but I don't think not already having it is purely because they're dummies or "not ambitious enough".
In a world where profile-guided fuzzers are getting better every day, maintaining a high-profile library in a non-memory-safe language can be nerve-wracking enough with just a single implementation of your array-processing loops.
- dundarious 3y agoI don't agree with this argument at all when the domain is a relatively fixed format with a specification.
- Blackthorn 3y agoNot that I'm necessarily disagreeing, but simde has made life a lot easier here. If you want to write SIMD stuff and don't want to use a higher-level library or depend on the compiler to optimize the code down to SIMD (both of which honestly seem like pretty good options these days), simde is worth taking a look at. You write code that looks like native code, but it'll translate it to the other archs for you. Well, "other archs" just means partially-supported NEON here, but I'd still consider that quite a win. No relation with the project, just a user.
- wyldfire 3y agoI was curious about these libraries a few weeks ago and did some searching. Is there one that's got a clearly dominating set of users or contributors? I don't know what a good way to compare these might be, other than perhaps activity/contributor count. [1] https://github.com/simd-everywhere/simde https://github.com/simd-everywhere/simde [2] https://github.com/ermig1979/Simd https://github.com/ermig1979/Simd [3] https://github.com/google/highway https://github.com/google/highway [4] https://gitlab.com/libeigen/eigen https://gitlab.com/libeigen/eigen [5] https://github.com/shibatch/sleef https://github.com/shibatch/sleef
- Blackthorn 3y agoI'm not familiar with all of those but simde is afaict lower level than the others. It doesn't attempt to build any higher level abstractions, just translate between the existing standards and instruction sets.
- kolbe 3y agoThis entire space has a lot of options, and it's still an open question about what is the right approach. The fact is that C++ is a scalar language, with scalar guarantees and scalar assumptions built into their types. Hacking around this requires sacrifices, and different libraries are choosing different approaches. EVO (not featured in your list) is trying to approach it from an algorithms point of view. Highway is portable wrappers around common functionality. I think this is inherently a language-compiler problem. When libraries are using intrinsic functions to override compiler decisions, they interact poorly. One example: here's comparing a scalar versus an avx-512 intrinsic version of accumulate. https://godbolt.org/z/cMafh4Pnv https://godbolt.org/z/cMafh4Pnv IDK if you can read assembly, but the intrinsic version is waaaay worse, and about 4x slower. This is because using intrinsics makes it very hard for the compiler to do other optimizations it otherwise would (e.g. unrolling and reordering).
- deleted 3y ago[deleted]
- janwas 3y agoI agree compiler transforms sometimes do not interact well, but some still do help intrinsics. Interestingly, even `#pragma clang unroll(4)` is ignored here. Often in such cases we have to manually unroll, e.g. introduce multiple accumulators.
- kolbe 3y agoCan I pick your brain? I'm working on yet another C++ based SIMD library, and my head is sort of boiling over.
- tarnith 3y agoI've also run into this thinking, and have been looking to solve it in codebases I'm working on. I've run across: https://github.com/aff3ct/MIPP https://github.com/aff3ct/MIPP but have not worked with it extensively yet. It looks to be a solution to the rewriting X parallel pipeline into Y SIMD extensions. Perhaps something like this, or languages introducing something similar into their standard libraries/modules would be a solution. None of this of course solves the run-time detection of capability/growing binary size to support such.
- janwas 3y agoI totally agree we don't want multiple optimized implementations, but that is no longer necessary. Highway (mentioned below) lets you write your code once but be compiled+specialized for the major SIMD targets. If you're writing say a thousand lines of SIMD, the binary size impact is a few dozen KiB. This is also a purely library solution, so no build system magic is required.