3 ms·
It has the opposite problem; it's drawing from a small pool of skilled contributors,.. The project has 2000+ direct contributors and even more indirect contrib
by SubjectToChange 2y ago
It has the opposite problem; it's drawing from a small pool of skilled contributors,..
The project has 2000+ direct contributors and even more indirect contributors on its mailing lists.
...because so much other incorrect advice thinks it's fine to use autovectorization that doesn't work.
There are few high performance programmers who genuinely believe that autovectorization can compete with hand written assembly.
The recent article here about ffmpeg's use of assembly exclusively these comments, or people thinking it was a joke, even though everyone replying who'd actually used it explained why it was good.
I don't see anyone thinking it was a "joke". Comments range from std::simd to SIMD support in Java/C#. A few others quibble over the problems of hand written assembly, but only one or two users genuinely push back against the assembly. This is hardly persecution.
That said, I don't exactly understand your gripe with those people. Should they be showering ffmpeg et al. with praise or something? Like, it's great that the ffmpeg developers can afford to duplicate the same routines across different architectures and SIMD instruction sets, but hardly anyone else can justify doing that. For everyone else the best they can hope for are custom languages and/or better optimizing compilers.
- astrange 2y ago> The project has 2000+ direct contributors and even more indirect contributors on its mailing lists. I'm one of them, so please just believe me instead of trying to correct me ;) It's an ongoing problem the project talks about that there aren't enough newcomers ready to write more SIMD code with good enough quality. > Should they be showering ffmpeg et al. with praise or something? The top reply is "just do this other thing that the article said was unworkable", so not doing that would be a start. Though, the article could've spent some more time explaining why intrinsics don't work well enough. > but hardly anyone else can justify doing that Other people mostly only target one CPU architecture as they're less important, but they also get paid and ffmpeg developers largely didn't. (These days more of them do, but those people are contributing security work more than performance work I think.) It's similar to how x264 was better than every commercial competitor while working for free, simply because they took more time to think about what they were doing.
- dzaima 2y agoIf ffmpeg can't pull together enough good SIMD developers from its thousands of contributors, then most projects won't be able to get any. Having a problem of "not enough" is already miles better the problem of "having none".
- SubjectToChange 2y agoI'm one of them, so please just believe me instead of trying to correct me ;)… No you aren’t. Or rather, there’s absolutely no reason for me to believe you are. It's an ongoing problem the project talks about that there aren't enough newcomers ready to write more SIMD code with good enough quality. Good programmers are in short supply across the entire industry. Like anything else it’s just a matter of practice. The top reply is "just do this other thing that the article said was unworkable", so not doing that would be a start. A) Get a thicker skin. It’s not the end of the world when people leave comments related to the topic at hand. B) x86inc.asm isn’t a particularly interesting approach to programming assembly. Other people mostly only target one CPU architecture as they're less important,… If hardware portability is a goal then handwritten assembly is even more wasteful. It's similar to how x264 was better than every commercial competitor while working for free, simply because they took more time to think about what they were doing. A lot of it simply comes down to sheer man hours and the quantity/quality of bug reports. No 4d chess, no great geniuses, just “good enough” persistence. Anyway, the problem with handwriting assembly is that such programs are trivial in their complexity and/or given unusually strong guarantees.
- neonsunset 2y agoYes, I support your sentiment. On the topic, in C# (*), it is currently strongly recommended to rewrite any code that used to rely on specific ISA and ISA extensions (SSE4.2, AVX, NEON) to cross-platform methods on `Vector128/256/512<T>` and `Vector<T>` themselves. Also, .NET is getting support for SVE2 now on top of `Vector<T>`, which is nice. * Which has proper portable SIMD API, unlike Java panama vectors in their current shape (codegen and API limitations makes them currently an unsuitable counterpart)
- astrange 2y agoIs it strongly recommended by people who've successfully written performant software on multiple platforms (like ffmpeg), or by compiler engineers? Because this is the kind of thing compiler engineers would like to believe is true, but in practice isn't unless you have the power to make them fix it for you. There are actual differences between ISAs here and if your case deoptimizes on one of them then there wasn't a point in using SIMD at all. (Examples: whether it has a full permute, whether unaligned loads are fast or unusably slow, whether it supports half-floats.) And of course nobody can do an abstraction for MMX, so they just pretend it doesn't exist.
- neonsunset 2y agoThe quote reads as "In C#, it is recommended..." Which means that, in the past, C# did not have cross-platform SIMD abstractions aside from a very limited set of arithmetic operations on Vector<T> so that manually vectorized code had to rely on SIMD intrinsics introduced earlier (AVX2, AdvSimd, etc.). However, .NET 7 introduced a set of common arithmetic, logical and bitwise operations on Vector128/256/512<T>[0] (VectorXXX are common vector width types that used to be consumed by intrinsic APIs exclusively), and subsequently both 7 and 8 also included QoL improvements for more high-level Vector<T>. This change rendered the code that duplicated SIMD paths per-platform mostly obsolete save for certain operations like `Shuffle` (which then got addressed by introducing ShuffleUnsafe which is just a raw platform-specific shuffle with the expectation that the users will account for those manually, or the set of outputs they care about has sufficiently common behavior everywhere). CoreLib itself relies on these APIs now and there is no reason to use platform-specific intrinsics in most situations over cross-platform API. Now, one of the reasons .NET can do this is because it does not have to target such a wide range of platforms regular C code has to deal with: most code out there only ever cares about x86, x86_64 (multiple flavors due to SSE2/4, AVX/2 and AVX512), armv7, armv8a and now also wasm (with packed SIMD), and maybe riscv in the future. Out of those, x86_x64 and armv8a receive most attention and performance investment, which are sufficiently similar save for movemask workhorse emulation of which is suboptimal (community has learned vshrn[1] and other tricks for common operations since then to avoid the issue). With that said, C++ has its own experimental cross-plat SIMD abstraction, and there are many high-quality frameworks that allow to abstract away writing SIMD code manually completely. There is also a Rust crate[2] that offers C#-style SIMD abstraction, so it's not exclusive to C# (or particularly difficult in systems programming languages) but C# is probably the one and only high-level language to offer it with an assurance to emit good codegen (unless you abuse it too hard). [0] https://github.com/dotnet/runtime/blob/main/docs/coding-guidelines/vectorization-guidelines.md https://github.com/dotnet/runtime/blob/main/docs/coding-guid... [1] https://github.com/U8String/U8String/blob/main/Sources/U8String/Helpers/VectorExtensions.cs#L143-L156 https://github.com/U8String/U8String/blob/main/Sources/U8Str... [2] https://github.com/Lokathor/wide https://github.com/Lokathor/wide