4 ms·
Writing good assembly is a niche skill, especially SIMD assembly. Projects like ffmpeg are able to do it because they're pulling from a massive pool of contribu
by SubjectToChange 2y ago
Writing good assembly is a niche skill, especially SIMD assembly. Projects like ffmpeg are able to do it because they're pulling from a massive pool of contributors. In general writing raw assembly should be avoided unless you're genuinely in a position of knowing better.
...people constantly replying "um actually you never need to write anything in assembly" to them.
Honestly, who is saying that?
- astrange 2y ago> Projects like ffmpeg are able to do it because they're pulling from a massive pool of contributors. It has the opposite problem; it's drawing from a small pool of skilled contributors, because not enough people have learned it, because so much other incorrect advice thinks it's fine to use autovectorization that doesn't work. > Honestly, who is saying that? The recent article here about ffmpeg's use of assembly exclusively these comments, or people thinking it was a joke, even though everyone replying who'd actually used it explained why it was good. https://news.ycombinator.com/item?id=39813724 https://news.ycombinator.com/item?id=39813724 (note asm vs intrinsics is a different tradeoff - it doesn't use intrinsics because they aren't actually easier to work with; they are not faster, not more portable, and on Intel not even more readable because of Hungarian notation. Although they are easier to debug.)
- SubjectToChange 2y agoIt has the opposite problem; it's drawing from a small pool of skilled contributors,.. The project has 2000+ direct contributors and even more indirect contributors on its mailing lists. ...because so much other incorrect advice thinks it's fine to use autovectorization that doesn't work. There are few high performance programmers who genuinely believe that autovectorization can compete with hand written assembly. The recent article here about ffmpeg's use of assembly exclusively these comments, or people thinking it was a joke, even though everyone replying who'd actually used it explained why it was good. I don't see anyone thinking it was a "joke". Comments range from std::simd to SIMD support in Java/C#. A few others quibble over the problems of hand written assembly, but only one or two users genuinely push back against the assembly. This is hardly persecution. That said, I don't exactly understand your gripe with those people. Should they be showering ffmpeg et al. with praise or something? Like, it's great that the ffmpeg developers can afford to duplicate the same routines across different architectures and SIMD instruction sets, but hardly anyone else can justify doing that. For everyone else the best they can hope for are custom languages and/or better optimizing compilers.
- astrange 2y ago> The project has 2000+ direct contributors and even more indirect contributors on its mailing lists. I'm one of them, so please just believe me instead of trying to correct me ;) It's an ongoing problem the project talks about that there aren't enough newcomers ready to write more SIMD code with good enough quality. > Should they be showering ffmpeg et al. with praise or something? The top reply is "just do this other thing that the article said was unworkable", so not doing that would be a start. Though, the article could've spent some more time explaining why intrinsics don't work well enough. > but hardly anyone else can justify doing that Other people mostly only target one CPU architecture as they're less important, but they also get paid and ffmpeg developers largely didn't. (These days more of them do, but those people are contributing security work more than performance work I think.) It's similar to how x264 was better than every commercial competitor while working for free, simply because they took more time to think about what they were doing.
- dzaima 2y agoIf ffmpeg can't pull together enough good SIMD developers from its thousands of contributors, then most projects won't be able to get any. Having a problem of "not enough" is already miles better the problem of "having none".
- SubjectToChange 2y agoI'm one of them, so please just believe me instead of trying to correct me ;)… No you aren’t. Or rather, there’s absolutely no reason for me to believe you are. It's an ongoing problem the project talks about that there aren't enough newcomers ready to write more SIMD code with good enough quality. Good programmers are in short supply across the entire industry. Like anything else it’s just a matter of practice. The top reply is "just do this other thing that the article said was unworkable", so not doing that would be a start. A) Get a thicker skin. It’s not the end of the world when people leave comments related to the topic at hand. B) x86inc.asm isn’t a particularly interesting approach to programming assembly. Other people mostly only target one CPU architecture as they're less important,… If hardware portability is a goal then handwritten assembly is even more wasteful. It's similar to how x264 was better than every commercial competitor while working for free, simply because they took more time to think about what they were doing. A lot of it simply comes down to sheer man hours and the quantity/quality of bug reports. No 4d chess, no great geniuses, just “good enough” persistence. Anyway, the problem with handwriting assembly is that such programs are trivial in their complexity and/or given unusually strong guarantees.
- janwas 2y ago> it doesn't use intrinsics because they aren't actually easier to work with; they are not faster, not more portable Well golly, I'll just have to disagree based on >20 years of experience, including several in assembly. asm is only (maybe) faster for the code we manage to get written. From where I sit, video codecs are a rare special case in that the format is standardized, changes only every few years, and has only a few but super-time-critical kernels. For many many other use cases, the situation looks different and productivity matters more. Would you rather get a 10x speedup on 40% of the cycles, or 8x on 80%? BTW the "not more portable" comment is a strawman because intrinsics themselves indeed aren't portable, but a wrapper library on top (such as our Highway) is.
- astrange 2y ago> BTW the "not more portable" comment is a strawman because intrinsics themselves indeed aren't portable, but a wrapper library on top (such as our Highway) is. That's not intrinsics then, it's different abstraction. You could write a wrapper library over inline assembly if you wanted to. (And of course the intrinsics themselves could almost all be implemented as a header using inline assembly too. Since you're probably not relying on the compiler to optimize your intrinsic math. But optimization would be a bit worse because it doesn't know the byte size of each instruction.)
- janwas 2y agohm, sounds almost like macros wouldn't make it assembly anymore :) I'm curious whether you know of any such inline asm wrapper? Seems that this gives the compiler less information than the intrinsics, which largely expand to builtins.
- dzaima 2y agoThe compiler can still do optimizations on intrinsics - clang passes most through its regular optimizations, so you get things like loop unrolling, CSE (quite powerful if you have multiple invocations of the same SIMD thing, deduplicating constant loads or whatnot), and some genuine improvements/reducing what you need to pay attention to (don't need to manually merge to 'vpandn', 'vpand a,b,c; vptest a,a' → 'vptest b,c', sometimes improving shuffles, moving out negation from movmsk of negation of vpcmpeq), though it can of course make things worse too as regular compiler tax. An example of something that inline assembly would handle badly would be broadcast, which x86 pre-AVX-512 only can do with a value already in a SIMD register, or directly from memory, but the programmer almost always will want to provide it as a regular scalar variable, i.e. GPR.