5 ms·
The paper suggests FFmpeg uses intrinsics which is not correct. There have been many SIMD abstraction layers created in the past but none of them will beat the
by kierank 3y ago
The paper suggests FFmpeg uses intrinsics which is not correct.
There have been many SIMD abstraction layers created in the past but none of them will beat the raw speed of handwritten assembly. Try and implement something like vpternlogd in one of these abstraction layers.
- dist1ll 3y agoThe main abstraction of intrinsics is register allocation, right? Is there anything else that can be gained by handwritten asm?
- camel-cdr 3y agoFor rvv specifically there are a few things that aren't possible using the intrinsics abstraction. E.g. in asm you can run the same instruction sequence with different vtype (element width and LMUL).
- janwas 3y agoWe are able to do the same with Highway's RVV :)
- dzaima 3y agoI believe what camel-cdr is saying is being able to run the same code without duplication (say, a loop) which has no vsetvl-s inside, by conditionally choosing either an initial "vsetvli x0,x0,e32,m1" or "vsetvl x0,x0,e16,m1" or "vsetvl x0,x0,e32,m2" etc, which is just unachievable with the intrinsics as they hard-code vtype in each intrinsic. It's an extremely fun idea (primarily just for code size though), but thankfully (?) its usability is restricted by load/store instrs hard-coding the element type, so the main use of this would end up for switching LMUL, which has very limited usefulness. What Highway can support is generating multiple loops of different vtype from the same code, which effectively achieves the same thing, at the cost of machine code duplication.
- camel-cdr 3y ago> for switching LMUL, which has very limited usefulness I currently have a quite usefull use case for it, I'm concerting utf8 to utf32 and if I've got an average utf8 character size of above 2 I could reduce the LMUL for that loop iteration. This shouldn't actually improve performance that much in good rvv implementations, since you can use vl and not LMUL to schedule your execution units. Sadly this is currently not the standard, and ara is the only implementation, that does this I know of. I think this wouldn't even be about code size reduction, consider an input, where there is basically a 50/50 probability LMUL can be reduced, that would be horrible for the branch predictor, but with only a branch over vsetvl, this could behave as a conditional vsetvl via instruction fusion. We'll have to see if such optimization become relevant once there is more hardware out there.
- dzaima 3y agoThat's an interesting use-case, though I wouldn't be surprised if some impls really wouldn't like LMUL dynamically switching at runtime a lot (i.e. something like LMUL being forwarded at decode-time, so it couldn't decode after an unknown-LMUL vsetvl, ruining perf)
- janwas 3y agoOh, I see, thanks for clarifying. Yes, I was referring only to "same source code" and agree our approach would generate multiple copies of the instructions.
- anonymoushn 3y agoI think so. For the most part I'm happy to have the compiler rearrange my intrinsics and decide where the spills should go (inevitably you have spills if you unroll the loop enough times to avoid stalls, because these instructions tend to be like "you can have 6 in flight at a time but that would take 18 registers lol". The main problem I've met here is that the compiler will emit useless instructions to narrow or widen integers in general-purpose registers (e.g. the result of movemask) but this can be solved by looking at the generated assembly and fixing the high-level code.
- nkurz 3y agoIf you are trying to maximize port utilization, sometimes the exact instruction ordering can make a big difference. Compilers often want to "hoist" loads to the top, which can sometimes reduce performance. And sometimes you need a particular addressing mode to avoid a bottleneck. In general, kierank is right: if you want to full optimize something, and you know what you want the actual code to look like, just write it in assembly. Nothing else gives you full control over loads and stores, and anything else leaves you at the mercy of some future compiler "optimization" stepping in to defeat you.
- dist1ll 3y ago> If you are trying to maximize port utilization, sometimes the exact instruction ordering can make a big difference. Interesting. I thought you'd be at the mercy of the instruction scheduler for aggressive OoO cores.
- brigade 3y agoCustom function ABIs. Though the maintenance overhead really isn't worth the reduced cache footprint, especially since asm writers are allergic to leaving behind any comments. Also instruction scheduling. Low-end Cortex will probably be in-order till the end of time...
- brigade 3y agoIt's technically correct; FFmpeg has a tiny amount of NEON intrinsics for no particularly good reason. (well, if it had a lot the good reason would have been to avoid writing everything twice between A32 and A64...) Despite all the other comments, this doesn't appear to be intended to be used to write SIMD across multiple platforms? Rather, it's to quickly port codebases with lots of existing platform-specific intrinsics to a new platform? For this paper in particular, so that RISC-V can run somewhat optimized code without having to spend thousands of man-years writing new RVV code.
- cyber_kinetist 3y agoI still think SIMD helper libraries (like xsimd or highway) have some good use in numerical computation and graphics, since you have so many complex equations to optimize that it's basically unrealistic to write all of it in assembly. And it's much better in terms of readability, xsimd has lots of operator overloading built in so you can still get readable math equations in SIMD code. Even if you get up to 80% of the achievable performance of assembly it's still a much better improvement then plain scalar code or relying on auto-vectorization. (And if that isn't enough you can start optimizing in assembly for only the most frequently used functions)
- janwas 3y agoWe actually do use ternlog in several places in Highway :) Whenever we want to use a new immediate arg, we add new ops such as Not, Xor3, Or3, OrAnd, IfVecThenElse that also do something reasonable on other platforms. BTW this reminds me of a colleague grumbling that what should have been a 20-minute patch to ffmpeg took a day, because it was written in assembly. It is also quite possible to have large slowdowns due to assembly - all it takes is to forget a v prefix (VEX encoding), whereas intrinsics take care of that.
- kierank 3y agoDo you actually implement all permutations of vpternlogd? The lightweight macro layer in ffmpeg takes care of v prefixes. In FFmpeg, x264 and dav1d there are many different examples of code that couldn't be written in intrinsics or other abstraction layer. https://twitter.com/FFmpeg/status/1705543447245988245?t=Ul9ePc-raW4R5Acu8rhRSA&s=19 https://twitter.com/FFmpeg/status/1705543447245988245?t=Ul9e...
- anonymoushn 3y agoYeah, I think non-asm users have to use inlining to avoid clobbering all the vector registers on sysv ABI. I haven't really encountered cases where avoiding inlining is super important though.
- janwas 3y agoAs mentioned, we implement what applications are using/requesting. Do we know how many permutations are used in ffmpeg? hm, I vaguely remember there was a vzeroupper problem, perhaps one fell through the cracks. Interesting, can you share more details on the magic? Looks mainly like function call overhead. If functions aren't called often, we can inline (by moving into headers or enabling LTCG/LTO). If they are called often, are visible to the compiler, have internal linkage, but shouldn't be inlined, I'd be curious to learn why, and also why the compiler is then generating the full prolog/epilog.