4 ms·
Do you actually implement all permutations of vpternlogd? The lightweight macro layer in ffmpeg takes care of v prefixes. In FFmpeg, x264 and dav1d there are
by kierank 3y ago
Do you actually implement all permutations of vpternlogd?
The lightweight macro layer in ffmpeg takes care of v prefixes.
In FFmpeg, x264 and dav1d there are many different examples of code that couldn't be written in intrinsics or other abstraction layer.
https://twitter.com/FFmpeg/status/1705543447245988245?t=Ul9ePc-raW4R5Acu8rhRSA&s=19 https://twitter.com/FFmpeg/status/1705543447245988245?t=Ul9e...
- anonymoushn 3y agoYeah, I think non-asm users have to use inlining to avoid clobbering all the vector registers on sysv ABI. I haven't really encountered cases where avoiding inlining is super important though.
- janwas 3y agoAs mentioned, we implement what applications are using/requesting. Do we know how many permutations are used in ffmpeg? hm, I vaguely remember there was a vzeroupper problem, perhaps one fell through the cracks. Interesting, can you share more details on the magic? Looks mainly like function call overhead. If functions aren't called often, we can inline (by moving into headers or enabling LTCG/LTO). If they are called often, are visible to the compiler, have internal linkage, but shouldn't be inlined, I'd be curious to learn why, and also why the compiler is then generating the full prolog/epilog.