5 ms·
To clarify, it's 94x the performance of the naive C implementation in just one type of filter. On the same filter, the table posted to twitter shows SSSE3 at 40
by cjensen 2y ago
To clarify, it's 94x the performance of the naive C implementation in just one type of filter. On the same filter, the table posted to twitter shows SSSE3 at 40x and AVX2 at 67x. So maybe just a case where most users were using AVX/SSE and there was no reason to optimize the C.
But this is just one feature of FFmpeg. Usually the heaviest CPU user is encode and decode, which is not affected by this improvement.
It's interesting and good work, but the "94x" statement is misleading.
- Temporary_31337 2y agoHow does this compare to the naive GPU performance on an RTX GPU?
- H8crilA 2y agoThe GPU is a lot slower than the CPU - per core/thread. Otherwise you would need to select different core/thread counts on both sides, or set something like a power limit in watts. When it comes to watts the process node (nanometers) will largely determine the outcome.
- astrange 2y agoIf the entire program was on the GPU, that specific part would be extremely fast. But the rest of it is on the CPU because GPU cores aren't any good at largely serial things like video decoding. So it doesn't matter.
- mywittyname 2y ago> GPU cores aren't any good at largely serial things like video decoding I'm surprised to hear this, considering GPUs are often used for video encoding and decoding. For Nvidia cards, this is called NVDEC, and AMD/Intel both have corresponding features for their video cards.
- astrange 2y agoThat doesn't use the GPU cores. They just happen to package dedicated video hardware alongside it, but it's more similar to a CPU with a specialized DSP attached. Any kind of decompression isn't fully parallelizable. If you've found any opportunities, that means the compression wasn't as efficient as it theoretically could be. Most codecs are merciful and eg restart the entropy coder across frames, which is why the multithreaded decoding in ffmpeg is able to work. (But it comes with a lossless video codec called ffv1 that doesn't allow this.)
- robbie-c 2y agoYeah, I wrote my CS dissertation on this. It started as me writing a GPGPU video codec (for a simplified h264), and turned into me writing an explanation of why this wouldn't work. I did get somewhere with a hybrid approach (use the GPU for a first pass without intra-frame knowledge, followed by a CPU SIMD pass to refine), but it wasn't much better than a pure CPU SIMD implementation and used a lot more power.
- astrange 2y agox264 actually gets a little use out of GPGPU - it has a "lookahead" pass which does a rough estimate of encoding over the whole video, to see how complex each scene is and how likely parts of the picture are to be reused later. That can be done in CUDA, but IIRC it has to run like 100 frames ahead before the speed increase wins over the CPU<>GPU communication overhead.
- renhanxue 2y agoNVENC/NVDEC and similar technologies use what is effectively an ASIC that happens to be integrated on the GPU. It doesn't use the general purpose GPU cores. This is the reason why the limitations on e.g. which codecs it supports are so strictly tied to the hardware generation; it's just fixed function hardware. There's no reason you can't bundle the same kind of ASIC on a CPU too, and indeed Intel does do that with QuickSync. For video game capture/screen recording though (which is a big part of what people tend to do with NVEnc) it might be a bit more convenient for the chip to be on the GPU? I don't know, not a GPU expert.
- zbobet2012 2y agoYou've mis-understood. The 8tap filter is part of the HEVC encode loop and is used for sub pixel motion estimation. This is likely an improvement in encoding performance, but it's only in one specific coding tool.
- chrisco255 2y agoFFMPEG uses hand written assembly across their code base, and while 94x may not be representative everywhere, it's generally true that they regularly outperform the compiler with assembly: https://x.com/FFmpeg/status/1852913590258618852 https://x.com/FFmpeg/status/1852913590258618852 https://x.com/FFmpeg/status/1850475265455251704 https://x.com/FFmpeg/status/1850475265455251704
- deleted 2y ago[deleted]
- refulgentis 2y ago^ this, took my biggest step entry into programming via learning how to get my videos on an iPod Video in 2004, that eventually required compiling ffmpeg and keeping up with it. I'll bet money, sight unseen, that poster above is right its used for HEVC. I'll bet even more money its not some massive out of nowhere win, hand-writing assembly for popular codecs was de rigeur for ffmpeg. Thrust of the article, or at least the headline, is clickbait-y.
- jsheard 2y agoAccording to someone in the dupe thread, the C implementation is not just naive with no use of vector intrinsics, it also uses a more expensive filter algorithm than the assembly versions, and it was compiled with optimizations disabled in the benchmark showing a 94x improvement: https://news.ycombinator.com/item?id=42042706 https://news.ycombinator.com/item?id=42042706 Talk about stacking the deck to make a point. Finely tuned assembly may well beat properly optimized C by a hair, but there's no way you're getting a two orders of magnitude difference unless your C implementation is extremely far from properly optimized.
- evoke4908 2y agoIf and only if someone has spent the time to write optimizations for your specific platform. GCC for AVR is absolutely abysmal. It has essentially no optimizations and almost always emits assembly that is tens of times slower than handwritten assembly. For just a taste of the insanity, how would you walk through a byte array in assembly? You'd load a pointer to a register, load the value at that pointer, then increment the pointer. AVR devices can load and post-increment as a single instruction. This is not even remotely what GCC does. GCC will load your pointer into a register, then for each iteration it adds the index to the pointer, loads the value with the most expensive instruction possible, then subtracts the index from the pointer. In assembly, the correct AVR method takes two cycles per iteration. The GCC method takes seven or eight. For every iteration in every loop. If you use an int instead of a byte for your index, you've added two to four more cycles to each loop. (For 8 bit architectures obviously) I've just spent the last three weeks carefully optimizing assembly for a ~40x overall improvement. I have a *lot* to say about GCC right now.
- deleted 2y ago[deleted]
- deleted 2y ago[deleted]