4 ms·
Very interesting description. Are you familiar at all with the details of FPGAs for these very same tasks, especially the EV family of Xilinx Zynq Ultrascale+ M
by jng 5y ago
Very interesting description. Are you familiar at all with the details of FPGAs for these very same tasks, especially the EV family of Xilinx Zynq Ultrascale+ MPSoC? They include hardened video codec units, but I don't know how they compare quality/performance-wise. Thanks!
- izacus 5y agoI'm afraid I don't have any experience with those devices. Most HW encoders however struggle with one thing - the fact that encoding is very costly when it comes to memory bandwidth. The most important performance/quality related process in encoding is having the encoder take each block (piece) of previous frame and scan the current frame to see whether it still exists and where it moved. The larger area the codec scans, the more likely it'll find the area where the piece of image moved to. This allows it to write just a motion vector instead of actually encoding image data. This process is hugely memory bandwidth intensive and most HW encoders severely limit the area each thread can access to keep memory bandwidth costs down and performance up. This is also a fundamental limitation for CUDA/gpGPU encoders, where you're also facing a huge performance loss if there's too much memory accessed by each thread. Most "realtime" encoders severely limit the macroblock scan area because of how expensive it is - which also makes them significantly less efficient. I don't see FPGAs really solving this issue - I'd bet more on Intel/nVidia encoding blocks paired with copious amount of onboard memory. I heard Ampere nVidia encoding blocks are good (although they can only handle a few streams).
- spuz 5y agoThat is interesting context for this quote from the article: > "each encoder core can encode 2160p in realtime, up to 60 FPS (frames per second) using three reference frames." Apparently reference frames are the frames that a codec scans for similarity in the next frame to be encoded. If it really is that expensive to reference a single frame then it puts into perspective how effective this VPU hardware must be to be able to do 3 reference frames of 4K at 60 fps.
- daniellarusso 5y agoI always thought of reference frames as like the sampling rate, so in that sense, is it how few reference frames can it get away with, without being noticeable? Would that also depend on the content? Aren’t panning shots more difficult to encode?
- izacus 5y ago> I always thought of reference frames as like the sampling rate, so in that sense, is it how few reference frames can it get away with, without being noticeable? Actually not quite - "reference frames" means how far back (or forward!) the encoded frame can reference other frames. In plain words, "max reference frames 3" means that frame 5 in a stream can say "here goes block 3 of frame 2" but isn't allowed to say "here goes block 3 of frame 1" because that's out of range. This has obvious consequences for decoders: they need to have enough memory to keep "reference frames" decoded uncompressed frames around in a chance that a future frame will reference them. It also has consequences for encoders: while they don't have to reference frames far back, it'll increase efficienty if they can reuse the same stored block of image data across as much frames as possible. This of course means that they need to scan more frames for each processed input frames to try to find as much reusable data as possible. You can easily get away with "1" reference frame (MPEG-2 has this limit for example), but it'll encode same data multiple times, lowering overall efficiency and leaving less space to store detail. > Would that also depend on the content? It does depend on the content - in my testing it works best for animated content because the visuals are static for a long time so referencing data from half a second ago makes a lot of sense. It doesn't add a lot for content where there's a lot of scenecuts and actions like a Michael Bay movie combat scene.
- bick_nyers 5y agoOnly relatively recently has NVIDIA implemented B-Frames into NVENC. If I am not mistaken, AMD still does not have this capability. I am not deeply well versed in this space though, but if memory bandwidth is such a huge bottleneck, how does the CPU do it so efficiently comparatively? GPU surely wins in this area? Is it just designed that way so that consumer cards can offer realtime speeds? I'm not sure why this couldn't be configurable in some way.
- bick_nyers 5y agoI am by no means an expert, and this by no means is indicative of a video compression FPGA, but I've been looking at .GZIP and .PNG accelerators and it seems that while they deliver incredible speed, it is done so at the worst compression ratio you can fit in the compression spec. Equivalent to .GZIP setting 2, or maybe equivalent to a "super fast" or "ultra fast" video preset. It is important to note that these are lossless algorithms though. Still, it may not make sense to utilize if your application is bandwidth sensitive. If 4k Netflix doubled it's bitrate by switching to FPGA solutions, that would probably be too high of a cost, even for a 20x speedup. At least until high quality internet speeds become more universal of course.