4 ms·
So to clarify, this still leverages the fixed function hardware video encoders present alongside many GPUs from eg. Intel/NVIDIA/AMD, and would not (for example
by dishsoap 3y ago
So to clarify, this still leverages the fixed function hardware video encoders present alongside many GPUs from eg. Intel/NVIDIA/AMD, and would not (for example) be possible to use on just a GPU itself (without said encoding hardware)?
- shmerl 3y agoIt is for using GPU video related ASICs (same as VAAPI let's say). If you want to use shaders for some video filters, scaling and such - you don't need those extensions. mpv for example can use Vulkan for both at the same time.
- monocasa 3y agoInterestingly, quite a bit of video decode isn't as parallelizable as you might think and isn't a great fit for GPUs. For instance the initial Huffman decoding of the stream is essentially an intrinsically sequential process.
- 0x1ceb00da 3y agoIs there a reason we don't have parallelizable video formats?
- dagmx 3y agoLargely because a lot of compression is contextual to the data around it. That makes most decompression a poor fit for parallelism. The primary choices are space savings and power efficiency. Dedicated hardware and serial decoding/encoding often win out as a result.
- throwup238 3y agoModern video codecs use motion prediction algorithms to encode the information between key frames which depends on the entire frame and the ones immediately preceding it. You can encode the independent groups of pictures (the key frame and all the subsequent predictive frames until the next key frame) in parallel but at least with 4k video you hit memory bandwidth limitations quite quickly.
- magicalhippo 3y agoAlso, to get best compression you don't want to do key frames at even intervals, but rather when it pays off (ie large changes between two frames, like a scene change).
- geraldhh 3y agoseems rather obvious and i would imagine most encoding algos act accordingly
- magicalhippo 3y agoThe point was rather that it makes parallel encoding more difficult. With a fixed interval it is trivial, you just split the incoming stream based on the fixed rate, and let each thread work on separate intervals. With adaptive key frame placement you don't know the intervals up front, and they might have wildly different lengths. Some might be hundreds of frames, some might be just a few.
- geraldhh 3y agogood point
- bayindirh 3y agoNot always.DivX, XviD, etc. preferred constant interval key framing to battle with corruption and other transfer hazards.
- geraldhh 3y agointeresting, thanks
- dataking 3y agoAV1 is parallelizable at the thread and SIMD level. Other formats too I’m sure but I’ve only looked at AV1.
- anvuong 3y agoI guess because information of pixels in a video are not independent in either time or space, while parallelization relies on the independence of either temporal or spatial axis.
- mratsim 3y agoConvolutions and Discrete Cosine Transforms and many matrix operations (rotations, shifts) depends on operations around and ARE parallelizable. It's the non-linear-algebra stuff that is usually problematic.
- pjc50 3y agoWhat does that mean? A video is inherently a series of frames. And using that sequential nature is critical to achieving good compression. However, most formats allow for some parallelism at the "macroblock" level - you can usually decode all 16x16 pixels simultaneously. To some extent you can decode macroblocks separately, but "intra prediction" requires you to have the ones above and to the left available.
- _flux 3y agoIt means that e.g. in H265 you can split the frame into tiles that are much larger than macroblocks: it's like the video is a set of n videos that are then composed to build a complete frame. This allows both encoding and decoding to work in parallel even inside a single frame. In addition, in VR 360 context, it allows decoding only frames the viewer can see. It reduces the compression efficiency, but only a little, because the tiles will be quite large anyway, such as 6 tiles per frame (but can be more for 360 video applications). It can be limited by the number of decoding instances a piece of hardware can have concurrently.
- pjc50 3y agoSo the answer to the original comment "Is there a reason we don't have parallelizable video formats?" is "we do have them".
- bmicraft 3y agoThe answer seems to be "parallel videos can be decoded in parallel"
- philsnow 3y ago> 6 tiles per frame (but can be more for 360 video applications) Seems like for VR video, if it’s split up longitudinally you would only have to decode the tiles that are in the direction the user is looking at any given time, and just let the other bits pass by undecoded.
- p0nce 3y agoThey are, H.265 have "slices" and "tiles" that allow parallel decoding.
- kevincox 3y agoAV1 also has "tiles". This is basically slicing each frame into N separate frames and encoding each mostly separately. You lose a little efficiency (as you can't exploit cross-tile redundancy very well or at all) but unlock a lot of parallelism on both encode and decode.
- bonton89 3y agoDuring the video card shortage I was kind of lamenting that straight encoder cards just don't exist anymore. To get the best encoder acceleration you were stuff fighting with gamers and crytominers over the same video cards even if you didn't need the majority of the hardware for your task.