4 ms·
Does all the 1.7GB of the decoded video get copied to the GPU? Or is there some playback controller that knows how to read the “delta format” from the codecs an
by choxi 5y ago
Does all the 1.7GB of the decoded video get copied to the GPU? Or is there some playback controller that knows how to read the “delta format” from the codecs and only copies deltas to the framebuffer?
It still blows my mind that we can stream video at 60FPS. I was making an animation app that did frame-by-frame playback and 16.6ms goes by fast! Just unpacking a frame into memory and copying it to the GPU seemed like it took a while.
- pavlov 5y agoYou shouldn’t copy the frame data to the GPU (assuming that’s literally what your code was doing). Instead create a GPU texture that’s backed by a fixed buffer in main memory. Decode into that buffer, unlock it, and draw using the texture. The GPU will do direct memory access over PCIe, avoiding the copy. The CPU can’t be writing into the buffer while the GPU may be reading from it, so you can either use locks or double buffering to synchronize access.
- danachow 5y ago> Instead create a GPU texture that’s backed by a fixed buffer in main memory. Decode into that buffer, unlock it, and draw using the texture. The GPU will do direct memory access over PCIe, avoiding the copy. With a dedicated GPU with its own memory there still is usually a memory to memory copy, it just doesn’t have to involve the CPU.
- pavlov 5y agoYeah, like so many other things in the GPU world, main RAM texture storage is more of a hint to the graphics card driver — "this buffer isn't going away and won't change until I explicitly tell you otherwise". It definitely used to be that GPUs did real DMA texture reads though, at least in the early days of dedicated GPUs with fairly little local RAM. I'm thinking back to when the Mac OS X accelerated window compositor was introduced — the graphics RAM simply wouldn't have been enough to hold more than a handful of window buffers.
- zoenolan 5y agoI would build on that saying, you should look at double buffering or maybe triple buffering. Frame A is copied to the GPU. Frame B is being decoded into Frame B is copied to the GPU. Frame C is being decoded into Frame C is copied to the GPU. Frame A is being decoded into
- pjc50 5y agoSmart people chuck the encoded video at the GPU and let that deal with it: e.g. https://docs.nvidia.com/video-technologies/video-codec-sdk/nvdec-video-decoder-api-prog-guide/ https://docs.nvidia.com/video-technologies/video-codec-sdk/n... ; very important on low end systems where the CPU genuinely can't do that at realtime speed. Raspberry Pi and so on. > 16.6ms That's sixteen million nanoseconds, you should be able to issue thirty million instructions in that time from an ordinary 2GHz CPU. A GPU will give you several billion. Just don't waste them.
- cogman10 5y agoAgreed. GPUs support decoding a wide range of codecs (even though you are probably using something like H.264). So it doesn't make sense wasting the time to both decode the data and pipe it out to the GPU.