4 ms·
What is a reasonable lower limit on latency that I could expect from a GPU pipeline? For example, I know that if I use WebGL and issue a `readPixels` call then
by Strilanc 6y ago
What is a reasonable lower limit on latency that I could expect from a GPU pipeline?
For example, I know that if I use WebGL and issue a `readPixels` call then that seems to take on the order of 1 to 10 milliseconds no matter what I do. This suggests that if I needed to react to something in under 100 microseconds, I probably shouldn't have the computation of the reaction involve WebGL. But is that true of GPU compute in general, or just an artifact of the abstraction exposed by WebGL?
- jdashg 6y agoThere's a bunch of intractable latency, but there are patterns that can help maintain pipelining, if you don't need the data immediately: https://developer.mozilla.org/en-US/docs/Web/API/WebGL_API/WebGL_best_practices#Non-blocking_async_data_downloadreadback https://developer.mozilla.org/en-US/docs/Web/API/WebGL_API/W...
- Strilanc 6y agoRight, but if I go closer to the metal how much of the latency is really truly intractable? 100 microseconds? 1 millisecond? 10 milliseconds?
- bl0b 6y agoCan't speak for WebGL APIs, but in OpenGL (and other C++ graphics/GPGPU platforms), it is certainly possible (and necessary) to write fully interactive real-time apps that rely on a few rounds of communication between GPU and CPU code every frame. At 60 fps, you have 16 ms per frame. As long as you aren't copying huge arrays back and forth, I'd say you can budget in the 1 ms order-of-magnitude per data transfer. I don't know much of that would be avoidable if you were actually closer to the metal than the raw OpenGL C++ API call
- kllrnohj 6y agoYou do have to design around the GPU typically being almost like a server. It's very fast, but also "very far" away - both in latency & in bandwidth. PCI-E 3.0 x16 is 'only' 16GB/s. That's fast for I/O, but still in the ballpark of being like I/O & not like RAM. And that's if you even get x16 electrical - x8 for GPUs is pretty common in eg. laptops where the other lanes go to things like thunderbolt 3. Integrated graphics are a different story, they often just have a direct connection to DRAM and CPU<->GPU communication is therefore super fast & low latency. In theory, anyway, if the API abstractions even let you share memory between the two. So that means even for things where GPUs should do really well at it, like summing two arrays together, doing it on the CPU can still be a lot faster if the inputs & outputs are local to the CPU.
- rrss 6y agowith cuda or opencl, you can launch a kernel and get results back in something like 10 microseconds (for trivial 'results' like setting a flag or something). Varies by hardware, platform, etc, but that's the ballpark.