3 ms·
Something I want to point out is _why_ you would want a graph for GPU work. The existing programming model is based on passes. You have one dispatch spin up th
by jms55 2y ago
Something I want to point out is _why_ you would want a graph for GPU work.
The existing programming model is based on passes. You have one dispatch spin up thousands of workgroups, each of which do some computation. Then you issue a barrier, change your shader, and issue another dispatch doing more work, often reading the previous pass's result.
At first glance this seems fine, but it has some inefficiencies.
1. Often the size of the second pass depends on the output of the first. The first pass might determine that some work is unneeded, and therefore the second pass would have less to do. But from the CPU side, you need still need to allocate memory for the worst-case amount of output from the first pass.
2. Towards the end of the first pass, you'll have only a couple of workgroups left running. But you can't start running the second pass yet - you need to wait for _all_ workgroups in the first pass to finish, and then issue a barrier for memory synchronization.
Workgraphs let you fix these problems. You define two nodes A->B that would correspond to your old passes. Then, you only need to allocate a reasonable amount of scratch memory for the first pass's output - you don't need to allocate for the worst case. When A is running, B can immediately run and pickup the results from A via the scratch buffer as A's workgroups complete.
Workgraphs also give you a lot more flexibility with branching based on work type. A common technique is to split the screen into fixed-size tiles, and classify each tile as "cheap" or "expensive" to shade. Then you can run a cheap shading pass only on the cheap tiles, and an expensive shading pass only on the expensive tiles. The problem here is the same as we saw earlier - you need a full barrier and wait for the previous pass to complete before the next can be issued, even though they're independent. And that's with only 2 possible shaders. With more passes, or different length chains of work (e.g. this tile just needs some quick shadows, this other needs shadows and then 3 passes of denoising), it gets more and more inefficient. Workgraphs let you turn all that into nodes, and schedule them much more efficiently.
- flohofwoe 2y agoTo me the downside is that it's basically building an AST with a cumbersome 'AST builder API' instead of designing a new specialized programming language where the dependency tree and 'node payload' is all integrated into a single language (instead of having the traditional split between CPU- and GPU-side languages, and a 3D API as glue sitting inbetween - as more and more work moves to the GPU that traditional split makes less and less sense IMHO). The idea to build an AST with an API instead of a dedicated language feels a lot like GLTF's new interactivity extension, which describes an AST in JSON (blech): https://github.com/KhronosGroup/glTF/blob/interactivity/extensions/2.0/Khronos/KHR_interactivity/Specification.adoc https://github.com/KhronosGroup/glTF/blob/interactivity/exte... PS: also, it would be nice to get an idea what the debugging situation is for workgraphs.