4 ms·
GPUs are setup as large amounts of SIMD blocks of threads with some shared units like registers, cache, ALU, etc for every few blocks of threads. The typical w
by jms55 2y ago
GPUs are setup as large amounts of SIMD blocks of threads with some shared units like registers, cache, ALU, etc for every few blocks of threads.
The typical way you schedule GPU work is to dispatch a very large amount of work all running the same program. So you'd spawn 10 million units of work, to run on 10 thousand thread blocks. Due to differing memory accesses per unit of work, each threadblock will complete their work at a different time. Whenever a threadblock is stuck waiting for memory accesses to complete, or has finished it's work, the GPU scheduler gives it a new unit of work to work on.
If you graph "occupancy" as the percentage of threadblocks busy doing work, then you'd see a spin-up period as threadblocks are filled with work, a steady period where all threadblocks are busy, and then a spin-down period as there's gradually less work available than the number of threadblocks.
If you wanted to run two programs (e.g., check meshes to determine what's visible and cull invisible meshes, then draw the remaining visible meshes), then you'd have a hard gap in between. Threadblocks would spin-up with culling work, work steadily, spin-down until 0 work is left, spin-up with mesh drawing work, work steadily, and then spin-down again. The spin-down period in between the two passes is bad for performance.
Rather than having the GPU go completely idle, wouldn't it be better if, as the GPU is running out of culling work to execute (available work is less than the number of threadblocks available to perform work), you could fill in the gaps by immediately starting the mesh drawing work for meshes that have already passed culling? That way there would be no idle time between passes. Workgraphs let you do this, by specifying execution not as monolithic passes, but as nodes that perform 1 unit of work and produce output for other nodes to consume.
Another benefit is memory allocation required. With the two distinct passes model, you need to allocate the worse-case amount of memory to hold the first pass output (input to the second pass). If you have 10,000 meshes, then you need to allocate space to be able to draw 10,000 meshes for the worst case that they're all visible - there is no runtime memory allocation on the GPU. With workgraphs, the GPU can allocate a reasonable estimate of how much memory it needs for node 1's output (input for node 2), and if the output buffer is full, the GPU can simply stop scheduling node 1 work, and start scheduling node 2 work to pop from the buffer and free up space.
As for whether you can do this with the existing GPU pass-based model, more or less yeah. You can build your own queue and use global atomics to control producer/consumer synchronization. It's called persistent threads. You might even do better than the GPU's built-in scheduling depending on the task, if you hyper-optimize and tune your code. However, it's a lot harder, and dependent on specific implicit behavior of the GPU's scheduler. If the GPU is not smart enough to sleep threadblocks waiting for the atomic "lock" and schedule threadblocks that are holding the lock, then you get a deadlock and your computer freezes until the GPU driver kills your program.
The big reason workgraphs are getting introduced is for Unreal Engine 5's Nanite renderer that came out a few years ago. It uses the persistent threads technique to traverse a BVH, doing cull checks on each node, with passing nodes getting their children pushed onto the queue. BVHs (trees) can be unbalanced, so Nanite uses the persistent threads technique to dynamically load-balance nodes amongst the available threadblocks. With workgroups, Nanite could instead define it as a graph, with culling nodes outputting work for more culling nodes, and have the driver handle the load balancing.