4 ms·
"Last Fall, Nvidia released the Geforce GTX 970. It has 5.2 BILLION transistors on it. It already supports DirectX 12. Right now. It has thousands of cores in
by jra101 12y ago
"Last Fall, Nvidia released the Geforce GTX 970. It has 5.2 BILLION transistors on it. It already supports DirectX 12. Right now. It has thousands of cores in it. And with DirectX 11, I can talk to exactly 1 of them at a time."
That's not how it works, the app developer has no control over individual GPU cores (even in DX12). At the API level you can say "draw this triangle" and the GPU itself splits the work across multiple GPU cores.
- kllrnohj 12y agoMoreover the actual GPU doesn't work like that either. GPUs do not have the capability to run more than one work-unit-thing at a time. They have thousands of cores, yes, but much more in a SIMD-style fashion than in a bunch of parallel threads. They cannot split those cores up into logical chunks that can then individually do independent things. The whole post isn't just oversimplified, it's just wrong. Across the board wrong wrong wrong. The point of Mantle, of Metal, and of DX12 is to expose more of the low level guts. The key thing is that those low level guts aren't that low level. The threading improvements come because you can build the GPU objects on different threads, not because you can talk to a bunch of GPU cores from different threads. The majority of CPU time these days in OpenGL/DirectX is in validating and building state objects. DX12 and others now lets you take lifecycle control of those objects. Re-use them across frames, build them on multiple threads, etc... Then talking to the GPU is a simple matter of handing over an already-validated, immutable object to the GPU. Which is fast. Very fast.
- mattnewport 12y agoYeah, I'd have to agree it's hard to describe this post in more generous terms than just flat out wrong. DX12 is making it more efficient to spread CPU side rendering work across multiple cores but it's not about letting individual CPU cores talk to individual GPU cores. That isn't even really a coherent concept. The whole digression on lighting is mostly just wrong too. Deferred renderers have been rendering with 100s of dynamic lights for years. DX12 may make it a bit more efficient to deal with the large amount of constant data that needs to be updated when dealing with 100s of dynamic lights but it isn't introudcing any fundamental changes to dynamic lighting.
- jra101 12y agoMost GPUs group together individual ALUs into larger units (sometimes called compute units or clusters or SMs) and each compute unit can run independent work. http://www.anandtech.com/show/8526/nvidia-geforce-gtx-980-review/3 http://www.anandtech.com/show/8526/nvidia-geforce-gtx-980-re...
- deleted 12y ago[deleted]
- bhouston 12y agoI believe that is true but I haven't seen this exposed via DX. Does dx12 expose this functionality? Does cuda or OpenGL expose this type of xontrol.?
- jra101 12y agoNo, this is something the GPU front end controls.
- jeremiep 12y agoYou still have to be aware of it when optimizing the shaders and workloads though. On consoles where the hardware is fixed this is easily profiled. The GPU is creating threads and tasks internally and it's not always easy to balance this workload so no parts of the GPU becomes saturated while following parts in the chip's pipeline are idly waiting for work. The PowerVR chips we're working with have dozens and dozens of different profile metrics corresponding to the different areas of its pipeline, each one being a potential bottleneck. You could do something as silly as render a ball with 12k vertices instead of 24 and expecting the vertex processing to be much slower, but after profiling you find out its the fragment part lagging way behind because the data sequencer is overloaded trying to generate fragment tasks. In both cases you're rendering about the same amount of pixels. With unified shader architectures, its very frequent for vertex and fragment tasks from different draw calls to overlap simultaneously. We're even seeing tasks from different render targets overlapping! Such as fragment tasks from the shadow pass still running when the solid geometry pass is processing its vertices.
- Arwill 12y ago>build them on multiple threads That was the point of the blog post. Maybe described wrong, but that seems to be the point. The possibility of uploading stuff to the GPU on multiple threads on both CPU side, and the ability of the GPU to store those uploads in parallel. Maybe even render shadow-maps in parallel? I wonder is there already the OpenGL equivalent of this? One of the main hurdles of OGL was issuing all the calls from the main thread, if that is gone now, that would be awesome.
- bhouston 12y agoIt isn't possible to render shadow maps in parallel in dx12 I believe.
- kllrnohj 12y agoUploads are already done asynchronously, that's driver optimization 101 level stuff. It's also one of the very few operations that has an independent core to handle it (the copy engine) Rendering shadow-maps in parallel would be pointless. If you render 2 at the same time, then each map gets half the GPU so an individual render takes twice as long, resulting in the same total time as if you gave each render 100% of the GPU and rendered in sequence. > I wonder is there already the OpenGL equivalent of this? One of the main hurdles of OGL was issuing all the calls from the main thread, if that is gone now, that would be awesome. Yes, NV_command_list extension: http://www.slideshare.net/tlorach/opengl-nvidia-commandlistapproaching-zerodriveroverhead http://www.slideshare.net/tlorach/opengl-nvidia-commandlista...