5 ms·
Most GPUs group together individual ALUs into larger units (sometimes called compute units or clusters or SMs) and each compute unit can run independent work.
by jra101 12y ago
Most GPUs group together individual ALUs into larger units (sometimes called compute units or clusters or SMs) and each compute unit can run independent work.
http://www.anandtech.com/show/8526/nvidia-geforce-gtx-980-review/3 http://www.anandtech.com/show/8526/nvidia-geforce-gtx-980-re...
- deleted 12y ago[deleted]
- bhouston 12y agoI believe that is true but I haven't seen this exposed via DX. Does dx12 expose this functionality? Does cuda or OpenGL expose this type of xontrol.?
- jra101 12y agoNo, this is something the GPU front end controls.
- jeremiep 12y agoYou still have to be aware of it when optimizing the shaders and workloads though. On consoles where the hardware is fixed this is easily profiled. The GPU is creating threads and tasks internally and it's not always easy to balance this workload so no parts of the GPU becomes saturated while following parts in the chip's pipeline are idly waiting for work. The PowerVR chips we're working with have dozens and dozens of different profile metrics corresponding to the different areas of its pipeline, each one being a potential bottleneck. You could do something as silly as render a ball with 12k vertices instead of 24 and expecting the vertex processing to be much slower, but after profiling you find out its the fragment part lagging way behind because the data sequencer is overloaded trying to generate fragment tasks. In both cases you're rendering about the same amount of pixels. With unified shader architectures, its very frequent for vertex and fragment tasks from different draw calls to overlap simultaneously. We're even seeing tasks from different render targets overlapping! Such as fragment tasks from the shadow pass still running when the solid geometry pass is processing its vertices.
- bhouston 12y agoThis is fascinating. I would love to know more. You should write an article about advanced optimization for mobile GPUs or something.
- m_mueller 12y agoYou can have multiple CUDA kernels running simultaneously on the same GPU, but you have no direct control over which SM(X) core is assigned to which kernel AFAIK. So it actually works pretty similarly to multi threaded programming on CPU if you take away thread affinity. In general I find a good way to approximate a top-of the line NVIDIA GPU is to think of it as an ~8 core with a vector length of 192 for single precision and 96 (half of full length) for double precision. It has a high memory bandwidth which has the limitation of requiring memory accesses to 32 neighbours simultaneously in order to make full use of the performance. The CUDA programming model is particularly set up in a way that the programmer doesn't have to handle this manually - (s)he just needs to be aware of it. I.e. you program everything scalar, map the thread indices to your data accesses and make sure that the first thread index (x) maps to the fastest varying index of your data.
- kllrnohj 12y ago> In general I find a good way to approximate a top-of the line NVIDIA GPU is to think of it as an ~8 core with a vector length of 192 for single precision and 96 (half of full length) for double precision. But that's not entirely correct either. Yes you can use it like that, but you can also use it as a single core with a vector length of 1536. In the context of over-simplification these are better thought of as single-core processors. The reason being if you have method foo() that you need to run 10,0000 times, it doesn't matter if you use 1 thread, 2 threads, or 8 threads - the total time it will take to complete the work will be identical. This is very different from an 8-core CPU where using 8 threads will be 8x faster than using 1 thread (blah blah won't be perfectly linear etc, etc).