3 ms·
I worked in HPC when NVIDIA started taking serious market share from Intel. My memory of Intel’s performance comparisons were that they were often technically u
by moconnor 8y ago
I worked in HPC when NVIDIA started taking serious market share from Intel. My memory of Intel’s performance comparisons were that they were often technically unsupportable once you scratched the surface.
In one case a third party who were demonstrating how much faster Intel Xeon Phi was for deep learning admitted that they were comparing highly-optimised code to unoptimised code in their results.
This doesn’t surprise me at all.
- m_mueller 8y agoI've been in the same boat and I completely agree. One thing that's unexpected to people is that getting decent performance out of a GPU is actually easier than CPUs - vectorization and multithreading is unified in the parallel programming model, cache optimizations are mostly not needed. These are the two biggest time sinks you have when optimizing for CPU, solved right there. What you instead have to care about is resource utilization per thread, and that is IMO way easier to reason about and optimize for.
- celrod 8y agoAre there any good guides or tutorials? I've found GPUs difficult, in part because I don't really know where to start. FWIW, I have an AMD GPU with ROCm. HIP it's a lot like CUDA, so NVidea-focused tutorials ought to be fine. With the caveat that I'd have to be aware of hardware differences.
- dragontamer 8y agoThe main issue IMO is thread-divergence. Because "threads" on a GPU are really SIMD-elements, things work very differently. Lets use a simple example: for(int i=0; i<10000000; i++){ if(someCondition) break; else doStuff(); } In a CPU case, the thread will break out of the loop early on "someCondition". But in the GPU case, it will only break out of the loop when "someCondition" holds for the entire SIMD-group. GPUs execute roughly 32-threads with the same instruction pointer. Lets say thread#0 had "someCondition" to be true. Then thread#0 will be set to "disabled", but otherwise, it will have to wait for the 31-other threads to be done with the loop before continuing. Even if 31-threads have hit "someCondition" and have broken out of the loop, the 32nd thread will keep executing the loop until it is done (and threads 0-through-30 will "execute with" the 32nd thread, but will throw away the results). That's the key with SIMD. Threads are run in groups of ~32ish at a time, at the same time. All 32-threads must execute if statements together and loops together. In most cases, an if/else statement will be executed by BOTH threads (but the results "thrown out" by the GPU engine, through execution masks)
- dragontamer 8y agoDepends on how many if-statements / branches your code takes. If you have simple if-statements and all branches are grouped together to be SIMD'd easily... then yeah. GPU threads kind of are like normal CPU threads. But as soon as you have a serious degree of thread-divergence, your performance tanks. That's why things like Chess engines (which ARE parallel problems at heart), execute poorly on GPUs. Because even though it is massively parallel, chess has too many if-statements and can't extend to SIMD very easily. -------- Raytracing algorithms are funny: they group rays together so that the GPU can SIMD over them more easily. But without the "re-grouping" step, its bad performance. Ex: A bunch of rays start at the camera. Some might hit a diffuse surface like wood... some might hit a subsurface-scattering surface like skin, and others might hit a metalic surface. GPU-raytracing algorithms then save off all the rays, and then processes all "diffuse" rays together, to minimize divergence. You can't just follow a singular ray in raytracing on a GPU. You gotta re-group to SIMD units for maximum performance.