4 ms·
How GPU Computing Works [video]
- deleted 4y ago[deleted]
- einpoklum 4y agoThis seems like this year's version of the talk given last year, which was just recently posted here on HN as "How CUDA Programming works": https://news.ycombinator.com/item?id=31983460 https://news.ycombinator.com/item?id=31983460
- wrs 4y agoThe Wheel of Reincarnation continues. [0] (Though it’s sort of turning the other way, this time around?) [0] http://www.catb.org/jargon/html/W/wheel-of-reincarnation.html http://www.catb.org/jargon/html/W/wheel-of-reincarnation.htm...
- dragontamer 4y agoThe opposite. GPUs became more general purpose. Old, vector processors from the 1980s served as inspiration. Even in 90s commercials its obvious that the GPU / SIMD-compute similarities were all over the 3dfx cards. In the 00s, GPUs became flexible enough to execute arbitrary code for vector-effects and pixel-effects. Truly general purpose, arbitrary code, albeit in the SIMD methodology. -------- Today, your Direct2D windowing code in Windows is largely running on the GPU, and has been migrated away from the CPU. In fact, your video decoders (Youtube) and video games (shaders) are all GPU code. GPUs have proven themselves to be a general purpose processor, albeit with a strange SIMD-model of compute rather than the traditional Von Neumann design. We're in a cycle where more-and-more code is moving away from CPUs into GPUs, and permanently staying in GPU space. This is the opposite effect of the cycle of reincarnation (CPUs may have gotten faster, but GPUs have become not only faster at a higher rate, but also more general purpose and generic allowing for more general code to be run on them). Code successfully ported over (ex: Tensorflow), may never return back to CPU-side. SIMD-compute is just superior underlying model for a large set of applications.
- nayuki 4y ago> In fact, your video decoders (Youtube) [...] are all GPU code. I believe this is false; video decoders like H.264/AVC have a significant ASIC component that cannot be expressed as general-purpose SIMT code. I think this is because the entropy coding portion (arithmetic coding, Huffman coding, etc.) needs to be decoded serially. Some stuff like macroblock-to-macroblock prediction in I-frames is serial as well. But IDCT is indeed parallelizable.
- boberoni 4y ago> (Almost) Nobody (really) cares about flops ...because we should really be caring about memory bandwidth In university, I was shocked to learn in a database class that CPU costs are dwarfed by the I/O costs in the memory hierarchy. This was after spending a whole year on data structures and algorithms, where we obsessed over runtime complexity and # of operations. It seems that the low-hanging fruit of optimization is all gone. New innovations for performance will have to happen in transporting data.
- malnourish 4y agoAt the risk of being flippant, I hope you learned about space complexity and the lessons behind how the algorithms and data structures you use impact performance via the cache.
- vladf 4y agoCompression converts I/O bottlenecks to compute ones again.
- Sherl 4y agoThere is an entire field of parallel algorithms which makes use of sequential algorithms to overcome some of these issues. So no it's not wasted. You would apply the knowledge to build parallel algorithms. There are projects in my school where we implemented a combination of CUDA and openMP in some and MPI+OpenMP in a few. I think the bottleneck is always gonna be there, its just how much and how you deal with it in hardware from the software front.
- bogomipz 4y agoI don't really understand your comment. Databases are generally I/O bound workloads, almost by definition of what they do. Regardless, data structures and algorithms are equally important in databases. B-Trees, linked-lists, buffers, LSM Trees, Bloom filters, caching strategies etc are all fundamental to databases. At any rate for a long time now the low hanging fruit of optimization has been throwing money and hardware at problems - NAND Flash, more cores, larger caches, tons of memory, edge networks etc. Those options are all still there.
- Lichtso 4y agoIf I understand correctly: CPUs do minimize latency by: - Register renaming - Out of order execution - Branch prediction - Speculative execution They should not be over subscribed as they have to context switch by storing / loading registers and the cache coherence protocols scale badly with more threads. GPUs on the other hand maximize throughput by: - A lot more memory bandwidth - Smaller and slower cores, but more of them - Ultra threading (the massively over subscribed hyper threading the video mentions) - Context switching between wavefronts (basically the equivalent of a CPU thread), just shifts the offset into the huge register file (no store and load) The one area in which CPUs are getting closer to GPUs is SIMD / SIMT. CPUs used to be able to apply one instruction to a vector of elements without masking (SIMD). In ARM SVE and x86 AVX-512 they can now (like GPUs) mask out individual lanes (SIMT) for ALU operations and memory operations (gather load / scatter store).
- oddity 4y agoThe difference is much more nuanced than this. A modern GPU can (and probably does) do most of what you've listed for a CPU. Speculative execution and branch prediction are a bit less likely to be invested in (because they don't need it as much due to oversubscription), but that's increasingly true for CPUs as well for high-efficiency cores. The difference (at a category vs category level and not specific microarch) is mostly a matter of tuning for particular workloads. I'm increasingly souring on SIMD/SIMT being a useful distinction now that bleeding-edge CPUs are widening in the microarch and bleeding-edge GPUs are getting better at handling thread divergence in the microarch. There is a difference, certainly, but it's difficult to describe in a few bullet points. GPUs are more likely to have more exotic features than you'll see on a CPU to deal with things like thread coordination and cache coherence, but there's nothing fundamentally stopping CPUs from adding that (or wanting that) as well.
- Lichtso 4y ago> GPUs are getting better at handling thread divergence in the microarch That is an interesting point, how does that work (especially with the dynamics of ray tracing)? Do they recombine under utilized wavefronts or something?
- oifjsidjf 4y agoHere is another interesting series of articles which describes in more details how GPUs draw: https://fgiesen.wordpress.com/2011/07/09/a-trip-through-the-graphics-pipeline-2011-index/ https://fgiesen.wordpress.com/2011/07/09/a-trip-through-the-...
- deleted 4y ago[deleted]