5 ms·
I'm wondering since GPUs keep getting more performant, is there a middle ground between CPU and GPU that would still get performance increases? In my understan
by BenoitP 4y ago
I'm wondering since GPUs keep getting more performant, is there a middle ground between CPU and GPU that would still get performance increases?
In my understanding GPUs do mostly parallel worloads: SIMD with limited code size, operating on the same batch loaded arrays, but also with massive local memory.
And CPUs do mostly serial workloads: code size can be gigantic, lots of indirection, can pipeline a handful of operations (computations but several inflight small data loads), complicated cache herarchy; with massive silicon areas dedicated to shortening this serial critical path (branch prediction, instruction reordering, speculative execution)
Is there a new paradigm that we could see emerge and go mainstream in a few years? We do away with branch prediction, instruction reordering, speculative execution, and instead we can have massive amount of small dumb independent cores. These could do their share of small independent loads and would have small amounts of local working memory, could share some data with nearby cores. It would be common for these to halt for a long time waiting for a load from RAM, but that'd be ok.
Could we see that appear? Would it get the performance increases GPUs get? What language and coding style would be best suited for that silicon? What would be the market drivers (certainly not matrix multiplication/AI, nor single-threaded performance which are GPU and CPU's unfair advantage respectively)?
- dragontamer 4y ago> Is there a new paradigm that we could see emerge and go mainstream in a few years? We do away with branch prediction, instruction reordering, speculative execution, and instead we can have massive amount of small dumb independent cores. These could do their share of small independent loads and would have small amounts of local working memory, could share some data with nearby cores. It would be common for these to halt for a long time waiting for a load from RAM, but that'd be ok. You're gonna have to explain why a GPU is illsuited for this task. GPUs even have __shared__ memory space that passes data between local threads at extreme speeds (comparable to GPU L1 cache). > but also with massive local memory. On the contrary. CPU has more "local" memory if we're talking about L1, L2, and L3 caches. GPUs have massive register space, but very little high-speed memory (aka: cache). CPUs probably win if the data fits inside of say, 10MBs with hot-data within 512kB. GPUs win if you can get the entire computation inside of register space. But if you have lots of lookups, GPUs have very small caches, and the latency to GPU cache is an order of magnitude slower than CPU caches. GPU vRAM is just RAM. GPUs have much faster vRAM. GDDR6x has well over 500GBps bandwidth, while DDR4 and DDR5 will be in the 50GBps to 100GBps on typical machines, maybe 200GBps to 400GBps on servers.
- Veliladon 4y agoJust an FYI but GPUs have been beefing up L2 caches over generations. Ada has a fat unified 96MB L2 cache smack in the middle of the crossbar.
- dragontamer 4y agoI'd say the per-core L2 cache of 2MB is a bigger deal for computations than a last-level-cache of size 96MB on the GPU that's shared between all cores. The per-core (or SM / WGP) L1 or L0 cache of GPUs is quite anemic. The GPU designers clearly "intend" for the programmer to hold as much state in registers as possible, rather than in cache, for their computations. You really want to keep the state at 1024 bytes or less in practice for high speed GPU computations. CPUs with 2MB L2 cache per core (alder lake from Intel), or 1MB cache per core (Apple M2 and/or AMD Zen4), are simply designed to handle bigger "states per thread". Using the full 1MB to 2MBs for this L2 cache is reasonable.
- ls612 4y agoWhat sort of latency does the GPU have to VRAM compared to the latency of system memory to the CPU?
- dragontamer 4y agoThat's a difficult question to answer correctly. Not only because of the multitude of hardware differences between GPUs, but also because GPUs have a different strategy of latency-hiding than CPUs. ... And if you can successfully "hide the latency", then you kind of don't care about the latency... I'd estimate CPU RAM latency to be 50ns to 100ns and GPU vRAM latency to be 100ns to 500ns, depending on architecture. (There's many more GPU architectures out there, and they all have very different latency characteristics) But again, GPUs "don't care" about the latency to some extent due to their massive number of concurrent wavefronts/blocks. While CPUs also "don't care" because of L1/L2/L3 cache and out-of-order execution. Its probably more important to understand the different strategies of latency-hiding between CPU vs GPU, rather than knowing the actual latency figure.
- nynx 4y agoYou’re describing a GPU.
- jayd16 4y agoI don't think so. Not unless we start seeing workloads that are parallelizable enough to use all the CPU cores but not enough that it's worth it to farm the work off to the GPU. I suppose one possible gain could come from using integrated GPUs that can access CPU cache directly. Reducing the data upload to the GPU would really open up the more workloads to GPU type work.
- pjfin123 4y agoI'd like to see hardware more suited for running sparse ML models. You don't need branch prediction like on a CPU but you also don't need as much shared memory as a GPU.