3 ms·
> If RAM is the main bottleneck then CPUs should be on the table That's certainly not the case. The graphics memory model is very different from the CPU memory
by IX-103 2y ago
> If RAM is the main bottleneck then CPUs should be on the table
That's certainly not the case. The graphics memory model is very different from the CPU memory model. Graphics memory is explicitly designed for multiple simultaneous reads (spread across several different buses) at the cost of generality (only portions of memory may be available on each bus) and speed (the extra complexity means reads are slower). This makes then fast at doing simple operations on a large amount of data.
CPU memory only has one bus, so only a single read can happen at a time (a cache line read), but can happen relatively quickly. So CPUs are better for workloads with high memory locality and frequent reuse of memory locations (as is common in procedural programs).
- dragontamer 2y ago> CPU memory only has one bus If people are paying $15,000 or more per GPU, then I can choose $15,000 CPUs like EPYC that have 12-channels or dual-socket 24-channel RAM. Even desktop CPUs are dual-channel at a minimum, and arguably DDR5 is closer to 2 or 4 buses per channel. Now yes, GPU RAM can be faster, but guess what? https://www.tomshardware.com/pc-components/cpus/amd-crafts-custom-epyc-cpu-for-microsoft-azure-with-hbm3-memory-cpu-with-88-zen-4-cores-and-450gb-of-hbm3-may-be-repurposed-mi300c-four-chips-hit-7-tb-s https://www.tomshardware.com/pc-components/cpus/amd-crafts-c... GPUs are about extremely parallel performance, above and beyond what traditional single-threaded (or limited-SIMD) CPUs can do. But if you're waiting on RAM anyway?? Then the compute-method doesn't matter. Its all about RAM.
- ryao 2y agoWhere are these GPUs with multiple buses? I only know of GPUs with wide buses.