5 ms·
Please explain, why GPUs cannot take weights directly from main memory? Why keep weights in VRAM? Why cannot I buy 256 Mb GPU and use it with main RAM?
by codedokode 3y ago
Please explain, why GPUs cannot take weights directly from main memory? Why keep weights in VRAM? Why cannot I buy 256 Mb GPU and use it with main RAM?
- dragontamer 3y agoBecause GPUs only have VRAM on board. For GPUs to take data from CPU RAM, it must traverse PCIe which is on the order of ~100x slower than through the on board VRAM.
- haltist 3y agoIt's called direct memory access. The CPU doesn't have to do all I/O. Any subsystem can access RAM with DMA. The main issue is scheduling the loads, stores, and arithemtic operations to avoid stalling the pipeline.
- jamesfmilne 3y agoYes, NVIDIA GPUs can indeed read directly from main memory. It's called Heterogeneous Memory Management. https://developer.nvidia.com/blog/simplifying-gpu-application-development-with-heterogeneous-memory-management/ https://developer.nvidia.com/blog/simplifying-gpu-applicatio... This would allow you to mmap() the weights file into CPU memory, and pass that CPU pointer to the GPU and allow it to fault the data in on-demand. The GPU will have less bandwidth accessing this memory than if you copied it onto the GPU though, as it will need to take a trip over PCIe to read it. But it would eliminate the need to manually upload data to the GPU and the synchronisation involved with that.
- wongarsu 3y agoBecause VRAM is faster than main memory, and a lot faster than ferrying data over PCIe. Nvidia supports reading data directly from main memory. They introduced it around 2013 as Unified Memory, and it occasionally gets updates that make it more usable. It's probably faster than loading each layer into VRAM, computing it, discarding the weights and loading the next layer; but it's still a lot slower than having the weights in VRAM and often slower than just doing the work on the CPU.
- hmottestad 3y agoApple kinda does something like this with their GPU. They put it on the same chip as their CPU, added RAM very close to it with a lot of bandwidth and then use the RAM interchangeably between the CPU and GPU. Downside is that you can’t upgrade the RAM since it’s all stuck together in the same package as the CPU/GPU. It does allow you to get a CPU/GPU with 192GB of memory for pennies compared to something similar from Nvidia, but definitely not going to be as fast. I don’t think the GPU can address all the memory though, I know that the 128GB is limited to 96GB that can be shared with the GPU.
- hmottestad 3y agoYou can get two NVIDIA A40 cards with 48GB of memory each, second hand on eBay for 12 000 USD. You can connect them with NVLink to get 96GB effectively. Or you could get a Mac Studio for 4 800 USD.
- codedokode 3y agoAMD (and Intel) has CPUs with integrated GPUs that use system memory. Cannot one use them as a reasonably priced alternative? I think that buying 64 Gb of ordinary memory is much cheaper than buying Apple's proprietary memory.
- codedokode 3y agoNot only Apple does that. Intel and AMD have CPUs with integrated GPUs that do not have their own VRAM and use system memory. Why cannot they be used for LLMs? They don't cost an arm and a leg.
- jamesTee49 3y ago[dead]
- hmottestad 3y agoIntel Iris Xe can use half of the available memory. I don’t think it will help much though, the GPU is not very powerful and the bandwidth is around 90 GB/s. The M2 Ultra in the Mac Studio has 800 GB/s. The Nvidia A40 has 696 GB/s while the Nvidia H100 SMX has 3.35 TB/s, which is probably the best you could get today. https://www.intel.com/content/www/us/en/support/articles/000020962/graphics.html https://www.intel.com/content/www/us/en/support/articles/000...