4 ms·
In the absence of hardware unified memory, CUDA will automatically copy data between CPU/GPU when there are page faults.
by zcbenz 1y ago
In the absence of hardware unified memory, CUDA will automatically copy data between CPU/GPU when there are page faults.
- fenced_load 1y agoThere is also NVLink c2c support between Nvidia's CPUs and GPUs that doesn't require any copy, CPUs and GPUs directly access each other's memory over a coherent bus. IIRC, they have 4 CPU + 4 GPU servers already available.
- benreesman 1y agoYeah NCCL is a whole world and it's not even the only thing involved, but IIRC that's the difference between 8xH100 PCI and 8xH100 SXM2.
- nickysielicki 1y agoSee also: https://www.kernel.org/doc/html/v5.0/vm/hmm.html https://www.kernel.org/doc/html/v5.0/vm/hmm.html
- saagarjha 1y agoThis seems like it would be slow…
- freeone3000 1y agoMatches my experience. It’s memory stalls all over the place, aggravated (on 12.3 at least) there wasn’t even a prefetcher.
- deleted 1y ago[deleted]