3 ms·
This article contains good tips for building a GPU cluster with RDMA. One thing I would like to add is that there are two types of GPUDirect depending on CUDA v
by nemonemo 12y ago
This article contains good tips for building a GPU cluster with RDMA. One thing I would like to add is that there are two types of GPUDirect depending on CUDA versions. Previous CUDA supported GPUDirect through CPU memory, and now CUDA supports "true" GPUDirect between the RDMA device and the GPU memory. However, some chipsets may not support the "true" GPUDirect very well, and two of our old machines had up to 20x times of throughput asymmetry with GPUDirect (which is, send was much slower than recv.) There are several papers that discuss this limitation. Our work, GPUnet[1], overcame this performance issue with GPUDirect by using fairly recent chipsets, but you can probably imagine our pain when we saw around 150MB/s throughput with GPUDirect, when 3GB/s is the expected one.
[1] GPUnet: Networking Abstractions for GPU Programs, OSDI 2014
https://sites.google.com/site/silbersteinmark/GPUnet https://sites.google.com/site/silbersteinmark/GPUnet