4 ms·
https://docs.nvidia.com/dgx-superpod/reference-architecture-scalable-infrastructure-h100/latest/network-fabrics.html https://docs.nvidia.com/dgx-superpod/refere
by rlupi 2y ago
https://docs.nvidia.com/dgx-superpod/reference-architecture-scalable-infrastructure-h100/latest/network-fabrics.html https://docs.nvidia.com/dgx-superpod/reference-architecture-...
NVIDIA large GPU supercomputers have separate compute-networking (between GPUs) and storage-networking (storage to GPUs, or storage to SSD, SSD to GPUs with CPU assistance). This helps avoid networking issues, even more if not using Infiniband.
From what I read here and on your website, you don't go that route.
I haven't found the equivalent system level reference architecture for MI300x from AMD. I wonder if you have a link to a public document where AMD provides guidance about this choice?
- derefr 2y ago> even more if not using Infiniband It's interesting that the above "HPC reference architecture" shows a GPU-to-GPU Infiniband fabric, despite Nvidia also nominally pushing NVLink Switch (https://www.nvidia.com/en-us/data-center/nvlink/ https://www.nvidia.com/en-us/data-center/nvlink/) for the HPC use-case.
- bee_rider 2y agoHow does NVLink work? Because I already know MPI and I’m not going to learn anything else, lol. Edit: after googling it looks like OpenMPI has some NVLink support, so maybe it is OK.
- zxexz 2y agoI use OpenMPI with no issues over multiple H100 nodes and A100 nodes, with multiple infiniband 200G and ethernet 100G/200G networks, and RDMA (though using mellanox instead of broadcom cards, but afaik broadcom supports this just the same). Side note, make sure you compile nvidia_peermem correctly if you want GDRMA to work :)
- latchkey 2y agoNo issues, except this minor bit of arcane knowledge that is missing from SO. :)
- kcb 2y agoThere is "CUDA-aware MPI" which would let you RDMA from device to device. But the more modern way would be MPI for the host communication and their own library NCCL for the device communication. NCCL has similar collective functions a MPI but runs on the device which makes it much more efficient to integrate in the flow of your kernels. But you would still generally bootstrap your processes and data through MPI.
- latchkey 2y agoWe have a separate OOB/east-west network which is 100G and would be used for external storage. We're spending an absurd amount of money on just cables. It is documented on the website [0], but I do see that I did not document the actual cards for that, will add when I wake up tomorrow. The card is: Broadcom 57504 Quad Port 10/25GbE,SFP28, OCP NIC 3.0 As far as I know, AMD doesn't really have the docs, it is Dell. Their team actively helped us design this whole cluster. We haven't decided on which type storage we want to get yet. It'll really depend on customer demand and since we haven't deployed quite yet, we are punting that can down the road a bit. Our boxes do all have 122TB in them and we have some additional servers not listed as well with 122TB... so for now I think we can cobble something useful together. [0] https://hotaisle.xyz/networking/ https://hotaisle.xyz/networking/
- latchkey 2y agowebsite updated with more details