3 ms·
packet sojurn time is bounded by the latency of the GPU memory architecture. which as I understand has the design dial cranked to ten for parallelism and not so
by pyvpx 3y ago
packet sojurn time is bounded by the latency of the GPU memory architecture. which as I understand has the design dial cranked to ten for parallelism and not so much for expediency
- touisteur 3y agoPeople have been using GPU + DMA for low latency / real-time / high compute intensity applications for some time (using them for adaptive optics of all things). My PhD student's been cranking it to 100/200/400G, with 'just' DPDK, gpudev, and persistent cuda kernels.. Depends on the application, batching policy, compute intensity, etc. But you can put 8 NICs and 8 GPUs in one node (and have them communicate through nvlink, so huge intergpu bandwidth!) which I can't for CPUs. You can maybe also get some unobtainium A100X or CX7+H100 to skimp on PCIe if you're well funded...