5 ms·
The RAM bandwidth is so slow on this that you can barely train or do inference or do anything on it. I think the only use case they have in mind for this is fin
by ComplexSystems 1y ago
The RAM bandwidth is so slow on this that you can barely train or do inference or do anything on it. I think the only use case they have in mind for this is fine tuning pretrained models.
- wmf 1y agoIt's the same as Strix Halo and M4 Max that people are going gaga about, so either everyone is wrong or it's fine.
- gardnr 1y agoMemory Bandwidth: Nvidia DGX: 273 GB/s M4 Max: (up to) 546 GB/s M3 Ultra: 819 GB/s RTX 5090: ~1.8 TB/s RTX PRO 6000 Blackwell: ~1.8 TB/s
- 7thpower 1y agoThe other ones are not framed as an “AI Supercomputer on your desk”, but instead are framed as powerful computers that can also handle AI workloads.
- aurareturn 1y agoM4 max has more than double the bandwidth. Strix Halo has the same and I agree it’s overrated.
- Rohansi 1y agoI would expect/hope that DGX would be able to make better use of its bandwidth than the M4 Max. Will need to wait and see benchmarks.
- aurareturn 1y agoIt should. It has tensor cores which should drastically improve prompt processing. It should also be highly optimized for most AI apps.
- woooooo 1y agoMatrix vector multiplication for feed forward layers is most of the bandwidth as I understand things, there's not really a way to do it "better", its just a bunch of memory-bound dot products. (Posting this comment in hopes of being corrected and learning something).
- Rohansi 1y agoThe problem is different parts of the SoC (CPU, GPU, NPU) may not actually be able to consume all of the bandwidth available to the system as a whole. This is why you'd need to benchmark - different chips may be able to feed the cores better than others.
- woooooo 1y agoAh, yeah. I guess as we venture further into SoCs that will be more common, I was just thinking "it's whatever the memory controller can do".
- imtringued 1y agoTraining is performed in parallel with batching and is more flops heavy. I don't have an intuition on how memory bandwidth intensive updating the parameters is. It shouldn't be much worse than doing a single forward pass though.
- littlestymaar 1y agoSame as Strix Halo, which is 30% cheaper and readily available, yes. Hence the disappointment.