4 ms·
I’m really curious what training is going to be like on it, though. If it’s good, then absolutely! :) But it seems more aimed at inference from what I’ve read?
by jbarrow 2y ago
I’m really curious what training is going to be like on it, though. If it’s good, then absolutely! :)
But it seems more aimed at inference from what I’ve read?
- bmenrigh 2y agoI was wondering the same thing. Training is much more memory-intensive so the usual low memory of consumer GPUs is a big issue. But with 128GB of unified memory the Digits machine seems promising. I bet there are some other limitations that make training not viable on it.
- tpm 2y agoIt will only have 1/40 performance of BH200, so really not enough for training.
- jbarrow 2y agoPrimarily concerned about the memory bandwidth for training. Though I think I've been able to max out my M2 when using the MacBook's integrated memory with MLX, so maybe that won't be an issue.
- ryao 2y agoTraining is compute bound, not memory bandwidth bound. That is how Cerebras is able to do training with external DRAM that only has 150GB/sec memory bandwidth.
- jdietrich 2y agoThe architectures really aren't comparable. The Cerebras WSE has fairly low DRAM bandwidth, but it has a huge amount of on-die SRAM. https://www.hc34.hotchips.org/assets/program/conference/day2/Machine%20Learning/HC2022_Cerebras_Final_v02.pdf https://www.hc34.hotchips.org/assets/program/conference/day2...
- ryao 2y agoThey are training models that need terabytes of RAM with only 150GB/sec of memory bandwidth. That is compute bound. If you think it is memory bandwidth bound, please explain the algorithms and how they are memory bandwidth bound.