4 ms·
Primarily concerned about the memory bandwidth for training. Though I think I've been able to max out my M2 when using the MacBook's integrated memory with MLX
by jbarrow 2y ago
Primarily concerned about the memory bandwidth for training.
Though I think I've been able to max out my M2 when using the MacBook's integrated memory with MLX, so maybe that won't be an issue.
- ryao 2y agoTraining is compute bound, not memory bandwidth bound. That is how Cerebras is able to do training with external DRAM that only has 150GB/sec memory bandwidth.
- jdietrich 2y agoThe architectures really aren't comparable. The Cerebras WSE has fairly low DRAM bandwidth, but it has a huge amount of on-die SRAM. https://www.hc34.hotchips.org/assets/program/conference/day2/Machine%20Learning/HC2022_Cerebras_Final_v02.pdf https://www.hc34.hotchips.org/assets/program/conference/day2...
- ryao 2y agoThey are training models that need terabytes of RAM with only 150GB/sec of memory bandwidth. That is compute bound. If you think it is memory bandwidth bound, please explain the algorithms and how they are memory bandwidth bound.