5 ms·
Sounds like you might be more the target for the $3k 128GB DIGITS machine.
by iandanforth 2y ago
Sounds like you might be more the target for the $3k 128GB DIGITS machine.
- jbarrow 2y agoI’m really curious what training is going to be like on it, though. If it’s good, then absolutely! :) But it seems more aimed at inference from what I’ve read?
- bmenrigh 2y agoI was wondering the same thing. Training is much more memory-intensive so the usual low memory of consumer GPUs is a big issue. But with 128GB of unified memory the Digits machine seems promising. I bet there are some other limitations that make training not viable on it.
- tpm 2y agoIt will only have 1/40 performance of BH200, so really not enough for training.
- jbarrow 2y agoPrimarily concerned about the memory bandwidth for training. Though I think I've been able to max out my M2 when using the MacBook's integrated memory with MLX, so maybe that won't be an issue.
- ryao 2y agoTraining is compute bound, not memory bandwidth bound. That is how Cerebras is able to do training with external DRAM that only has 150GB/sec memory bandwidth.
- jdietrich 2y agoThe architectures really aren't comparable. The Cerebras WSE has fairly low DRAM bandwidth, but it has a huge amount of on-die SRAM. https://www.hc34.hotchips.org/assets/program/conference/day2/Machine%20Learning/HC2022_Cerebras_Final_v02.pdf https://www.hc34.hotchips.org/assets/program/conference/day2...
- ryao 2y agoThey are training models that need terabytes of RAM with only 150GB/sec of memory bandwidth. That is compute bound. If you think it is memory bandwidth bound, please explain the algorithms and how they are memory bandwidth bound.
- gpm 2y agoWeirdly they're advertising "1 petaflop of AI performance at FP4 precision" [1] when they're advertising the 5090 [2] as having 3352 "AI TOPS" (presumably equivalent to "3 petaflops at FP4 precision"). The closest graphics card they're selling is the 5070 with a GPU performing at 988 "AI TOPS" [2].... [1] https://nvidianews.nvidia.com/news/nvidia-puts-grace-blackwell-on-every-desk-and-at-every-ai-developers-fingertips https://nvidianews.nvidia.com/news/nvidia-puts-grace-blackwe... [2] https://www.nvidia.com/en-us/geforce/graphics-cards/50-series/ https://www.nvidia.com/en-us/geforce/graphics-cards/50-serie...
- MyFedora 2y agoIt seems like they're entirely different units? TOPS means Tera Operations Per Second. Petaflops means Peta Floating Point Operations Per Second. ... which, at least to my uneducated mind here, doesn't sound comparable.
- gpm 2y agoI'm assuming that the operations being referred to in both cases are fp4 floating point operations. Mostly because 1. That's used for AI, so it's plausibly what they mean by "AI OPS" 2. It's generally a safe bet that the marketing numbers NVIDIA gives you is going to be for the fastest operations on the computers, and that those are the same for both computers when they're based on the same architecture. Other than that, Terra is 10^12, Peta is 10^15, so 3352 Tera ops is 3.352 Peta ops and so on.