3 ms·
Cerebras's "wafer scale engine" [1] takes some of these ideas and applies them narrowly to deep learning training. 400,000 cores, 18GB of on-chip memory, 9.6 pe
by localhost 7y ago
Cerebras's "wafer scale engine" [1] takes some of these ideas and applies them narrowly to deep learning training. 400,000 cores, 18GB of on-chip memory, 9.6 petabytes of memory bandwidth in a 1.2 trillion transistor package in a gigantic 46,225mm^2 package.
[1] https://www.cerebras.net/cerebras-wafer-scale-engine-why-we-need-big-chips-for-deep-learning/ https://www.cerebras.net/cerebras-wafer-scale-engine-why-we-...