3 ms·
nah, it was designed for hpc and raw flops. llm inference really requires memory bandwidth.
by throwawaymaths 1y ago
nah, it was designed for hpc and raw flops. llm inference really requires memory bandwidth.
- HumanOstrich 1y agoMemory bandwidth, eh? You should learn the basics about what Cerebras does. https://www.cerebras.ai/chip https://www.cerebras.ai/chip
- adamtaylor_13 1y agoSheeeesh. 21 petabytes per second of memory bandwidth? That’s bonkers.
- cgdl 1y agoI'd say llm inference requires both memory capacity and bandwidth. Cerebras provides bandwidth with on-chip SRAM, but not capacity (an entire wafer has only 44GB SRAM).