3 ms·
> Only on smaller models, their numbers are all 70b in the article. No, they are 5x-10x faster for all the model sizes (because it's all just running from SRAM
by krasin 2y ago
> Only on smaller models, their numbers are all 70b in the article.
No, they are 5x-10x faster for all the model sizes (because it's all just running from SRAM and they have more of it than NVIDIA/AMD), even though they benchmarked just up to 70B.
> Those numbers also need to be adjusted for the comparable amounts of capex+opex costs. If the costs are so high that they have to subsidize the usage/results, then they are just going to run out of money, fast.
True. Although, for some workloads, fast enough inference is a strict prerequisite and GPUs just don't cut it.
- cma 2y agoCerebras had less chip perimeter to hook up external memory I/O and is memory capacity limited with just SRAM. SRAM circuit size hasn't been scaling nearly as well as logic on recent nodes, but if scaling there had continued to from when Cerebras started it may have worked out better. They'll probably still have to do advanced packaging putting HBM on top to save things. They could maybe enable some cool real time inference stuff like VR SORA, but that doesn't seem like much of a product market for the cost yet. Maybe something heavy on inference iteration like an o1 style model that trades training time for more inference, used to process earnings reports the fastest or something zero sum latency war like that will be a viable market. A real time use case that may be viable with cerebras first could be with flexible robotics in ad hoc latency sensitive environments, maybe warfare. If models keep lasting ~year timescales could we ever see people going with ROM chips for the weights instead of memory? Has density and speed kept up there? Lots of stuff uses identical elements to help make the masks more cheaply, so I don't think you could use something like EUV for ROM where every few um^2 of die is distinct.
- krasin 2y ago> If models keep lasting ~year timescales could we ever see people going with ROM chips for the weights instead of memory? Before ROM, there's a step where HBM for weights is replaced with Flash or Optane (but still high bandwidth, on top of the chip) and KV cache lives in SRAM - for small batch sizes, that would actually be decently cheap. In this case, even if weights change weekly, it's not a big deal at all.
- rbanffy 2y ago> They'll probably still have to do advanced packaging putting HBM on top to save things. This is where the interesting wafer-size packaging TSMC does for the Dojo D1 supercomputer comes in - Cerebras has demonstrated what can be a superior process for inter-element bandwidth, because connections can be denser than they are with an interposer, but the ability to have different elements coming from different processes is also important, and used on the D1 slab. Stacking HBM modules on top of a Cerebras wafer might help with that. I'm sure the smart people there are not sleeping on these ideas. For ultra low latency uses such as robotics or military applications, I believe a more integrated approach similar to the Telum processors from IBM is better - putting the inference accelerator on the same die as the CPUs gives them that, and they are also much smaller than a Cerebras wafer (and it's cooling). Gene Amdahl would have loved to see them.
- latchkey 2y ago576 CS-3 nodes costs around $900 million, which is $1.56 million per node. It takes 4 nodes to serve one 70B model or $6.24m. It is unclear how many requests they can serve concurrently for this, they only report on tokens. A cluster of 128 MI300x, which has a combined 24,576GB and can serve up a whole ton of models and users, 4 racks total, is in the ~$5m range, if you don't go big on networking/disk/ram (which you don't need for inference anyway). While speed might be an issue here, I don't think people are going to be able to justify the price for the speed (always the tradeoff) unless they can get their costs down significantly.
- modeless 2y agoYou are right assuming that model capabilities are determined only by model size. But consider that OpenAI is saying they have a way of scaling intelligence with inference time compute, not just model size. If that proves out, reducing latency per output token potentially becomes as valuable as or even possibly more valuable than scaling model size. Speed becomes intelligence. And Cerebras has 1/10 the latency per token of anything else.
- krasin 2y agoYou're correct on $/bandwidth. The point about low latency continues to be ignored, though.
- menaerus 2y agoIt's maybe because the assumption about low latency because everything fits in SRAM is not valid? CS-1 had 18G of SRAM, CS-2 extended it to 40G and CS-3 has 44G of SRAM. None of these is sufficient to run the inference of Llama 70B and much less so of even larger models.
- latchkey 2y agoExactly. Latency is less relevant if you have to have 4 literal servers (each taking up a whole rack) to push out one single 70B model and we don't know how many concurrent user requests that actually services (probably 1).