7 ms·
Not knowing much about special-purpose chips, I would like to understand whether chips like this would give Google a significant cost advantage over the likes o
by nehalem 1y ago
Not knowing much about special-purpose chips, I would like to understand whether chips like this would give Google a significant cost advantage over the likes of Anthropic or OpenAI when offering LLM services. Is similar technology available to Google's competitors?
- baby_souffle 1y agoThere are other ai/llm ‘specific’ chips out there, yes. But the thing about asics is that you need one for each *specific* task. Eventually we’ll hit an equilibrium but for now, the stuff that Cerebras is best at is not what TPUs are best at is not what GPUs are best at…
- monocasa 1y agoI don't even know if eventually we'll hit an equilibrium. The end of Moore's law pretty much dictates specialization, it's just more apparent in fields without as much ossification first.
- avrionov 1y agoNVIDIA operates at 70% profit right now. Not paying that premium and having alternative to NVIDIA is beneficial. We just don't know how much.
- kccqzy 1y agoI might be misremembering here, but Google's own AI models (Gemini) don't use NVIDIA hardware in any way, training or inference. Google bought a large number of NVIDIA hardware only for Google Cloud customers, not themselves.
- heymijo 1y agoGPUs, very good for pretraining. Inefficient for inference. Why? For each new word a transformer generates it has to move the entire set of model weights from memory to compute units. For a 70 billion parameter model with 16-bit weights that requires moving approximately 140 gigabytes of data to generate just a single word. GPUs have off-chip memory. That means a GPU has to push data across a chip - memory bridge for every single word it creates. This architectural choice, is an advantage for graphics processing where large amounts of data needs to be stored but not necessarily accessed as rapidly for every single computation. It's a liability in inference where quick and frequent data access is critical. Listening to Andrew Feldman of Cerebras [0] is what helped me grok the differences. Caveat, he is a founder/CEO of a company that sells hardware for AI inference, so the guy is talking his book. [0] https://www.youtube.com/watch?v=MW9vwF7TUI8&list=PLnJFlI3aINuUNaznonoFEgMUGLBiLmB-w&index=5 https://www.youtube.com/watch?v=MW9vwF7TUI8&list=PLnJFlI3aIN...
- hanska 1y agoThe Groq interview was good too. Seems that the thought process is that companies like Groq/Cerebras can run the inference, and companies like Nvidia can keep/focus on their highly lucrative pretraining business. https://www.youtube.com/watch?v=xBMRL_7msjY https://www.youtube.com/watch?v=xBMRL_7msjY
- latchkey 1y agoCerebras (and Groq) has the problem of using too much die for compute and not enough for memory. Their method of scaling is to fan out the compute across more physical space. This takes more dc space, power and cooling, which is a huge issue. Funny enough, when I talked to Cerebras at SC24, they told me their largest customers are for training, not inference. They just market it as an inference product, which is even more confusing to me. I wish I could say more about what AMD is doing in this space, but keep an eye on their MI4xx line.
- heymijo 1y ago> they told me their largest customers are for training, not inference That is curious. Things are moving so quickly right now. I typed out a few speculative sentences then went ahead and asked an LLM. Looks like Cerebras is responding to the market and pivoting towards a perceived strength of their product combined with the growth in inference, especially with the advent of reasoning models.
- latchkey 1y agoI wouldn't call it "pivoting" as much as "marketing".
- usatie 1y agoThank you for sharing this perspective — really insightful. I’ve been reading up on Groq’s architecture and was under the impression that their chips dedicate a significant portion of die area to on-chip SRAM (around 220MiB per chip, if I recall correctly), which struck me as quite generous compared to typical accelerators. From die shots and materials I’ve seen, it even looks like ~40% of the die might be allocated to memory [1]. Given that, I’m curious about your point on “not enough die for memory” — is it a matter of absolute capacity still being insufficient for current model sizes, or more about the area-bandwidth tradeoff being unbalanced for inference workloads? Or perhaps something else entirely? I’d love to understand this design tension more deeply, especially from someone with a high-level view of real-world deployments. Thanks again. [1] Think Fast: A Tensor Streaming Processor (TSP) for Accelerating Deep Learning Workloads — Fig. 5. Die photo of 14nm ASIC implementation of the Groq TSP. https://groq.com/wp-content/uploads/2024/02/2020-Isca.pdf https://groq.com/wp-content/uploads/2024/02/2020-Isca.pdf
- xnx 1y agoGoogle has a significant advantage over other hyperscalers because Google's AI data centers are much more compute cost efficient (capex and opex).
- claytonjy 1y agoBecause of the TPUs, or due to other factors? What even is an AI data center? are the GPU/TPU boxes in a different building than the others?
- xnx 1y ago> Because of the TPUs, or due to other factors? Google does many pieces of the data center better. Google TPUs use 3D torus networking and are liquid cooled. > What even is an AI data center? Being newer, AI installations have more variations/innovation than traditional data centers. Google's competitors have not yet adopted all of Google's advances. > are the GPU/TPU boxes in a different building than the others? Not that I've read. They are definitely bringing on new data centers, but I don't know if they are initially designed for pure-AI workloads.
- nsteel 1y agoWouldn't a 3d torus network have horrible performance with 9,216 nodes? And really horrible latency? I'd have assumed traditional spine-leaf would do better. But I must be wrong as they're claiming their latency is great here. Of course, they provide zero actual evidence of that. And I'll echo, what even is an AI data center, because we're still none the wiser.
- xnx 1y ago> what even is an AI data center A data center that runs significant AI training or inference loads. Non AI data centers are fairly commodity. Google's non-AI efficiency is not much better than Amazon or anyone else. Google is much more efficient at running AI workloads than anyone else.
- pkaye 1y agoAnthropic is using Google TPUs. Also jointly working with Amazon on a data center using Amazon's custom AI chips. Also Google and Amazon are both investors in Anthropic. https://www.datacenterknowledge.com/data-center-chips/ai-startup-anthropic-to-use-google-chips-in-expanded-partnership https://www.datacenterknowledge.com/data-center-chips/ai-sta... https://www.semafor.com/article/12/03/2024/amazon-announces-new-rainier-ai-compute-cluster-with-anthropic https://www.semafor.com/article/12/03/2024/amazon-announces-...
- cavisne 1y agoNvidia has ~60% margins in their datacenter chips. So TPU's have quite a bit of headroom to save google money without being as good as Nvidia GPU's. No one else has access to anything similar, Amazon is just starting to scale their Trainium chip.
- buildbot 1y agoMicrosoft has the MAIA 100 as well. No comment on their scale/plans though.
- deleted 1y ago[deleted]