8 ms·
Cerebras achieves 2,500T/s on Llama 4 Maverick (400B)
- thawab 1y agoare the Llama 4 issues fixed? what is it good at? coding is out of the window after the updated R1.
- NitpickLawyer 1y agoYes, the issues were fixed ~1-2 weeks after release. It's a good "all-rounder" model, best compared to 4o. Good multilingual capabilities, even in languages not specifically highlighted. Fast to run inference on it. Code is not one of its strong suits at all.
- y2244 1y agoInvestors list include Altman and Ilya https://www.cerebras.ai/company https://www.cerebras.ai/company
- ryao 1y agoTheir CEO is a felon who plead guilty to accounting fraud: https://milled.com/theinformation/cerebras-ceos-past-felony-conviction-hangs-over-ipo-listing-qolWXXFud-5GfTpy https://milled.com/theinformation/cerebras-ceos-past-felony-... Experienced investors will not touch them: https://www.nbclosangeles.com/news/business/money-report/cerebras-ipo-has-too-much-hair-as-ai-chipmaker-tries-to-sell-wall-street-on-nvidia-alternative/3533423/ https://www.nbclosangeles.com/news/business/money-report/cer... I estimated last year that they can only produce about 300 chips per year and that is unlikely to change because there are far bigger customers for TSMC that are ahead of them in priority for capacity. Their technology is interesting, but it is heavily reliant on SRAM and SRAM scaling is dead. Unless they get a foundry to stack layers for their wafer scale chips or design a round chip, they are unlikely to be able to improve their technology very much past the CSE-3. Compute might somewhat increase in the CSE-4 if there is one, but memory will not increase much if at all. I doubt the investors will see a return on investment.
- pinoy420 1y ago[dead]
- impossiblefork 1y agoWhile the CEO stuff is a problem, I don't think the other stuff matters. Per chip area WSE-3 is only a little bit more expensive than H200. While you may need several WSE-3s to load the model, if you have enough demand that you are running the WSE-3 at full speed you will not be using more area in the WSE-3. In fact, the WSE-3 may be more efficient, since it won't be loading and unloading things from large memories. The only effect is that the WSE-3s will have a minimum demand before they make sense, whereas an H200 will make sense even with little demand.
- ryao 1y agoI did the math last year to estimate how many wafers per year Nvidia had, and from my recollection it was >50,000. Cerebras with their ~300 per year is not able to handle the inference needs of the market. It does not help that all of their memory must be inside the wafer, which limits the amount of die area they have for actual logic. They have no prospect for growth unless TSMC decides to bless them or they switch to another foundation. > While you may need several WSE-3s to load the model, if you have enough demand that you are running the WSE-3 at full speed you will not be using more area in the WSE-3. You need ~20 wafers to run the Llama 4 Behemoth model on Cerebras hardware. This is close to a million mm^2. The Nvidia hardware that they used in their comparison should have less than 10,000 mm^2 die area, yet can run it fine thanks to the external DRAM. How is the CSE-3 not using more die area? > In fact, the WSE-3 may be more efficient, since it won't be loading and unloading things from large memories. This makes no sense to me. Inference software loads the model once and then uses it multiple times. This should be the same for both Nvidia and Cerebras.
- impossiblefork 1y agoYes, on an ordinary GPU it loads the weights to GPU memory, but then these weights must be moved from GPU memory onto the chip. But on these the weights can presumably be kept on chip entirely-- that's basically their whole point, so with the Cerebras there's no need to ever move weights to the chip. Of course these guys depend on getting chips, but so does everybody. I don't know how difficult it is, but all sorts of entities get TSMC 5nm. Maybe they'll get TSMC 3nm and 2nm later than NVIDIA, but it's also possible that they don't.
- MangoToupe 1y ago[flagged]
- ryao 1y ago> At over 2,500 t/s, Cerebras has set a world record for LLM inference speed on the 400B parameter Llama 4 Maverick model, the largest and most powerful in the Llama 4 family. This is incorrect. The unreleased Llama 4 Behemoth is the largest and most powerful in the Llama 4 family. As for the speed record, it seems important to keep it in context. That comparison is only for performance on 1 query, but it is well known that people run potentially hundreds of queries in parallel to get their money out of the hardware. If you aggregate the tokens per second across all simultaneous queries to get the total throughput for comparison, I wonder if it will still look so competitive in absolute performance. Also, Cerebras is the company that not only was saying that their hardware was not useful for inference until some time last year, but even partnered with Qualcomm with the claim that Qualcomm’s accelerators had a 10x price performance improvement over their things: https://www.cerebras.ai/press-release/cerebras-qualcomm-announce-10x-inference-performance https://www.cerebras.ai/press-release/cerebras-qualcomm-anno... Their hardware does inference with FP16, so they need ~20 of their CSE-3 chips to run this model. Each one costs ~$2 million, so that is $40 million. The DGX B200 that they used for their comparison costs ~$500,000: https://wccftech.com/nvidia-blackwell-dgx-b200-price-half-a-million-dollars-top-of-the-line-ai-hardware/ https://wccftech.com/nvidia-blackwell-dgx-b200-price-half-a-... You only need 1 DGX B200 to run Llama 4 Maverick. You could buy ~80 of them for the price it costs to buy enough Cerebras hardware to run Llama 4 Maverick. Their latencies are impressive, but beyond a certain point, throughput is what counts and they don’t really talk about their throughput numbers. I suspect the cost to performance ratio is terrible for throughput numbers. It certainly is terrible for latency numbers. That is what they are not telling people. Finally, I have trouble getting excited about Cerebras. SRAM scaling is dead, so short of figuring out how to 3D stack their wafer scale chips, during fabrication at TSMC, or designing round chips, they have a dead end product since it relies on using an entire wafer to be able to throw SRAM at problems. Nvidia, using DRAM, is far less reliant on SRAM and can use more silicon for compute, which is still shrinking.
- littlestymaar 1y ago> This is incorrect. The unreleased Llama 4 Behemoth is the largest and most powerful in the Llama 4 family. Emphasis mine. Behemoth may become the largest and most powerful llama model, but right now it's nothing but vaporware. Maverick is currently the largest and more powerful llama model today (and if I had to bet, my money would be on Meta discarding Llama4 Behemoth entirely it eventually without having released it, and moving on to the next version number).
- MangoToupe 1y ago[flagged]
- turblety 1y agoMaybe one day they’ll have an actual api that you can pay per token. Right now it’s the standard “talk to us” if you want to use it.
- twothreeone 1y agoHuh? Just make an account, get your API key, and try out the free tier.. works for me. https://cloud.cerebras.ai https://cloud.cerebras.ai
- iansinnott 1y agoAlthough not obvious, you _can_ pay them per token. You have to use OpenRouter or Huggingface as the inference API provider. https://cerebras-inference.help.usepylon.com/articles/1925546995-usage-based-billing-options https://cerebras-inference.help.usepylon.com/articles/192554...
- lordofgibbons 1y agoVery nice. Now for their next trick they should offer inference on actually useful models like DeepSeek R1 (not the distills).
- bob1029 1y agoI think it is too risky to build a company around the premise that someone won't soon solve the quadratic scaling issue. Especially, when that company involves creating ASICs. E.g.: https://arxiv.org/abs/2312.00752 https://arxiv.org/abs/2312.00752
- qeternity 1y agoAttention is not the primary inference bottleneck. For each token you have to load all of the weights (or activated weights) from memory. This is why Cerebras is fast: they have huge memory bandwidth.
- Havoc 1y agoYeah also strikes me as quite risky. Their gear seems very focused on llama family specifically. Just takes one breakthrough and it's all different. See the recent diffusion style LLMs for example
- tryauuum 1y agoyes, was not obvious it's not terabytes per second
- Alifatisk 1y agoIn the context of LLMs, the unit is token and to measure the output it's tokens per second (T/s)
- diggan 1y ago> The most important AI applications being deployed in enterprise today—agents, code generation, and complex reasoning—are bottlenecked by inference latency Is this really true today? I don't work in enterprise, so don't know how things look like, but I'm sure lots of people here do, and it feels unlikely that inference latency is the top bottleneck, even above humans or waiting for human input? Maybe I'm just using LLMs very differently from how they're deployed in a enterprise, but I'm by far the biggest bottleneck in my setup currently.
- baq 1y agoIt is if you want good results. I’ve been giving Gemini pro prompts for 200+ seconds multiple times per day this week and for such tasks I really like to make it double/triple check and sometimes give the results to Claude for review, too (and vice versa). Ideally I can just run the prompt 100x and have it pick the best solution later. That’s prohibitively expensive and a waste of time today.
- diggan 1y ago> That’s prohibitively expensive Assuming you experience is working within enterprise, you're then saying that cost is the biggest bottleneck currently? Also surprising to me that enterprises would use out-of-the-box models like that, I was expecting at least fine-tuned models be used most of the time, for very specific tasks/contexts, but maybe that's way optimistic.
- threeseed 1y agoCost is irrelevant when compared to the salaries of the people using them so they will do basic cost controls but nothing too onerous. And cost is never a reason to prevent solutions being built and deployed. And most enterprises aren't even doing anything advanced with AI. Just doing POCs with chat bots (again) which will likely fail (again). Or trying to do enterprise search engines which are pointless because most content is isolated per team. Or a few OCR projects which is pretty boring and underwhelming.
- 1y ago
- deleted 1y ago[deleted]
- bravesoul2 1y agoI tried some Llama 4s on Cerebras and they were hallucinating like they were on drugs. I gave it a URL to analyse a post for style and it made it all up and didn't look at the url (or realize that it hadn't looked at it).
- geor9e 1y agoI love Cerebrus. 10-100x faster than the other options. I really wish the other companies realized that some of us prefer our computer be instant. I use their API (with Qwen3 reasoning model) for ~99% of my questions, and the whole answer finishes in under 0.1 seconds. Keeps me in a flow state. Latency is jarring. Especially the 5-10 seconds most AIs take these days, where it's just enough to make switching tasks not worth it. You just have to sit there in statis. If I'm willing to accept any latency, might as well make it a couple minutes in the background, and use a full agent mode or deep research AI at that point. Otherwise I want instant.