11 ms·
Cerebras systems raises $1.1B Series G
- fcpguru 1y agoTheir core product is the Wafer Scale Engine (WSE-3) — the largest single chip ever made for AI, designed to train and run models much faster and more efficiently than traditional GPUs. Just tried https://cloud.cerebras.ai https://cloud.cerebras.ai wow is it fast!
- OGEnthusiast 1y agoI'm surprised how under-the-radar Cerebras is. Being able to get near-instantaneous responses from Qwen3 and gpt-oss is pretty incredible.
- data-ottawa 1y agoI wish I could invest in them. Agree they're under the radar.
- redwood 1y agoWould be interesting if IBM were to acquire. Seems like the big iron approach to GPUs
- maz1b 1y agoCerebras has been a true revelation when it comes to inference. I have a lot of respect for their founder, team, innovation, and technology. The colossal size of the WS3 chip, utilizing DRAM to mind-boggling scale, it's definitely ultra cool stuff. I also wonder why they have not been acquired yet. Or is it intentional? I will say, their pricing and deployment strategy is a bit murky and unclear. Paying $1500-$10,000 per month plus usage costs? I'm assuming that it has to do with chasing and optimizing for higher value contracts and deeper-pocketed customers, hence the minimum monthly spend that they require. I'm not claiming to be an expert, but as a CEO/CTO, there were other providers in the market that had relatively comparable inference speed (obviously Cerebras is #1), easier onboarding, better response from people that worked there (all of my experience with Cerebras have been days/weeks late or simply ignored). IMHO, if Cerebras wants to gain more mindshare, they'll have to look into this aspect.
- oceanplexian 1y agoI’ve been using them as a customer and have been fairly impressed. The thing is, a lot of inference providers might seem better on paper but it turns out they’re not. Recently there was a fiasco I saw posted on r/localllama where many of the OpenRouter providers were degraded on benchmarks compared to base models, implying they are serving up quantized models to save costs, but lying to customers about it. Unless you’re actually auditing the tokens you’re purchasing you may not be getting what you’re paying for even if the T/s and $/token seems better.
- dlojudice 1y agoOpenRouter should be responsible for this quality control, right? It seems to me to be the right player in the chain with the duties and scale to do so.
- teruakohatu 1y ago> many of the OpenRouter providers were degraded on benchmarks compared to base models, implying they are serving up quantized models to save costs, Do you have information on this? This seems like brand destroying for both OpenRouter and the model providers.
- liuliu 1y agoThey were acquisition target since 2017 (from the OpenAI internal emails). So lacking of acquisition is not because lacking of interests. Let you wonder what happened in these due-diligence.
- OkayPhysicist 1y agoThe UAE has sunk a lot of money into them, and I suspect it's not purely a financial move. If that's the case, an acquisition might be more complicated than it would seem at first glance.
- throw123890423 1y ago> I will say, their pricing and deployment strategy is a bit murky and unclear. Paying $1500-$10,000 per month plus usage costs? I'm assuming that it has to do with chasing and optimizing for higher value contracts and deeper-pocketed customers, hence the minimum monthly spend that they require. Yeah wait, why rent chips instead of sell them? Why wouldn't customers want to invest money in competition for cheaper inference hardware? It's not like Nvidia has a blacklist of companies that have bought chips from competitors, or anything. Now that would be crazy! That sure would make this market tough to compete in, wouldn't it. I'm so glad Nvidia is definitely not pressuring companies to not buy from competitors or anything.
- ramshanker 1y agoI am not able to guess, what is preventing Cerebras from replacing few of the cores in the Wafer-Scale package with HBM memory? It seems the only constraint with their WSE3 is memory capacity. Considering the size of NVDA chips, Only a small subset of wafer area should easily exceed the memory size of contemporary models.
- reliabilityguy 1y agoDRAMs (core of the HBM memories) use different technology nodes than logic and SRAM. Also, stacking that many DRAMs on waver will complicate the packaging quite a bit I think.
- xadhominemx 1y agoI don’t think so. The reason why Cerebras is so fast for inference is that the KV cache sits in the SRAM.
- aurareturn 1y agoIf you replace some cores with HBM on package, you basically get the traditional GPU + HBM model.
- Shakahs 1y agoSonnet/Claude Code may technically be "smarter", but Qwen3-Coder on Cerebras is often more productive for me because it's just so incredibly fast. Even if it takes more LLM calls to complete a task, those calls are all happening in a fraction of the time.
- nerpderp82 1y agoWe must have very different workflows, I am curious about yours. What tools are you using and how are you guiding Qwen3-Coder? When I am using Claude Code, it often works for 10+ minutes at a time, so I am not aware of inference speed.
- CaptainOfCoit 1y ago> When I am using Claude Code, it often works for 10+ minutes at a time, so I am not aware of inference speed. Indirectly, it sounds like you're aware about the inference speed? Imagine if it took 2 minutes instead of 10 minutes, that's what the parent means.
- yodon 1y ago2 minutes is the worst delay. With 10 minutes, I can and do context switch to something else and use the time productively. With 2 min, I wait and get frustrated and bored.
- dataangel 1y agoContext switching makes you less productive compared to if you could completely finish one task before moving to the other though. in the limit an LLM that responds instantly is still better.
- solarkraft 1y agoYou must write very elaborate prompts for 10 minutes to be worth the wait. What permissions are you giving it and how much do you care about the generated code? How much time did you spend on initial setup? I‘ve found that the best way for myself to do LLM assisted coding at this point in time is in a somewhat tight feedback loop. I find myself wanting to refine the code and architectural approaches a fair amount as I see them coming in and latency matters a lot to me here.
- mythz 1y agoRunning Qwen3 coder at speed is great, but would also prefer to have access to other leading OSS models like GLM 4.6, Kimi K2 and DeepSeek v3.2 before considering switching subs. Groq also runs OSS models at speed which is my preferred way to access Kimi K2 on their free quotas.
- JLO64 1y agoMy experience with Cerebras is pretty mixed. On the one hand for simple and basic requests, it truly is mind blowing how fast it is. That said, I’ve had nothing but issues and empty responses whenever I try to use them for coding tasks (Opencode via Openrouter, GPT-OSS). It’s gotten to a point where I’ve disabled them as a provider on Openrouter.
- divmain 1y agoI experienced the same, but I think it is a limitation of OpenRouter. When I hit Cerebra’s OpenAI endpoint directly, it works flawlessly.
- allisdust 1y agoIf the idiots at AMZN have any brains left, they would acquire this and make it the center of their inference offerings. But considering how lackluster their performance and strategy as a company has been off late, I doubt that. Disappointed quite a bit with this fund raise. They were expected to IPO this year and give us poor retail investors a chance at investing in them.
- reliabilityguy 1y agoAmazon has their own chips for inference and training: Trainium1/2.
- allisdust 1y agoNothing (may be except groq ?) comes even close to Cerebras in inference speed. I seriously don't get why these guys aren't more popular. The difference in using them as a inference provider vs anything else for any use case is like night and day. I hope more inference providers focus on speed. And this is where AMZN will benefit a lot since their entire cloud model is to have something people would anyway want and mark it up by 3x. God forbid if AVGO acquires this.
- xadhominemx 1y agoCerebras hasn’t made any technical breakthroughs, they are just putting everything in SRAM. It’s a brute force approach to get very high inference throughput but comes at extremely high cost per token per second and is not useful for batched inferencing. Groq uses the same approach. Memory hierarchy management across HBM/DDR/Flash is much more difficult but necessary to achieve practical inference economics.
- twothreeone 1y agoI don't think you realize the history of wafer-scale integration and what it means for the chip industry [1]. The approach was famously taken by Gene Amdahl's Trilogy Systems in the 80ies, but failed dramatically leading to (among others) deployment of "accelerator cards" in the form of.. the NVIDIA GeForce 256, the first GPU in 1999. It's not like NVIDIA hasn't been trying to integrate multiple dies in the same package, but doing that successfully has been a huge technological hurdle so far. [1] https://ieeexplore.ieee.org/abstract/document/9623424 https://ieeexplore.ieee.org/abstract/document/9623424
- rvz 1y agoSooner or later, lots of competitors including Cerebras are going to take apart Nvidia's data center market share and it will cause many AI model firms to question the unnecessary spend and hoarding of GPUs. OpenAI is still developing their own chips with Broadcom, but they are not operational yet. So for now, they're buying GPUs from Nvidia to build up their own revenue income (to later spend it on their own chips) By 2030, eventually many companies will be looking for alternatives to Nvidia like Cerebras or Lightmatter for both training and inference use-cases. For example [0] Meta just acquired a chip startup for this exact reason - "An alternative to training AI systems" and "to cut infrastructure costs linked to its spending on advanced AI tools.". [0] https://www.reuters.com/business/meta-buy-chip-startup-rivos-ai-effort-source-says-2025-09-30/ https://www.reuters.com/business/meta-buy-chip-startup-rivos...
- onlyrealcuzzo 1y agoThere's so much optimization to be made when developing the model and the hardware it runs on, most of the big players are likely to run a non-trivial percentage of their workloads on proprietary chips eventually. If that's 5 years into the future, that looks bad for Nvidia, if it's >10 years in the future, that doesn't affect Nvidia's current stock price very much.
- arjie 1y agoI just tried out Qwen-3-480B-Coder on them yesterday and to be honest it's not good enough. It's very fast but has trouble on lots of tasks that Claude Code just solves. Perhaps part of it is that I'm using Charm's Crush instead of Claude Code.
- arisAlexis 1y agoThey make chips. Potentially you could also test Claude on them in the future.
- tibbydudeza 1y agoDamm they are fast.
- dgfitz 1y agoValued at 8.1 billion dollars. https://www.cerebras.ai/pricing https://www.cerebras.ai/pricing $50/month for one person for code (daily token limit), or pay per token, or $1500/month for small teams, or an enterprise agreement (contact for pricing). Seems high.
- arisAlexis 1y agoThe valuation is for the best inference chips. You know Nvidia cloud pricing is irrelevant and so it's here too.
- dgfitz 1y agoSo their angle is an exit?
- lvl155 1y agoLast I tried, their service was spotty and unreliable. I would wait maybe a year or so to retry.
- fcpguru 1y agodoes Guillaume Verdon from https://www.extropic.ai/ https://www.extropic.ai/ have thoughts on on cerebras? (or other people that read the litepaper https://www.extropic.ai/future https://www.extropic.ai/future)
- landl0rd 1y agoBeff has shipped zero chips and shitposted a lot. It is a cool idea but he has made tons of promises and it's starting to seem more like vaporware. Don't get me wrong, I hope it works, but doubt it will. Less podcasts more building please. He reads to me like someone who markets better than he does things. I am disinclined to take him as an authority in this space. How do you believe this is related to Cerebras?
- fcpguru 1y agothat's the first thing I thought of when I read "cerebras faster chip for ai". Beff "sold me" on it a year ago. I guess I drank the kool-aid. Thinking about un-drinking now...
- rbitar 1y agoCongrats to the team, I'm surprised the industry hasn't been as impressed with their benchmarks on token throughput. We're using the Qwen 3 Coder 480b model and seeing ~2000 tokens/second, which is easily 10-20x faster then most LLM models on the market. Even some of the fastest models still only achieve 100-150 tokens / second (see OpenRouter stats by provider). I do feel after around 300-400 tokens/second the gains in speed feel more incremental, so if there was a model at 300+ tokens/second, I would consider that a very competitive alternative.
- darkbatman 1y agoIts so useful to use Cerebras api for other tasks too not just coding with qwen coder but even simpler things like lets say analysing with gpt-120 oss or llama. Just plug it in with normal chat interface like Jan or Cherry studio and its incredibly fast.