13 ms·
Cerebras CS-4
- 9cb14c1ec0 1mo agoJust a reminder for everyone that we are only several years and 3 or 4 iterations into hardware being optimized for LLMs. We should all expect orders of magnitude improvement in speed and/or cost over the next 5 years. Then we can have fun conversations about "unlimited" "intelligence" and about what the price wars and profit margins of consumer AI products are when your average ChatGPT user costs the company $0.10 per month. > CS-4 delivers more than 1,000 tokens per second on models exceeding 10 trillion parameters Wow!
- api 1mo agoThis is part of why I think the data center build-out is a bubble. We've barely scratched the surface when it comes to hardware optimization. We'll see exponential improvements in energy efficiency and speed over the next decade. Exponential, not linear. GPUs really aren't that great for AI. They just happen to be the best chips we have in mass production right now for this work load, and it takes time to field new designs. Basically every chip engineer on the planet is working on this right now.
- jeffybefffy519 1mo agoExactly right, and nVidia is protecting their moat through business practices rather than genuine product innovation.
- winrid 1mo agoOn the plus side, lots of cheap servers to swoop up :)
- dgellow 1mo agoLook at their power supply, it’s not something you can run in a home lab. Unfortunately most of that will likely go to the bin eventually :(
- winrid 1mo agoNo but dedicated server prices will likely drop or at least you'll get more for your money
- sroussey 1mo agoBut power hungry. In that 5+ year timeline, the compute per watt could change by three orders of magnitude. GPUs are to LLMs what CPUs are to gaming — not a good fit.
- amluto 1mo agoA cursory estimate courtesy of ChatGPT suggests that there is a grand total of one order of magnitude or less of power efficiency improvement available compared to current Blackwell if the entire system’s power consumption outside the ALUs went all the way to zero. If you want three orders of magnitude improvement, you probably need to find two of those orders of magnitude somewhere else: process improvements, different ALU design, model architecture changes, etc.
- sroussey 1mo agoOh, you could go analog rather than digital. Ever look at ALU design? Nothing efficient about it!
- mindwok 1mo agoWhether it's a bubble or not depends on how much the demand for compute and the type of workload keeps growing, though. If AI tends to be something used mainly in ideation and development, which is how a lot of people use it today, then once consumer hardware gets good enough you could see a bunch of the current data centre workloads move onto consumer devices. But if AI starts being used more in repeatable, operational workloads I think it makes sense to have significant cloud infrastructure for it. TBH I haven't seen much of this, and I've been skeptical about people using agents for much of anything when it can be done with just software. But we are starting to see more of this kind of workload, like the taggable Claude in your slack etc that people seem to really love.
- petra 1mo agoI wonder: in world where inference is cheap, how many engineering agents that use simulation as their feedback we will use? In the scenario, engineering everything becomes so easy - so why not optimize everything? every component, every product, every system? And maybe llm's could invent. So even more to simulate. And simulation is inherently compute-heavy. So unless there are some other bottlenecks, we'll use a lot of simulation servers.
- RachelF 1mo agoTrue, I have to agree with you. The AI giants might be investing a huge amount of money in generation 1 technology. There might be a much better way to do it just around the corner. They might know this and thus the hurry to IPO. A rough analogy would be if the first generation of ISP's spent billions on dial-up exchanges, when fibre could be invented next year.
- __turbobrew__ 1mo agoBy the time these gigawatt datacenters are done being built the hardware will be so far behind state of the art they may be mostly useless.
- deleted 1mo ago[deleted]
- aurareturn 1mo agoBy the way, this is the same argument that Michael Burry used to short Nvidia. He claims that GPU depreciation/obsoletion is much faster than hyperscalers are assuming because new chips will be much better. He's being proved wrong right now because H200 rental prices have been claiming for the last 8 month despite B200 having 10-20x better inference efficiency.[0] The logic is fundamentally flawed in my opinion. Let's use future Nvidia chips being much better optimized for LLMs for example. New Nvidia chips 10x better than H200 --> data centers buy a lot --> Nvidia profits a lot. New Nvidia chips 10x better than H200 --> data centers don't buy --> no faster than expected obsoletion. In other words, the very act of buying many new Nvidia GPUs would be the event that causes faster than expected obsoletion. Yet, if you don't buy those new Nvidia GPUs, then there is no faster than expected obsoletion. We also live in a world where there is competition. If Amazon doesn't buy but Microsoft does, suddenly Microsoft can offer better $/token prices. [0]https://inferencex.semianalysis.com/inference https://inferencex.semianalysis.com/inference
- haldujai 1mo ago1. The same isn’t necessarily true of the rest of the hardware stack which may be reused between accelerator generations. 2. You’re missing the “New Nvidia chips 10x B200, compute requirement grows less than 10*software improvements YoY -> buy less Nvidia.” Valuations are based on forward projections (>1T annual for NVDA) which can be revised down leading to a drop in valuation. > If Amazon doesn't buy but Microsoft does The big 3 all have their own proprietary accelerators. Meta is buying TPUs as well for now. I would bet Nvidia’s major customers in 2 years are neoclouds and it seems that Jensen is making the same bet.
- aurareturn 1mo ago1. So this makes Burry’s argument even less convincing since those auxiliary hardware can last longer. 2. Jevons Paradox. More efficiency should lead to bigger models, faster inference, and more total tokens. 3. By all accounts, Trainium and Maia and Meta’s internal chip are struggling to keep up with Nvidia. That’s why they order as many Nvidia chips as possible. They’re not giving up but it isn’t as easy as buying stock Arm cores and taking them to TSMC. Neoclouds may very well be Nvidia’s biggest customers and this probably what Nvidia wants.
- trympet 1mo ago> We'll see exponential improvements in energy efficiency and speed over the next decade. Exponential, not linear. Why? Make your case.
- SwellJoe 1mo agoAnd, the software side isn't finished being optimized, either. We've seen with Qwen 3.8 27B and DeepSeek V4 Flash 0731 and GLM 5.3 that quite small models can pack a punch. Intelligence density will improve, efficiency of kernels will improve, efficiency of KV caching and MTP will improve, algorithms for splitting workloads across compute units will improve. It'll all be as cheap as DeepSeek was before the price hike. And, it'll become more and more realistic to run near-frontier intelligence on personal devices.
- rvz 1mo agoCongratulations! You have just realized that the AI data center build out is a total scam, built on both the insurmountable trillions of debt, and the assumption that only GPUs are all we need to continue scaling. There exist other AI accelerators (TPUs, ASICs) that perfectly exceed the throughput that LLMs need to scale as well. But the true solution is more software optimizations. There's a tiny handful of them but more needs to be discovered so that we can reduce building hundreds of more data centers as the alternatives mature. As better software becomes more useful for the alternative AI hardware for developers with LLMs running efficiently you then would have more choices of hardware to run your LLMs on rather than just only GPUs.
- blovescoffee 1mo agoTPUs and ASICs run in data centers too. Your argument only holds true if there's some satisfied limit to demand for inference. If not, data centers will continue to spring up to host more and more agents. Even if agents were running on hardware and software as efficient as the human brain, its conceivable we want trillions of them running at any given time which would require data center scale.
- georgeecollins 1mo agoEverything has some satisfied limit to demand, often depending on the price. If you assume there will never be any satisfied limit to demand for inference at any price you can justify any investment.
- aldonius 1mo agoYeah, but there's certainly a part of the curve where price drops by X OOMs and demand increases by much more than X OOMs. (Presumably some of that is substitution and some of that is new use cases.)
- adventured 1mo agoLooking back nearly 80 years, what has been the limit to transistor demand so far? Unlimited. What has been the limit to electricity demand globally? Unlimited. We can't get enough and never will. Costs have to become pretty severe to turn back the demand as well.
- moralestapia 1mo agoHence why taalas was one of the best strategic acquisitions of the year. I'm honestly baffled they were not acquired by somebody else (sorry AMD).
- adventured 1mo agoTaalas will be one of the great disaster investments of the early AI era. It'll be a near total write-down. The absolute worst market time to etch a model to a chip is right now (very rapid iteration). There is no scenario where they can keep up. The Taalas approach will be viewed as comically foolish within just a few years. Cerebras will win in terms of approach. It's 1998: hey, I can drastically speed up your web service, let's etch it right to silicon.
- moralestapia 1mo agoIt's 2026: let's etch nginx into silicon and get 10,000,000 rps at a cost of 0.1 US/day. Yes, please!
- redox99 1mo agoMost people probably don't care about nginx performance. It shouldn't be your bottleneck unless you serve massive amounts of static data.
- dyzone 1mo agoOk, how about postgres?
- mdp2021 1mo agoIn the case of needs to process natural language, instead, massive efficiency (esp. time) can be a game changer. It's like "you have two years to complete the project" vs "you have two hours to complete the project": if you can squeeze that "two years worth" into a negligible delay, it's a game changer.
- alightsoul 1mo ago
- dgellow 1mo ago> Then we can have fun conversations about "unlimited" "intelligence" and about what the price wars and profit margins of consumer AI products are when your average ChatGPT user costs the company $0.10 per month. We can have that discussion now: sounds like that would kill OpenAI and Anthropic
- brausepulver 1mo agoWhat LLM-specific hardware improvements should one expect? Seems to me that LLM inference is simple architecturally (matmul et al) so most scaling in hardware should come from general improvements (memory BW, packaging, interconnect, power).
- dorkypunk 1mo agoWhat you describe is basically Cerebras case, at the bottom it's just a really big die (about x28 an NVIDIA GB200) with a lot of work to reduce memory latency and improve throughput. What it's actually amazing is how can they make a chip so big and still have a decent yield to be commercially viable.
- LarsDu88 1mo agoA design that bakes the architecture into silicon would be 10x faster, and imagine a version that does all the multiplication ops using single log-amp addition versus dozens of transistors to cut down the amount of silicon used by 50x. The ceiling for AI optimized hardware is extremely high. Stack on top of that the fact that diffusion based models like the ones made by Inception Labs are far faster and more efficient than autoregressive LLMs and have an even higher ceiling of optimization (single step path prediction via model distillation versus 50 step denoise is currently an active area for image diffusion) The human brain is soon neither going to be more powerful nor energy efficient than the stuff we use to run AI.
- anonymous_user9 1mo agoConspicuously missing: power consumption figures
- wmf 1mo ago162 kW
- xattt 1mo agoI presume per rack? Can you imagine something radiating that much energy into a space in your home?
- wmf 1mo agoI guess because I have actually set foot in a data center I don't imagine literally every product in my home.
- walrus01 1mo agoIt's mandatory liquid cooling, so it's meant to be attached to a specialized liquid cooling loop that gets the heat outside the building. This is far beyond the practical maximums of like 10 to 15kW per 44U cabinet front to rear air cooling for 'regular' rackmount server stuff.
- ttul 1mo agoIndeed. You need 45 to 60 liters per second of cooling water flowing over a Cerebras wafer every minute to keep it under 90C. And that’s assuming the water leaves at 90C… More realistically, you need much more cooling water.
- 0xbadcafebee 1mo agoCS-3 used 100 liters of water with a cold plate (https://www.brownstoneresearch.com/bleeding-edge/ai-infrastructure-investment-continues-to-increase/ https://www.brownstoneresearch.com/bleeding-edge/ai-infrastr...). But you can also use refrigerants with a cold plate (https://eng.umd.edu/engineering-ai-public-good/cooling-data-centers-safer-two-phase-fluids https://eng.umd.edu/engineering-ai-public-good/cooling-data-...) or dielectric fluid in total immersion. Until they find a more power-efficient design, my guess is total immersion will come back in style.
- OutOfHere 1mo agoFive years from now, I don't know why anyone will still be using Nvidia for inference. Note that Cerebras is for inference only, not for training. I understand that Cerebras has competition, but this bodes even more poorly for Nvidia for inference. Nvidia may still have a role to play for training, however.
- wmf 1mo agoCerebras is only claiming ~2x the performance of Groqvidia which usually isn't enough for people to switch.
- kcb 1mo agoNvidia is at this time a pretty well run company tech wise. They are going to keep iterating on the inferencing hardware stack over the next five years too.
- OutOfHere 1mo agoThe only way I see in which Nvidia can catch up is by buying Cerebras.
- dgellow 1mo agoNVIDIA has the best supply chain in the entire game. They are the only ones who can produce at their scale. You really shouldn’t underestimate their position
- 4k0hz 1mo ago> Introducing the all new Cerebras CS-4, a revolutionary rack-scale solution that delivers upto 30x faster inference compared to GPUs, enhanced economics, and a simple path todeploy [sic] hyperscale capacity. Did nobody proofread this?
- SoMomentary 1mo agoSometimes I wonder if mistakes are now used to indicate the possibility that a human actually wrote it.
- ceejayoz 1mo agoThere’s been a spate of Reddit AI bots using all lower case in hopes of evading detection. It’s still incredibly obvious.
- VladVladikoff 1mo agoI don’t really visit Reddit much these days but would love to see an example.
- ceejayoz 1mo agohttps://www.reddit.com/r/ModSupport/comments/1tu67kn/influx_of_bot_comments_with_gen_z_wording_anybody/ https://www.reddit.com/r/ModSupport/comments/1tu67kn/influx_...
- geodel 1mo agoMaybe it is just part of their "compact design".
- algoth1 1mo agoIf they had ask Claude it would probably look like this: Introducing the all new Cerebras CS-4, a revolutionary rack-scale solution that delivers up to 30x faster inference compared to GPUs, enhanced economics, and a simple path to load-bearing hyper scale capacity.
- syntaxing 1mo agoI think the fun takeaway from this is that GPT 5.4 is probably 45B active parameters and GPT 5.6 Sol is closer to 50B.
- logicallee 1mo ago(Where did you see that?) This was also interesting: "CS-4 delivers more than 1,000 tokens per second on models exceeding 10 trillion parameters." Was it known that there were 10 trillion parameter models in use? I think the frontier providers keep the size of their models carefully hidden.
- ewild 1mo agoIt's rumored fable is around that 10T number
- walrus01 1mo agoIf this is true, it's even more impressive that some of the open weight models that are <3.5T in size, approx 33% of its size, are within a few points of it in the artificial analysis leaderboard.
- manquer 1mo agoNot necessarily, there could be diminishing returns on mere parameters count . There is nothing to say for example a 1 Quadrillion parameter model will be vastly more intelligent than current SOTA especially since new training data is largely synthetic today
- HDBaseT 1mo agoThat's precisely what he is saying, there is diminishing returns (or optimization left on the table).
- 1mo ago
- gpm 1mo agoIs it just me or is it bizarre that they're advertising old open-weight models. GLM 4.7 (December 2025) not 5 (Feb) 5.1 (April) or 5.2 (June). 5.3 (4 days ago) is, to be fair, not open weights yet... but there's a lot since 4.7. Kimi K2.7 (April) not K2.7-code (June) or K3 (July). Gemma 4 (April), Llama (April), and gpt-oss (August 2025) are up to date, but old (for models). Meanwhile the closed source GPT 5.6 sol is up to date (June)... Should potential purchasers take away from this that they're not going to be able to run recent models unless they front the cost of developing software or something?
- eli 1mo agoI think they run whatever models they get paid to run. But mostly from enterprise. They are clearly not interested in consumer dollars.
- gpm 1mo agoI mean the product is a server rack and while there's no advertised price I would assume it's six figures. So yes, an enterprise product. But even an enterprise is going to care about the difference between "we can run the model we want with support from the manufacturer" and "we have to purchase the product, and then spend another 6 figure sum having developers port a recent model to the product to use it".
- WarmWash 1mo agoI feel like 6-figures would be the clearance price on it...
- oceanplexian 1mo ago6 figures is a single mid range Xeon or Epyc server these days.
- kube-system 1mo agoYou’re at least an order or magnitude under… likely two. A single AI server with a mere 8 GPUs from Nvidia is already mid 6 digits. A rack system from Nvidia is mid 7 digits. There’s some info out there that suggests the CS1 had an 8 digits price tag, so it wouldn’t be surprising to see that here.
- sreekanth850 1mo agoAMD along with cerebras may probably compete with NVIDIA monopoly in near future. Also, NVIDIA will have competition form multiple companies. Just my prediction.
- eitally 1mo agoMaybe, but GPU is just one aspect of NVIDIA's dominance. If you are buying Vera Rubin GPUs, you're getting an NVL72 rack, which is only one of several racks that you're probably buying. You'll also need your NVIDIA racks with NVIDIA networking & storage gear, too. At the end of the day, they're "vertically integrated" for your accelerated computing data center (e.g. the "AI Factory"). This doesn't even count the software layer, where CUDA + CUDA-X (not to mention the software for all the sysadmin pieces) has a huge first mover advantage over anyone else.
- zarzavat 1mo agoIs CUDA still a moat? Are we not at the point where frontier models can reimplement software stacks, given you throw enough tokens at the problem.
- vatsachak 1mo agoYou just proved that AI cannot currently do that
- incrudible 1mo agoIt can definitely create a software stack for you if you hold it right, but the software stack supported by a trillion dollar company with decades of expertise, that also uses AI to improve its stack is probably gonna be better.
- tomrod 1mo agoOn the reverse side, there are fundamental limits to the number of ways you can perform certain actions, and agents are both diligent as well as able to swarm. If you have your tests beforehand, there is a chance.
- reilly3000 1mo ago> CS-4 delivers more than 1,000 tokens per second on models exceeding 10 trillion parameters Oops did they just out GPT-5.6 sol’s parameter count?
- whatever1 1mo agoI mean we kinda know the frontier models are multi trillion parameter models. The only open weights that are close to the frontier are that size too
- verdverm 1mo agosave qwen3.8 27B which is outclassing much larger models and is in spitting distance of the top 10 in https://artificialanalysis.ai/models#intelligence https://artificialanalysis.ai/models#intelligence
- Vax- 1mo agoI wonder why they removed DeepSWE from their incorporates evaluations
- verdverm 1mo agoThey didn't afaict https://artificialanalysis.ai/agents/coding-agents?coding-agents-performance-chart=deep-swe https://artificialanalysis.ai/agents/coding-agents?coding-ag... It seems it takes some time to run a new model on all the benchies, not sure they run all models on all of them either
- sho 1mo agoSol is supposed to be 5T according to rumour. The imminent Astra is allegedly 10
- nozzlegear 1mo agoRumors and allegations aren't worth much. Why don't they just tell us mere mortals?
- tamimio 1mo agoI wonder what are the benchmarks of hashcat on different hashes.
- kobe_bryant 1mo agocan these vibe coded sites please set a max width and overflow so their sites work fine on mobile
- dgellow 1mo agoGood news, future models will have your comment in their training set, making them slightly more likely to fix that problem!
- singingtoday 1mo agoWhat a time to be alive
- lostmsu 1mo agoKV caching status? What's the point of 1000tok/s if you have to do prefill on every agentic turn which at 100k depth would make it 1.5 min latency every turn?
- walrus01 1mo agoInformation about RAM type/size and connection topology of the RAM to be used for context cache seems to be conspicuously absent from the slick looking marketing materials.
- gpm 1mo agoThere's a few more details at the bottom of this page: https://www.cerebras.ai/blog/introducing-cerebras-cs-4 https://www.cerebras.ai/blog/introducing-cerebras-cs-4 44GB on-chip-sram * 3 chips. Per chip: 43.2 PB/s memory access + 53.5 PB/s on-chip fabric bandwidth + 2.4 Tbits/s "IO" bandwidth (I think that means their RoCE v2 RDMA over Ethernet interface). I suspect there might be a certain amount of customization for how much RAM they attach when you order it.
- porridgeraisin 1mo agoThey have managed to make the external link 300GBps/2us. Cs3 was 150/5. This is 1/3rd blackwells nvlink c2c bandwidth already. Not too bad. We can make KV cache offload work with that I suppose. If magically KV cache was not an issue, pipeline parallelism on cerebras can be quite pleasant. As for the KV cache offload, I have hopes their CPO solution they're trying with that canadian company ends up bearing fruit.
- ethanzhang1024 1mo agoIf cerebars is performing well, why didn't its predecessor, server S-3, become the largest API token provider on OpenRouter, surpassing the official model releases?
- doctorpangloss 1mo agoit only takes ~445 GB300 NVL72 (about $22b) to run ALL of openrouter demand for a year. Microsoft rolled out $32b of DC 2026Q1. imo the issue is that most openrouter demand is inauthentic activity (things that anthropic and openai models will refuse to do like pretend to not be bots when interacting with humans)
- senordevnyc 1mo agoI was curious so I looked it up: looks like a GB300 NVL72 is about $4M. So $22B would buy you 5500 such racks, no?
- selectodude 1mo agoGPU cost is about half of datacenter cost. Other half is cooling, power and networking.
- HDBaseT 1mo agoIt is worth mentioning, the OpenRouter demand isn't static though. It has increased week on week since early 2026.
- aurareturn 1mo agoI thought your numbers must be wrong. So I plugged 288 trillion tokens/month (OpenRouter's current rate), 500 billion MoE model average, and the math comes out to be around 620 B200 GPUs minimum. So basically, OpenRouter's volume must be absolutely tiny compared to the volume hyperscalers are getting.
- Mattwmaster58 1mo ago
- aneryu 1mo agoIt would be even better if a version available to individual users were released soon.
- ttul 1mo agoI’ll get that 250kW home power service dropped in next week!
- gpm 1mo agoThey do offer API services to individual users... though with a set of models that makes it unlikely that you want to use it. They are promising Qwen 3.8 27B any day now though*. if you have the money as an "individual user" to purchase one of their racks... save your money and retire. * Actually they sent out an email claiming they already have it, but I don't seem to have access, they're promising to release it to the "shared tier" any day now.
- fragmede 1mo ago> save your money and retire. Now that this hypothetical person has retired, what are they gonna do all day? Just sit on the beach and drink Mai Tais? If that's what they wanna do, sure, but nerds gonna nerd, and if I had that kind of money to retire on, I'd totally buy some ridiculously expensive AI box for fun.
- gpm 1mo ago
- sva_ 1mo ago> enabling massive clusters and models with more than 50 trillion parameters
- avantnyc 1mo agoCerebras should slowly also move to dgx/ryzen market for a desktop version for masses at affordable price yet providing substantial tokens/second on desktop
- denizay 1mo agoThe comparison seems incomplete. CS‑4 is a full rack-scale system with three wafer-scale processors, but the exact GPU models, GPU count, power consumption, price information are not disclosed. We still don't know if buying a multi-GPU rack (or racks) is cheaper and/or more efficient in power. The fact that they didn't disclose these numbers makes me believe that the numbers are not in their favor. And personally, makes me see them as disingenuous.
- adventured 1mo agoOpenAI needs to immediately move to acquire Cerebras. Nvidia's extreme margin is the opportunity for OpenAI's cost reduction. Buying Cerebras would pay for itself and they should take all of its future production (after filling required contracts). Right now China's models have no silicon moat. Cerebras as a drastic speed-up / cost-reduction potential, can assist in building a competitive moat. And every time a Cerebras pops up, OpenAI or Anthropic should eat them if at all possible. There's no stand-alone frontier AI company of great scale in the near future that doesn't have a large silicon advantage in-house. Apple knew it in smartphones, Google figured it out a long time ago as well.
- wmf 1mo agoDo you know about Jalapeno?
- mmmeff 1mo ago^ OpenAI is partnering with Cerebras while simultaneously investing in their own silicon play. Hedged bets. After sitting thru their keynote today, it makes sense. The main throughput speedups they tout are an obvious evolution of the GPU that all companies will be building in the next year. Wafer-scale interconnected memory and compute is just going to beat out mountains of network cabling any day on both cost and performance metrics.
- thefounder 1mo agoWith what? More debt? What will nvidia say?
- deleted 1mo ago[deleted]
- arthurcolle 1mo agoWhat's the sticker price? If I have 20 million in the bank can I just like buy one or what
- selimonder 1mo agoThat "GPU" comparison is the vaguest i seen so far
- epolanski 1mo agoTrue, it's also "per user", somehow, but I think it's a misleading metric. Cerebras chips take the whole wafer? A single TSMC wafer contains 60 to 65 B200s, assuming 70% yields that's 40ish wafers per die. Cerebras cannot redefine wafer economics.
- tjoff 1mo agoDepends on what you mean, they have more redundancy which means that the yield can be much higher.
- KronisLV 1mo agoGod I wish they'd back up all of those claims by offering a subscription of Kimi K3 and GLM 5.3, not some outdated GLM 4.7 instance that they then proceed to call a preview model and say that they'll remove it, leaving users only with GPT-OSS 120B which is nigh useless nowadays: https://support.cerebras.net/articles/9996007307-cerebras-code-faq https://support.cerebras.net/articles/9996007307-cerebras-co... and https://www.cerebras.ai/pricing https://www.cerebras.ai/pricing Guess they don't care about regular devs atm and are focused only on hardware sales.
- preommr 1mo ago> GPT-OSS 120B which is nigh useless nowadays: I still think that was a really great model that got overlooked. It was really great in terms of latency/throughput while still being fairly intelligent. I was planning on using it for a design tool, but moved over to luna since it's comparable speeds and cost for a lot more intelligence.
- KronisLV 1mo ago> I was planning on using it for a design tool, but moved over to luna since it's comparable speeds and cost for a lot more intelligence. Everyone should occasionally go back to the old models to see how much worse they were, like even a year ago you could generate results but they were typically full of bugs and you have to fix a non-insignificant amount of it all manually: https://blog.kronis.dev/blog/i-blew-through-24-million-tokens-in-a-day/ https://blog.kronis.dev/blog/i-blew-through-24-million-token... Admittedly that post was before agentic development truly took off and that 3k EUR figure when paying per API tokens would nowadays be closer to like 6k EUR for the volume of work I do, but still. It's the same how Qwen 2.5 was pretty problematic for anything remotely serious, same with Qwen 3 Coder Next (80B), and at least the most recent versions are getting better but still not quite good enough in real world use cases outside of benchmarks. They've come a long way, regardless!
- 5555watch 1mo ago>Everyone should occasionally go back to the old models to see how much worse they were, like even a year ago you could generate results but they were typically full of bugs Oh yeah, I'm still amazed how good the current iteration of models are for coding (I have a fear it's too good to be true - so will get taken away..). Exactly a year ago I switched from GPT 5 to Gemini just because the coding with R language was terrible; and even with Python it kept forgetting and mixing basic stuff. Gemini at the time had much longer context window and was miles ahead on R syntax. Current experience of just leaving a Codex Agent chug until a stable solution is completed is still mind blowing to me.
- rajnathani 1mo agoInterestingly they’re still on the WSE-3 (5nm TSMC) wafer chip and slightly bumped up the specs there (overlocking mostly it seems), for why it’s called WSE-3 Turbo now. I think people were also expecting WSE-4, as it’s been 2 years now since WSE-3 was launched.
- xmorse 1mo agoCerebras is very fast but you can basically never use it because of its scarcity
- rbanffy 1mo agoImpressive that this is an "interim" product, the start of a new line that ought to be continued with the WSE-4 family, where they are supposed to use a 3nm process and, maybe, 3D stacked SRAM. The modular architecture also points towards field upgrades that are badly needed for AI datacenter builders.
- charlielidbury 1mo agoonly 44GB * 3 of VRAM per rack :O I guess you'd need a DOZEN(s) of these to host a large model with long context KV caches?
- jaumesnts 1mo ago[dead]
- bearjaws 1mo agoI was hoping to see Cerebras launch something other than GPT-OSS-120b in production this week, especially with GLM4.7 going away. If they could launch Qwen 27b or Deepseek Flash that would be amazing.
- anarticle 1mo agoReally wish they’d host more models for us normies. My guess is OAI will buy / subsidize them with terms that will close off open models.
- LarsDu88 1mo agoThat's a Lord of the Rings length book from a frontier model (Sol) every 15 minutes.
- kitk 1mo agothis is cool!