4 ms·
If cerebars is performing well, why didn't its predecessor, server S-3, become the largest API token provider on OpenRouter, surpassing the official model relea
by ethanzhang1024 1mo ago
If cerebars is performing well, why didn't its predecessor, server S-3, become the largest API token provider on OpenRouter, surpassing the official model releases?
- doctorpangloss 1mo agoit only takes ~445 GB300 NVL72 (about $22b) to run ALL of openrouter demand for a year. Microsoft rolled out $32b of DC 2026Q1. imo the issue is that most openrouter demand is inauthentic activity (things that anthropic and openai models will refuse to do like pretend to not be bots when interacting with humans)
- senordevnyc 1mo agoI was curious so I looked it up: looks like a GB300 NVL72 is about $4M. So $22B would buy you 5500 such racks, no?
- selectodude 1mo agoGPU cost is about half of datacenter cost. Other half is cooling, power and networking.
- HDBaseT 1mo agoIt is worth mentioning, the OpenRouter demand isn't static though. It has increased week on week since early 2026.
- aurareturn 1mo agoI thought your numbers must be wrong. So I plugged 288 trillion tokens/month (OpenRouter's current rate), 500 billion MoE model average, and the math comes out to be around 620 B200 GPUs minimum. So basically, OpenRouter's volume must be absolutely tiny compared to the volume hyperscalers are getting.
- Mattwmaster58 1mo agoFor reference, Google serves >3 quadrillion/mo. [1] https://x.com/ren_stocks/status/2056946641815396718?s=20 https://x.com/ren_stocks/status/2056946641815396718?s=20
- aurareturn 1mo agoWhich is only 10x more than Open Router by the way.
- walrus01 1mo agoWithout having any inside information, one possible theory: All or a vast majority of of the cerebras manufacturing capacity was going to a few companies that aren't publicly available inference providers on openrouter, for their own internal use. or The asking price of the S-3, no matter how speedy it might be, for small/medium size customers made it economically prohibitive to purchase and use to sell public inference vs. buying more common nvidia b200 or whatever.
- smallerize 1mo agoIf you're willing to pay a significant premium for latency, why use openrouter? And anyway Cerebras only supported a few specific models.
- aseipp 1mo agoThe WSE is very expensive to build, and they have a waiting list of customers who are already willing to pay a lot of money for the available supply.
- wmf 1mo agoCerebras provides high-speed inference at high cost. It's never going to be the cheapest and thus it will probably remain niche.
- zurfer 1mo agoBut that's supply and demand, not technology. Right now a lot more people want their inference than they can supply. as supply catches up in the next 5-10 years, the underlying tech at scale is probably cheaper than GPUs per token produced.
- aurareturn 1mo agoProbably the same reason why there are more people who takes buses, subways, trains than drive Ferraris.
- epolanski 1mo agoI love how instead of comparing a Ferrari (fast and expensive) to some average car (not fast, not expensive) to make your point..you went for public transport where your comparison cracks from multiple angles.
- aurareturn 1mo agoBecause comparing it to an average car is wrong. It isn't about fast/expensive. It is about fast/capacity. A Ferrari can seat up to 4 people[0], which is about the same as the average car. Capacity doesn't change much. Meanwhile, a bus/subway system is meant to support millions of people. Tokyo's metro has to support up to 37 million people. You can't do that with Ferraris. [0]https://www.ferrari.com/en-EN/auto/ferrari-purosangue https://www.ferrari.com/en-EN/auto/ferrari-purosangue
- gampleman 1mo agoCerebras capacity was pretty much entirely bought out at some point. We needed it and couldn't get it.
- petesergeant 1mo agoI use Cerebras via OpenRouter. It’s every bit as fast and reliable for my needs as claimed. I suspect the reason is that they can either be making peanuts selling inference to plebs like me via OpenRouter, or making bank selling the more expensive models to businesses directly. In short: I would be very surprised if they have die capacity, and are at this point maximising revenue per chip.
- 0xbadcafebee 1mo agoBecause they aren't selling inference, they're selling hardware. The only reason they sell any tokens on OpenRouter is so they get on the benchmark that shows them as the fastest provider. It's free advertising.