5 ms·
I think there's clearly a "Speed is a quality of it's own" axis. When you use Cereberas (or Groq) to develop an API, the turn around speed of iterating on jobs
by estsauver 7mo ago
I think there's clearly a "Speed is a quality of it's own" axis. When you use Cereberas (or Groq) to develop an API, the turn around speed of iterating on jobs is so much faster (and cheaper!) then using frontier high intelligence labs, it's almost a different product.
Also, I put together a little research paper recently--I think there's probably an underexplored option of "Use frontier AR model for a little bit of planning then switch to diffusion for generating the rest." You can get really good improvements with diffusion models! https://estsauver.com/think-first-diffuse-fast.pdf https://estsauver.com/think-first-diffuse-fast.pdf
- refulgentis 7mo agoI'm very worried for both. Cerebras requires a $3K/year membership to use APIs. Groq's been dead for about 6 months, even pre-acquisition. I hope Inception is going well, it's the only real democratic target at this. Gemini 2.5 Flash Lite was promising but it never really went anywhere, even by the standards of a Google preview
- freeqaz 7mo agoYou can call Cerebras APIs via OpenRouter if you specify them as the provider in your request fyi. It's a bit pricier but it exists!
- andai 7mo agoI used their API normally (pay per token) a few weeks ago. Their Coding Plan appears to be permanently sold out though.
- nl 7mo agoTaalas is interesting. 16,000 TPS for Llama on a chip. https://taalas.com/ https://taalas.com/
- micw 7mo agoOn a very old model, it's more like 16.000 garbage words/s
- nl 7mo agoLlama 3.1 8B is pretty useful for some thing. I use it to generate SQL pretty reliably for example. They are doing an updated model in a month or so anyway, then a frontier level one "by summer".
- numeri 7mo agobut Taalas had to quantize Llama 3.1 8B to death to get it to fit. It can't produce coherent non-English text at all.
- patapong 7mo agoI do wonder if there are tasks where 16k garbage words/s are more useful than 200 good words per second. Does anyone have any ideas? Data extraction perhaps?
- pnocera 7mo agoA politician communication agent maybe...
- DeathArrow 7mo agoI wonder how many token per seconds can they get if they put Mercury 2 on a chip.
- replete 7mo agoIts exciting to see, but look at the die size for only an 8b model
- Nihilartikel 7mo agoNeat! I had been wondering if anyone was trying to implement a model in silico. We're getting closer to having chatty talking toasters every day now!
- 7thpower 7mo agoWhat do you mean by Grow is dead since about 6 months ago? Not refuting your point, but I’m curious.
- refulgentis 7mo agoNo new model since GPT-OSS 120B, er maybe Kimi K2 not-thinking? Basically there were a couple models it normally obviously support, and it didn't. Something about that Nvidia sale smelled funny to me because the # was yuge, yet, the software side shut down decently before the acquisition. But that's 100% speculation, wouldn't be shocked if it was: "We were never looking to become profitable just on API users, but we had to have it to stay visible. So, yeah, once it was clear an Nvidia sale was going through, we stopped working 16 hours a day, and now we're waiting to see what Nvidia wants to do with the API"
- vessenes 7mo agoThe groq purchase was designed to not trigger federal oversight of mergers, so you buy out the ‘interesting’ part, leave a skeleton team and a line of business you don’t care about -> no CFIUS, no mandatory FTC reporting -> smoother process.
- ainch 7mo agoI don't think it's a good comparison given Inception work on software and Cerebras/Groq work on hardware. If Inception demonstrate that diffusion LLMs work well at scale (at a reasonable price) then we can probably expect all the other frontier labs to copy them quickly, similarly to OpenAI's reasoning models.
- refulgentis 7mo agoDefinitely depends on what you're buying, maybe some of the audience here was buying Groq and Cerebras chips? I don't think they sold them but can't say for sure. If you're a poor schmoke like me, you'd be thinking of them as API vendors of ~1000 token/s LLMs. Especially because Inception v1's been out for a while and we haven't seen a follow-the-leader effect. Coincidentally, that's one of my biggest questions: why not?
- estsauver 7mo agoI am currently using their APIs on a paygo plan, I think it might just be a capacity issue for new sign ups.
- behnamoh 7mo agoOnce again, it's a tech that Google created but never turned into a product. AFAIK in their demo last year, Google showed a special version of Gemini that used diffusion. They were so excited about it (on the stage) and I thought that's what they'd use in Google search and Gmail.
- refulgentis 7mo agoGoogle did not create it; it is correct there was a Gemini that used diffusion, you could apply for access (not via API). It was okay
- Leynos 7mo agoCerebras are on OpenRouter.