3 ms·
I tried Cerebras with GLM-4.7 (not Flash) yesterday using paid API credits ($10). They have rate limits per-minute and it counts cached tokens against it so you
by HumanOstrich 9mo ago
I tried Cerebras with GLM-4.7 (not Flash) yesterday using paid API credits ($10). They have rate limits per-minute and it counts cached tokens against it so you'll get limited in the first few seconds of every minute, then you have to wait the rest of the minute. So they're "fast" at 1000 tok/sec - but not really for practical usage. You effectively get <50 tok/sec with rate limits and being penalized for cached tokens.
They also charge full price for the same cached tokens on every request/response, so I burned through $4 for 1 relatively simple coding task - would've cost <$0.50 using GPT-5.2-Codex or any other model besides Opus and maybe Sonnet that supports caching. And it would've been much faster.
- deleted 9mo ago[deleted]
- twalla 9mo agoI hope cerebras figures out a way to be worth the premium - seeing two pages of written content output in the literal blink of an eye is magical.
- Miraste 9mo agoI wonder why they chose per minute? That method of rate limiting would seem to defeat their entire value proposition.
- p91paul 9mo agoIn general, with per minute rate limiting you limit load spikes, and load spikes are what you pay for: they force you to ramp up your capacity, and usually you are then slow to ramp down to avoid paying the ramp up cost too many times. A VM might boot relatively fast, but loading a large model into GPU memory takes time.
- mlyle 9mo agoThe pay-per-use API sucks. If you end up on the $50/mo plan, it's better, with caveats: 1 million tokens per minute, 24 million tokens per day. BUT: cached tokens count full, so if you have 100,000 tokens of context you can burn a minute of tokens in a few requests.
- solarkraft 9mo agoIt’s wild that cached tokens count full - what’s in it for you to care about caching at all then? Is the processing speed gain significant?
- mlyle 9mo agoNot really worth it, in general. It does reduce latency a little. In practice, you do have a continuing context, though, so you end up using it whether you care or not.
- indigodaddy 9mo agoTry a nano-gpt subscription. Not going to be as fast as cerebras obviously but it's $8/mo for 60,000 requests
- Imustaskforhelp 9mo agoI know this might not be the most effective use case but I had ended up using the try AI feature in cerebras which opens up a window in browser Yes, it has some restrictions as well but it still works for free. I have a private repository where I ended up creating a puppeteer instance where I can just input something in a cli and then get output in cli back as well. With current agents. I don't see how I cannot just expand that with a cheap model like (think minimax2.1 is pretty good for agents) and get the agent to write the files and do the things and a loop. I think the repository might have gotten deleted after I resetted my old system or similar but I can look out for it if this interests you. Cerebras is such a good company. I talked to their CEO on discord once and have following it for >1-2 years now. I hope that they don't get enshittified with openAI deal recently & they improve their developer experience because people wish to pay them but now I had to do a shenanigan which was for free (but also its just that I was curious about how puppeteer works so I wanted to find if such idea was possible itself or not & I really didn't use it that much after building it)
- cmrdporcupine 9mo agoI use GLM 4.7 with DeepInfra.com and it's extremely reasonable, though maybe a bit on the slower side. But faster than DeepSeek 3.2 and about the same quality. It's even cheaper to just use it through z.ai themselves I think.