6 ms·
Tried this out with Cline using my own API key (Cerebras is also available as a provider for Qwen3 Coder via via openrouter here: https://openrouter.ai/qwen/qwe
by Flux159 1y ago
Tried this out with Cline using my own API key (Cerebras is also available as a provider for Qwen3 Coder via via openrouter here: https://openrouter.ai/qwen/qwen3-coder https://openrouter.ai/qwen/qwen3-coder) and realized that without caching, this becomes very expensive very quickly. Specifically, after each new tool call, you're sending the entire previous message history as input tokens - which are priced at $2/1M via the API just like output tokens.
The quality is also not quite what Claude Code gave me, but the speed is definitely way faster. If Cerebras supported caching & reduced token pricing for using the cache I think I would run this more, but right now it's too expensive per agent run.
- Havoc 1y agoThis seems to be rate limited by message not token so the lack of cache may matter less
- deleted 1y ago[deleted]
- Flux159 1y agoThe lack of caching causes the price to increase for each message or tool call in a chat because you need to send the entire history back after every tool call. Because there isn’t any discount for cached tokens you’re looking at very expensive chat threads.
- NitpickLawyer 1y agoYes, but the new "thing" now is "agentic" where the driver is "tool use". So at every point where the LLM decides to make a tool use, there is a new request that gets sent. So a simple task where the model needs to edit one function down the tree, there might be 10 calls - 1st with the task, 2-5 for "read_file", then the model starts writing code, 6-7 trying to run the code, 8 fixing something, and so on...
- itsafarqueue 1y agoYup. If you’ve ever watched a 60+ minute agent loop spawning sub agents, your “one message” prompt leaves you several hundred messages in the hole.
- andhuman 1y agoNo it’s by token. The FAQ says this: > Actual number of messages per day depends on token usage per request. Estimates based on average requests of ~8k tokens each for a median user. https://cerebras-inference.help.usepylon.com/articles/3468865440-how-do-you-calculate-messages-per-day https://cerebras-inference.help.usepylon.com/articles/346886...
- jtbayly 1y agoHow did you find that? Are you sure it applies to Cerebras Code Pro or Max?
- sysmax 1y agoAdding entire files into the context window and letting the AI sift through it is a very wasteful approach. It was adopted because trying to generate diffs with AI opens a whole new can of worms, but there's a very efficient approach in between: slice the files on the symbol level. So if the AI only needs the declaration of foo() and the definition of bar(), the entire file can be collapsed like this: class MyClass { void foo(); void bar() { //code } } Any AI-suggested changes are then easy to merge back (renamings are the only notable exception), so it works really fast. I am currently working on an editor that combines this approach with the ability to step back-and-forth between the edits, and it works really well. I absolutely love the Cerebras platform (they have a free tier directly and pay-as-you-go offering via OpenRouter). It can get very annoying refactorings done in one or two seconds based on single-sentence prompts, and it usually costs about half a cent per refactoring in tokens. Also great for things like applying known algorithms to spread out data structures, where including all files would kill the context window, but pulling individual types works just fine with a fraction of tokens. If you don't mind the shameless plug, there's a more explanation how it works here: https://sysprogs.com/CodeVROOM/documentation/concepts/symboledits/ https://sysprogs.com/CodeVROOM/documentation/concepts/symbol...
- postalcoder 1y agothis works if your code is exceptionally well composed. anything less can lead to looney tunes levels of goofiness in behavior, especially if there’s as little as one or two lines of crucial context elsewhere in the file. This approach saves tokens theoretically, but i find it can lead to wastefulness as it tries to figure out why things aren’t working when loading the full file would have solved the problem in a single step.
- sysmax 1y agoIt greatly depends on the type of work you are trying to delegate to the AI. If you ask it to add one entire feature at a time, file level could work better. But the time and costs go up very fast, and it's harder to review. What works for me (adding features to huge interconnected projects), is think what classes, algorithms and interfaces I want to add, and then give very brief prompts like "split class into abstract base + child like this" and "add another child supporting x,y and z". So, I still make all the key decisions myself, but I get to skip typing the most annoying and repetitive parts. Also, the code don't look much different from what I could have written by hand, just gets done about 5x faster.
- BenGosub 1y agoIf they say it costs $50 per month, why do you need to make additional payments?
- davidweatherall 1y ago$50 per month is their SaaS solution that let's you make 1000 requests per day. The openrouter cost is the raw API cost if you try to use qwen3-coder via the pay as you go model when using Cline
- seunosewa 1y agoThe Cerebras.ai plan offers a flat fee of $50 or $200. The API price is not a reason to reject the subscription price.
- dedene 1y agoThe flat fee is for a fixed max amount of tokens per day. Not requests, tokens.
- deleted 1y ago[deleted]
- deleted 1y ago[deleted]
- beastman82 1y agothe API price is not very relevant to this flat fee service announcement. In fact it seems obvious that you should use the flat fee model instead
- waldrews 1y agoDoes caching make as much sense as a cost saving measure on Cerebras hardware as it does on mainstream GPU's? Caching should be preferred if SSD->VRAM is dramatically cheaper than recalculation. If Cerebras is optimized for massively parallel compute with fixed weights, and not a lot of memory bandwidth into or out of the big wafer, it might actually make sense to price per token without a caching discount. Could someone from the company (or otherwise familiar with it) comment on the tradeoff?