4 ms·
I don’t think they will until they change the architecture. They don’t have a prefix cache like other providers, or at least don’t have a discount in their bil
by scosman 1mo ago
I don’t think they will until they change the architecture.
They don’t have a prefix cache like other providers, or at least don’t have a discount in their billing structure. Each message charges for the whole context window. It’s wildly more expensive for long multi turn scenarios with lots of tool calls (coding). It’s better for short few turn tasks.
Edit: I don’t know if they actually have a proper cache. This could just be a billing artifact.
- dannyw 1mo agoCerebras supports prompt caching and has a doc about it. A fairly standard automatic prefix-based implementation with 5min expiry. They do not seem to discount cached input for the self-serve Developer tier. Maybe they do for enterprise rate cards? https://inference-docs.cerebras.ai/capabilities/prompt-caching https://inference-docs.cerebras.ai/capabilities/prompt-cachi...
- scosman 1mo agoah, that makes it feasible! Okay, glad it's not technical limit. They should fix the pricing...