3 ms·
I wrote about this a couple of weeks ago. It's actually often the biggest cost and it tends to be hidden away on most platforms! https://martinalderson.com/pos
by martinald 1mo ago
I wrote about this a couple of weeks ago. It's actually often the biggest cost and it tends to be hidden away on most platforms!
https://martinalderson.com/posts/watch-out-for-cache-read-costs/ https://martinalderson.com/posts/watch-out-for-cache-read-co...
Btw I still haven't came across any decent model that is <$0.01/MTok cache costs apart from deepseek thru their official API (even with the price increases).
Seems like a bit of an opportunity for someone to take - drop cache read costs significantly.
- dakolli 1mo agoThat's because Deepseek invented the paradigm of prompt caching, they are the SOTA when it comes these techniques. Despite them open sourcing all their research, nobody beats them. edit: I do wish openrouter would let you sort providers by Cache Hit % and Cache cost. These are the only things that matter to me at this point when choosing a provider.
- minimaxir 1mo agoYou can click the table headers to sort Ascending/Descending.
- dakolli 1mo agoYou can sort by the cost cache cose, but you cannot by cache hit %. You have to click on the provider and see what their cache hit % is. A provider could have a super low cache cost, but a 50% cache hit percentage, making the cheap price of cache read's meaningless.
- andai 1mo agoWait, what does that number mean? I thought it always uses the cache price when the prefix matches.
- dakolli 1mo agoWhen the prefix matches a request sent to the same Providor. The thing is the TTL is different for each provider, some cache for 5 minutes some cache for 1hr. Its ideal to only use one provider per agent session / and per model with the best cache hit % if you care about costs.
- andai 1mo agoBut before 5m the hit rate is 100%, and after it's 0%? Why is there a probability? Is there some stochastic process that takes place during those 5 minutes that determines whether or not you get the discount?
- RussianCow 1mo agoPresumably the number that OpenRouter shows is averaged across all requests.
- Bolwin 1mo agoCache hit % on openrouter is not a good metric, it's mainly driven by openrouter's own provider juggling than the providers themselves
- dakolli 1mo agoThis is not true, there isn't even a way to see a cache hit % model for a specific model, that wouldn't make any sense. You are confusing what I'm saying with cache cost, that has nothing to do with effective cache hit %. I'm talking about when you click on a specific provider for a specific model, you can scroll down on the view and see their cache hit % for that model [0]. These cache Hit % are accurate, I've done a ton of testing of this myself. The cache hit % is one of the most important metrics as far as estimating cost. There are many providers with cheap cache reads, but have an effective cache hit % of 30%, making their cheaper cache pricing meaningless compared to another provider who charges more but has a 85% cache hit percentage. [0]: https://openrouter.ai/deepseek/deepseek-v4-flash-0731?endpoint=f954a48e-1c90-433b-9348-4720f8030331 https://openrouter.ai/deepseek/deepseek-v4-flash-0731?endpoi... scroll down on the provider/model card and you'll see a field called cache hit %, its different for every provider/model. I don't use routing on openrouter, I strictly use models with a single provider and no fallback, at least for use with harnesses its pretty dumb to route requests to multiple providers you are busting your cache every other request and increasing costs by 20-50%.
- Implicated 1mo agoI think you're arguing the same general point that the person you're responding to is. But you're saying he's not understanding - he understands that they report a cache hit % but you can't look at that public metric with any level of accuracy _because_ most people aren't pinning their providers and they _are_ getting juggled around which is bringing that metric down. That's not to say that specific providers might have issues or worse cache implementations - but it stands that if openrouter is juggling the requests back and forth by default then _that alone_ is breaking caches on those requests in huge numbers.
- 1mo ago
- orbital-decay 1mo ago>Deepseek invented the paradigm of prompt caching Caching was always here, you don't need to do anything special to get it on a single user local backend running a base model or a chatbot in the first place. Among commercial providers, OpenAI adopted it in 4o first.
- dominotw 1mo ago> open sourcing all their research, is this true?
- gpugreg 1mo agoNot all their research, but certainly a lot: https://github.com/orgs/deepseek-ai/repositories?q=sort%3Astars https://github.com/orgs/deepseek-ai/repositories?q=sort%3Ast...
- sieve 1mo agoFor me, an average long session results in about 200-300M cached input, 4-800K input, 2-400K output. Mostly the lower bound. Output depends on how much the model thinks. There are two problems here: - cache hit pricing (both Muse Spark 1.2 Contributor and MiMo 2.5 are around the $0.002-3/M mark) - cache persistence time Muse Spark drops the cache in less than 5m. MiMo keeps it around for at least an hour based on my experience with whoever is serving it for OpenCode. This difference itself will inflate bills massively. A 500K token input repeatedly read by MS 1.2 for full input price 12 times an hour = $0.60. You would be expecting $0.012. So a 50x difference. Same thing on MiMo 2.5 is $0.018 because of longer cache times.
- sourcecodeplz 1mo agoeven with the 5m cache, Muse Spark Contribs is still best bang for your buck for the intelligence you get. it is basically the old dsv4-flash prices, but even more smart.
- sieve 1mo agoI have used all three extensively. DS4 Flash is quite smart. I would rank MS 1.2 below it. MiMo is the dumbest of them all but good enough for basic stuff. MiMo wins handsomely if you want to think about your code for minutes at a time as you write. I use it to make changes as I think. I know it will screw up some stuff. I then switch to MS/DS4 once every few hours and have it do a code review and fix the broken stuff. So much cheaper than getting MS to do it on its own.