4 ms·
How could that possibly work? Deepseek was undercutting every other provider by an order of magnitude on cached tokens. Do they just set a super low caching ti
by Tharre 2mo ago
How could that possibly work? Deepseek was undercutting every other provider by an order of magnitude on cached tokens.
Do they just set a super low caching time and hope that drops effective cache rates low enough? Do all other providers somehow overcharge by that much? Are they just going to sell it as a loss leader?
- lcampbell 2mo ago> Do all other providers somehow overcharge by that much? This, I think. Cached inputs have an opportunity cost (keeping the KV cache until use) but a hit is basically free. “Basically” - if the cache is offloaded to system RAM or NVMe there’s some scheduling overhead. From a consumer viewpoint a more interesting metric than the raw costs is cached cost * hitrate + input cost * (1 - hitrate) from a personal standpoint rather than a per-provider one (e.g. if OpenRouter is blindly dispatching your requests you might have a bad time).
- Tharre 2mo agoSystem RAM and/or NVMe storage still has a real cost. And swapping out the context between VRAM and system RAM / NVMe still consumes bandwidth. I don't have a clue on what the real cost to inference providers comes out to, but it seems really weird that there would be such a big gap, in what should be a pretty competitive market.
- deleted 2mo ago[deleted]
- skeledrew 2mo ago> if OpenRouter is blindly dispatching your requests This can somewhat be the case, depending on your config. I updated mine to make DeepSeek high priority because I was having a lot of cache misses and reliability issues with the default (cheapest (at face value)) providers, and cost was actually higher overall than anticipated. Was smooth sailing from then; might have to tweak things again now pricing has changed though.
- nyargh 2mo agoI was having the same issue. Did not find any provider that had a cache hit rate anywhere close to DeepSeeks own API. Not sure if this was an OpenRouter issue or with the other inference providers.
- brandon997 2mo agoYou could also check out token.dance. They offer DeepSeek and other models through an OpenAI-compatible API, and the pricing seems pretty competitive. Might be worth comparing the cache rates.