4 ms·
If that wasn't impressive enough, it's actually ~60x cheaper if you take into account the typical cache-read/input/output split in agentic coding, and the deep
by xynelius 2mo ago
If that wasn't impressive enough, it's actually ~60x cheaper if you take into account the typical cache-read/input/output split in agentic coding, and the deep discount for cache reads offered by DeepSeek. Opencode has some public data on the typical split [1]:
For DeepSeek V4 Pro the typical split is 750 in, 290 out, 82k cached.
Cost per request for V4 Pro: $0.000875 per request.
Equivalent Opus cost (w/o taking into account cache write costs): $0.052 per request.
[1] https://opencode.ai/docs/go/#usage-limits https://opencode.ai/docs/go/#usage-limits
- taosx 2mo agoI created a simulation for coding harnesses based on my own pi sessions. When taking into account all factors, DS-v4-Pro is cheaper than gpt-5.6-luna due to caching. Look at the bill segments difference for cache read cost and uncached cost between deepseek and the other models. At this point is cheaper to use ds-v4-pro than the luna models from openai. ignore the numbers except the classic and keep in mind that classic is based on pi with the only change limiting tool output to 10kb https://harness.eveid.com/lazy-harness-cost-simulation https://harness.eveid.com/lazy-harness-cost-simulation * I built this for getting an initial estimate between different checkpoint/ compaction methods for the harness.
- RALaBarge 2mo agoHey this looks good! Maybe consider adding a hover-over popup for the rectangles explaining what each thing means to a lay person. I see it at the bottom, but that is below the fold.
- taosx 2mo agoDone, I'll take any other suggestions and apply them later, I will also split it a bit for different usecases as this was initially a throwaway prototype but found it useful. Basically it needs a bit more human touch.
- HDBaseT 2mo agoCan we have a conversation about subscription plans for a minute? I don't mean to hype up the US AI firms, but if a ChatGPT $200/m subscription can get you $16,000 in effective API costs, doesn't effectively every model get destroyed by the subsidized Claude/ChatGPT models? Both in price and intelligence.
- polski-g 2mo agoYeah pretty much. I spent half a billion in tokens one night on a huge refactor with DSFlash, cost $11. If I spent that every night it would be 3x my GPT subscription.
- nchmy 2mo agoThat seems too expensive to be honest. Did you do it with official deepseek api or a 3rd party provider? Because official has 10x cheaper cache reads than the rest. I've done similar sized chats for like $1
- HDBaseT 2mo agoThat pricing is going away soon. It's about to get 5x more expensive.
- xbmcuser 2mo agoThat is the problem currently the subscription plans are being subsidized by VC money and token buyers. When Open weight get good enough token buyers build their own servers instead of buying tokens then no one to subsidize the subscriptions
- adventured 2mo ago99.9%+ of the tech worker population will never be able to build their own servers to run future frontier models. Kimi 3 is an indication of what's coming. These models will keep getting drastically larger. The hardware isn't getting cheaper anytime soon (no matter what China does; that goes for memory and GPUs). Cycle forward to Fable 7, Kimi 5, GPT 7 a couple years out. Forget about it unless you own a datacenter.