5 ms·
I'm surprised more people aren't talking about the cache hit price: $0.003 per million tokens. I have a feeling that the price of 1 million tokens transmitted o
by k9294 23d ago
I'm surprised more people aren't talking about the cache hit price: $0.003 per million tokens. I have a feeling that the price of 1 million tokens transmitted over the internet is more expensive than cache hit. Are we close to making the chat completion API obsolete because the cost of context transfer over network is going to dominate the task total cost?
Here's the same token usage priced at different rates: a real long-running coding task, medium codebase, 447 turns.
Input 1,026,957
Output 164,667
Cache read 36,554,368
GPT-6-astra
Type Rate Cost Share
Input 10.000 10.270 19%
Output 50.000 8.233 15%
Cache 1.000 36.554 66%
Total 55.057 100%
DeepSeek v4.1 Flash, $0.003 cache hit
Type Rate Cost Share
Input 0.300 0.308 50%
Output 1.200 0.198 32%
Cache 0.003 0.110 18%
Total 0.615 100%
DeepSeek v4.1 Flash, $0.006 cache hit
Type Rate Cost Share
Input 0.300 0.308 42%
Output 1.200 0.198 27%
Cache 0.006 0.219 30%
Total 0.725 100%
Hypothetical: same DeepSeek input/output rates, but cache priced so it accounts for 66% of the bill.
Type Rate Cost Share
Input 0.300 0.308 21%
Output 1.200 0.198 13%
Cache 0.027 0.982 66%
Total 1.487 100%
This cache it improvement makes the model x2-x2.5 more efficient on a long horizon tasks in terms of cost.
- cbg0 23d ago$0.003 off-peak, not 0.003 cents.
- k9294 23d agoYep, but even 0.006 is quite a big improvement. I'm curious now to test the model on some token-heavy tasks, like code exploration before a coding session, to see whether it will decrease the total cost of the task in the end or not.
- czottmann 23d agoI think they meant it's 0.3c (= $0.003), not 0.003c.
- ForHackernews 23d agohttps://verizonmath.blogspot.com/ https://verizonmath.blogspot.com/
- deleted 23d ago[deleted]
- ponyous 23d ago> I have a feeling that the price of 1 million tokens transmitted over the internet is more expensive than cache hit. And this kinda makes sense. What is cheaper few KB of disk space or internet bandwidth?
- k9294 23d ago100%, but this means we are going to move to stateful APIs on the AI provider's end (like OpenAI already does with Codex and Responses API) to make this work.
- mmastrac 23d agoI ran the preview model around 2,126,605,070 tokens for $22.04 USD for the last couple of days. Kind of shocked. It did a decent job refactoring https://github.com/mmastrac/diffgemma https://github.com/mmastrac/diffgemma to create a CUDA support backbone, it's struggling a bit to port metal kernels to CUDA unattended (it hasn't managed to get numbers to match over >1 layer). It successfully ported a root exploit to an older Android phone that GLM5.3Flash and DSv4Flash were struggling a bit on, though I didn't start it from scratch and it picked up some of their work. FWIW it feels like a slightly-north of Opus 4.8 model, not quite Opus 5, not fable. It thinks in circles far less than DSv4F. The API version is insanely fast - was getting ~400 tok/s at times.
- dgacmu 23d agoBandwidth is really cheap in bulk. You can get a 100 gigabit internet connection for about $10,000 a month. If you were to somehow keep that saturated 24/7, you'd move about 30 petabytes in a month, so your per gigabit cost is only $0.0003. Realistically, if you had it 5% utilized, those million tokens would cost you about $0.0000062, which is pretty insignificant compared to what they charge you. (Assuming one byte per token, ignoring compression)
- sroussey 23d agoPeople read the AWS rate card for bandwith and think has something to do with reality. Even though it is 1000x higher!
- Onavo 23d agoMost web devs have never heard of Colo unfortunately, they only know Vercel and AWS. Hurricane electric should sponsor more booths at colleges. If they give out more swag maybe the millenials and gen Zs would finally understand bandwidth pricing.
- piyh 21d agoMy friend's vibe coded vercel site got hit by Meta for 21 million page views in 2 days. It cost him over $300. My self hosted compose stack running in my basement with two 9's of uptime was a 1 time cost of $600 between cat6e, refurb mini PCs and tons of time prompting for NixOS flakes that met my needs. I'm not sure who came out ahead.
- k9294 22d agoIt's not that I haven't heard of this, it's the reality that you have these constraints when you build applications on modern infrastructure. And let's face it, most of the applications use this infrastructure with these crazy prices for egress.