3 ms·
Sonnet : Cache Read $0.30 Gemini 3.5 flash : Cache Read $0.15
by maxdo 5mo ago
Sonnet :
Cache Read
$0.30
Gemini 3.5 flash :
Cache Read
$0.15
- maxdo 5mo agoGPT 5.4 Cache Read ≤272K $0.25 And it's multi modal, and available at whatever you might imagine rates limits.
- minimaxir 5mo agoFor Sonnet, that's 10% of input cost (and requires paying for the cache) For Gemini 3.5 Flash, it's also 10% of input cost. Which is why 2%/0.8% change the economics in a meaningful way, given the input/cache-heavy way agents operate.
- throwdbaaway 5mo agoAnd their disk-based caching is amazing. I got a long 700k context session spanning more than a week, with pauses in between that was longer than a day, and some rewinds mixed in as well. Stats from pi: ↑400k ↓438k R432M 71.9%/1.0M Half a billion tokens, $2.12
- kingstnap 5mo agoAnthropic's caching requires you to pay a $0.75/Mtok for Sonnet and $1.25/MTok for Opus as a surcharge on top of the original input token cost. It's not even automatic. If you are reading ~8 times (8 total back and forth tool calls) that means that cache reads in some sense cost ~$0.4 / M toks (Amortizing the write surcharge over all reads). It's really quite ridiculously expensive considering what you are paying for is some residence on a VRAM that sometimes gets offloaded to NVMe.