6 ms·
Show HN: Claude-thermos keeps your Claude session warm for you
- alukin 2mo agoFeels like will be shut down real quick
- l1n 2mo agoWhy? Seems fine if people want to burn more cache reads.
- deleted 2mo ago[deleted]
- agluszak 2mo agoBecause it makes Anthropic earn less money
- cadamsdotcom 2mo ago"Earn"? Cache duration is arbitrary. What it actually does (if used en masse) is decrease the amount of oversubscription their infra can handle..
- janderson215 2mo agoEarn is synonymous with profit, and this hurts their profits. Introducing latency can benefit much more people and actually make usage more efficient. Reducing latency can sieze up a functioning system. Lately, I’ve been thinking about how this related to fractional banking. If you were to eliminate fractional banking introduced in the US by Hamilton, you would destroy a lot of current prosperity.
- sillysaurusx 2mo agoIt’s unknown whether our current economic model will work out long term. It mostly relies on the US being the default reserve currency for most nations. Hopefully that lasts a long time.
- munk-a 2mo ago> Because it makes Anthropic burn more money Fixed that for you.
- gogobio 2mo agoYes, this wastes cycle. They dump your cache and deallocate the VMs so that others can use it. This will result in tokens wasted and downtime for others.
- supern0va 2mo agoCan you expand on this? I can't imagine this is holding up an entire machine or materially impacting capacity. Isn't it common to have a tiered cache and to ultimately evict (or move to another tier) if there's an active inference request and no available capacity elsewhere?
- Lalabadie 2mo agoYes, however if a tool like this commoditizes a way to stay "higher" in the cache hierarchy to avoid getting dropped, then for the same total capacity, Anthropic have to become stricter about retention.
- supern0va 2mo agoIt's unfortunate that there isn't some better way to signal that there's a high chance that a turn is coming (ie, due to active sub-agents). If CC could send a ping saying "I don't need a turn, but please keep this cached at some tier because there will be a turn soon." that would probably help with cache efficiency. Instead, it seems like there's just the shotgun one-hour solution.
- Cyberdog 2mo agoAgreed, the Thermos company's lawyers will probably fire off a terse letter once they find out.
- jonas21 2mo agoIs the 5-minute expiration correct? I thought it was more like ~1 hour.
- s0ck_r4w 2mo agoYou can set the retention globally (whole session) to 1 hour which will actually make it more expensive. The default is 5 minutes. *UPD:* actually it appears the default is authentication-dependent. API key gets 5 minutes, subscriptions - 1 hour.
- foota 2mo agoClaiming that it's more expensive isn't true, it's workload dependent. Filling the cache (worst case) 12 times in an hour is much more expensive than caching it for an hour, and the cheapest is to have your 5m cache refilled by a ping like this is doing. Imo it's their fault for not having pricing that aligns incentives.
- j45 2mo agoThis is a good idea if it could manage it in an acceptable way. For example, there might be something I intended to complete in one sitting, but took two sittings in the same day unexpectedly. Maybe it could just be a few cache delays per day or something, tagged in advance somehow.
- szin 2mo agoOn a subscription Claude Code already defaults to 1h TTL and on an API key you opt in with ENABLE_PROMPT_CACHING_1H=1. Was the ~22% claimed in the repo is measured against the 5m default, or with the 1h flag on?
- unholiness 2mo agoI've directly inspected calls for pro/Max plans and as of today they have 1hr cache expiries. This has definitely degraded to 5 min in the past but that's the behavior today. If you're paying API rates, you can choose 5m or 1hr yourself (and pay different rates). Keeping a 1hr cache warm could still be useful, sure, but outside that, I don't see much use of this today.
- s0ck_r4w 2mo agoWhen using API keys the default is still 5 minutes. Setting 1 hour for the entire session is actually a lot more wasteful considering the 2x rate that it comes with
- razodactyl 2mo agoDon't we pay for cache input though?
- cortesoft 2mo agoThe idea is this keeps it in cache so you don’t have to pay to re-input.
- broodbucket 2mo agoYou can send and receive 1 byte and refresh the cache. Keeping the cache alive isn't free, but it's close.
- leemoore 2mo agoif you are refreshing a cache of 120k tokens, you have to input the same 120k tokens. you can't bounce the cache with 1 byte. A proper cache hit requires your entire context that is cached. if you change 1 byte in it, everything after that byte is a cache write and no longer a cache hit
- broodbucket 2mo agoYes, you're paying the cache read, not the cache write, which is much more expensive. You can do the math for how many 5min refreshes you can be afk for until it starts costing instead of saving
- ATMLOTTOBEER 2mo agoGlad this exists. It will force anthropic to fix their flawed cache mechanism.
- cortesoft 2mo agoHow should they fix it?
- s0ck_r4w 2mo agoMake the cache TTL more adaptive. Have more tiers than just 5m and 1h. Long-running subagents become increasingly ubiquitous. There's no reason for Anthropic to not do a better job for this use case.
- addaon 2mo agoThere are a lot of workflows where cache is likely to be consumed exactly once (e.g. anything chat-like where a single thread is extended; if a prefix cache hit is found, and a set of new longer prefixes is inserted in the cache, it is unlikely that the original shorter prefix will be consumed in later turns). I could see allowing each session key a small number (1?) of 24-hour cache entries, where inserting a new cache entry (perhaps the maximum-length prefix, perhaps an API-tagged specified-length prefix) consumes that slot and either boots the rest of the items from cache, or demotes them to 5 minutes or something. Basically: workflow awareness, not uniform handling.
- randomblock1 2mo agoDefault to 1h. Allow setting it to longer or shorter, granularly. Add /pause to mark it for eviction, for a token refund.
- SwellJoe 2mo agoWhile Anthropic is far behind on caching and efficiency (made more dramatically apparent by how much cheaper it is to use GPT 5.6 Sol at API rates than even Opus 4.8, much less Fable), a bunch of people forcing their way to the front of the queue at the expense of everyone else isn't going to solve that.
- cosmotic 2mo agoHow will this not lead to tragedy of the commons?
- davesque 2mo agoExactly. I really wish people wouldn't use this. If this becomes popular, anthropic will just modify their cache policy to be much less fair. It's not like they have infinite cache.
- Wowfunhappy 2mo agoWell, it's not a free lunch, you're still paying for the cache-warming request. Most of that request will be cached by definition, but cached requests merely cost less, they aren't free.
- SwellJoe 2mo agoThis is just making it more expensive for everyone else, right? How Claude handles its sessions is none of my business. I'm going to let them do the best they can to provide good service for everyone, and if they can't/won't, I'll switch to a provider that can. Using these massive models is already pretty danged extravagant, I'm not going to demand to be at the front of the queue at all times, too.
- s0ck_r4w 2mo agoHow do a few extra requests with the same prefix make it more expensive for everyone else?
- sznio 2mo agoyou use more memory?
- SwellJoe 2mo agoKeeping a conversation with a very large model active requires hundreds of GB of memory. If my conversation can never be swapped out, like when I go to lunch or take an afternoon walk, that's several hundred GB not available for other users. If everybody does that, Anthropic needs even more infrastructure than the quarter trillion dollars in infra they're already using.
- boc 2mo agoInterested in how the critics of approaches like this defend an agentic session (with Fable, for example) that stops and runs a multi-hour ML training session. It's a script, so the actual LLM convo goes stale, but then when the results get returned to the main thread you get an expensive cache hit without doing anything. You would have avoided that cache hit if the LLM session was kept "alive" for those few hours. Why not automate the part where you keep the large main thread alive until you're ready to analyze the results?
- PcChip 2mo agoI think you mean miss?
- jaimehrubiks 2mo agoThey should increase the cache to 10 minutes. 5 is just too low, you can even miss it by taking time to select a response from a question.
- gabigrin 2mo ago10 mins would be amazing
- purpleidea 2mo agoI assumed (perhaps incorrectly, but it was a guess since I never dug into it) that less used "hot pockets" of previous inference gradually got more stale as time went on, and the conversation went elsewhere and didn't need those bits. Hearing one byte refreshes the whole thing is huge! 5min is wayy too slow, because sometimes I want to spent more than 5 min looking at a diff before choosing where to go next. Kind of outrageous, I hope this kind of feature gets built into claude code =D
- 2001zhaozhao 2mo agoI thought the cache length was 1 hour, not 5 minutes?
- leemoore 2mo agoit could be either. subscription defaults to 60 min with 2x cache writes. extra usage and api defaults to 5 min with 1.25 cache writes. You can override either with settings
- edot 2mo agoNice! Codex’s TTL on 5.6 is, I believe, 30 minutes, and when combined with how generous they are with resets of weekly limits and having completely gotten rid of 5hr limits, means Claude Code is just a total rip off right now. Not to mention their silent Fable nerfing and Chinese fearmongering.
- Wowfunhappy 2mo agoFYI, on Pro and Max plans caching lasts for one hour, not five minutes, unless you're currently using Extra Usage. https://code.claude.com/docs/en/prompt-caching#on-a-claude-subscription https://code.claude.com/docs/en/prompt-caching#on-a-claude-s... (Thank you to EliasWatson for giving me this link just a few days ago, as I was previously confused too.)
- skeledrew 2mo agoYeah I do this manually by compacting the convo. The TTL is about an hour though, definitely not 5 minutes.
- leemoore 2mo agoIf you're running the subscription, by default you are paying 2x for cache writes and you're getting an hour for expiry. So refreshing based on 5 min is wasteful. You need to detect whether you are in 5 min or 1 hour mode.
- sublinear 2mo agoYes! Accelerationism cuts the other way too.
- IrfanD 2mo ago[dead]
- smokeeaasd 2mo ago[flagged]
- hamza_ali_shah 2mo agothis is so cool. claude should actually opensource part of their harness so the community can improve it context windows and sessions already get maxed out faster with fable and opus 5 and cost a ton. this should lead to significant savings. checking it out *(and ingesting into mer personal ai builder /hamzaish. opensource)