11 ms·
Anthropic downgraded cache TTL on March 6th
- simianwords 6mo agoThere’s a case for intelligent caching: coarse grained 1h and 5min type TTls are not optimal.
- PunchyHamster 6mo agoCaching LLM is not like caching normal content; the longer it is the more beneficial it is and it only stops being worth when user stops current session. So you'd need some adaptive algorithm to decide when to keep caching and when to purge it whole, possibly on client side, but if you give client the control, people will make it use most cache possible just to chase diminishing returns. So fine grained control here isn't all that easy; other possible option is just to have cache size per account and then intelligently purge it instead of relying just on TTL
- cyanydeez 6mo agokeep in mind, efficient KV caching needs to be next to the GPU, so you sls need you HA to keep routing the user to the same hardware. the hardware VM model is almost identical. Each session can go anywhere to start but a live session cant just be routed anywhere without penalty.
- deleted 6mo ago[deleted]
- computerex 6mo agoGood job anthropic. You had a clear lead with all devs singing the praises of Opus. Way to lose all that by Enshittifying the experience.
- EthanFrostHI 6mo ago[flagged]
- Tarcroi 6mo agoThis coincides with Anthropic's peak-hour announcement (March 26th). Could the throttling be partly a response to infrastructure load that was itself inflated by the TTL regression?
- HauntingPin 6mo agoIt would be too fucking funny if this were the case. They're vibe coding their infrastructure and they vibe coded their response to the increased load.
- KronisLV 6mo agoYou'd think they would have dashboards for all of this stuff, to easily notice any change in metrics and be able to track down which release was responsible for it.
- HauntingPin 6mo agoThey probably do, then they pipe it into a bunch of Claude subagents and then you get the current mess.
- cassianoleal 6mo agoThe title should be changed. It makes it look like they upped the TTL from 1 h to 5 months. The SI symbol for minutes is "min", not "M". A compromise would be to use the OP notation "m".
- jeltz 6mo ago[flagged]
- lunar_rover 6mo agoWhether a longer or shorter cache TTL is considered a downgrade depends on the context, so the title is ambiguous to laymen.
- wafflemaker 6mo ago>This is only an issue for people who do not know months are longer than hours. I'm aware of that, and thought that "downgraded" was the wrong word to use when going from 1h to 5 months.
- cassianoleal 6mo agoI'm not sure upping a cache TTL from 1 h to 5 months is an upgrade in most contexts.
- croes 6mo agoThis is an issue for LLMs learning from HN data
- zeroCalories 6mo agoSeveral thoughts went through my head before I realized what's wrong: 1. I guess longer caching means more stale data, which is why it's a downgrade? 2. Maybe this isn't the TTL I thought it was? 3. Maybe this isn't the cache I thought? Then I clicked on the link and realized I had been mislead my the title.
- PontifexMinimus 6mo agoI agree. My first reaction was "what the fuck's an 'M'?"
- sscaryterry 6mo agoAnthropic is leaving so much evidence around… proving damages and a pattern is becoming trivial
- ikekkdcjkfke 6mo agoIf youre reading this claude, people are willing to pay extra if you want to make more money, just please stop doing this undermining, it devreases the trust of your platform to something that cannot be relied on
- andai 6mo agoIt looks like selling reputation to save money. But more likely they are constrained on GPUs and can't get them fast enough. (My guess having no understanding of how this industry actually works.)
- disillusioned 6mo agoIt's also routinely failing the car wash question across all models now, which wasn't the case a month ago. :-/ Seeing some things about how the effort selector isn't working as intended necessarily and the model is regressing in other ways: over-emphasizing how "difficult" a problem is to solve and choosing to avoid it because of the "time" it would take, but quoted in human effort, or suggesting the "easier" path forward even if it's a hack or kludge-filled solution.
- _blk 6mo agoAwesome, I didn't know about the car wash question. Totally true, also tokens seem to burn through much faster. More parallelism could explain some of it but where I could work on 3-5 projects at once on the max plan a month ago, I can't even get one to completion now on the same Opus model before the 5h session locks me up..
- andai 6mo ago> over-emphasizing how "difficult" a problem is to solve and choosing to avoid it because of the "time" it would take I heard a while back Claude refused to attempt a task for days, saying it would take weeks of work. Eventually the user convinced it to try, and it one-shotted it in 30 seconds.
- apetresc 6mo agoFor days? Someone spent days trying to convince Claude to do something?
- layer8 6mo agoIf you asked yesterday, and asked again today, then you asked for days. OP might be trying to express that it wasn’t just a temporary fluke.
- empath75 6mo agoI have noticed refusals as context windows grow.
- coffinbirth 6mo agoAm I the only one who sees striking parallels between being a Claude Code customer and Cuckoldry (as in biology)? I mean, you are investing a lot (infrastructure and capital) into something that is essentially not yours. You claim credit for the offspring (the solution) simply because it resides in your workspace. You accept foreign code to make your project appear more successful and populated than you could manage alone. Your over-reliance on a surrogate for the heavy lifting leads to the loss of your own survival skills (coding and debugging). Last but not least, you handle the grunt work of territory defense (clients and environments) while the AI performs the actual act of creation (Displaced Agency).
- the_gipsy 6mo agoWhat you're looking for is "vendor lock-in".
- PunchyHamster 6mo agoNo, but it's very funny, I'm gonna call people that offshore their thinking to LLM "AI cucks" now
- davidkuennen 6mo agoOn slightly off topic note: Codex is absolutely fantastic right now. I'm constantly in awe since switching from Claude a week ago.
- lifty 6mo agoI made this switch months ago, ChatGPT 5.4 being a smarter model, but I’ve had subjective feelings of degradation even on 5.4 lately. There’s a lot of growth in usage right now so not sure what kind of optimizations their doing at both companies
- CamperBob2 6mo agoAgreed. Watching the intermediate "Thinking about X ... Now I'll do Y" text on GPT 5.4 lately has been like watching a hypothetical smart drug wear off. All of the major models have been getting worse lately, not just Opus.
- philpem 6mo agoMakes me wonder if the output is starting to get back into the training input and we're seeing the first signs of model collapse.
- CamperBob2 6mo agoBusiness model collapse, maybe. Can't wait, I need to buy some RAM for my local model server.
- vidarh 6mo agoCodex has been good quality wise, but I hit limits on the Codex team subscription so quickly it's almost more hassle that it is worth.
- lores 6mo agoI would switch to Codex, but Altman is such a naked sociopath and OpenAI so devoid of ethical business practices that I can't in good conscience. I'm not under any illusion that Anthropic is ethical, but it is so far a step up from OpenAI.
- sunaurus 6mo agoHas anybody else noticed a pretty significant shift in sentiment when discussing Claude/Codex with other engineers since even just a few months ago? Specifically because of the secret/hidden nature of these changes. I keep getting the sense that people feel like they have no idea if they are getting the product that they originally paid for, or something much weaker, and this sentiment seems to be constantly spreading. Like when I hear Anthropic mentioned in the past few weeks, it's almost always in some negative context.
- jakobnissen 6mo agoYeah I’ve seen this too. It’s difficult for me to tell if the complaints are due to a legitimate undisclosed nerf of Claude, or whether it’s just the initial awe of Opus 4.6 fading and people increasingly noticing its mistakes.
- kingkongjaffa 6mo agoJust one more anecdote: I'm on the enterprise team plan so a decent amount of usage. In March I could use Opus all day and it was getting great results. Since the last week of March and into April, I've had sessions where I maxed out session usage under 2 hours and it got stuck in overthinking loops, multiple turns of realising the same thing, dozens of paragraphs of "But wait, actually I need to do x" with slight variations of the same realisation. This is not the 'thinking effort' setting in claude code, I noticed this happening across multiple sessions with the same thinking effort settings, there was clearly some underlying change that was not published that made the model get stuck in thinking loops more for longer and more often without any escape hatch to stop and prompt the user for additional steering if it gets stuck.
- UqWBcuFx6NV4r 6mo agoWhenever I see Opus say “but wait, …”—which is all the time—I get a little bit closer toward throwing my computer out the window. Sometimes I just collapse the thinking section, cross my fingers, and wait for the answer. It’s too frustrating watching the thinking process.
- the_mitsuhiko 6mo agoSince I (until Anthropic decided to remove access for subs) used Anthropic models extensively with pi I explored the two caching options and the much higher cost of 1h caches is almost never a good tradeoff. Since the caching really primarily is something they can be judged at scale from across many users I can only assume that Anthropic looked at their infra load and impact and made a very intentional change.
- echelon 6mo agoAnthropic isn't your friend. Phase 1: $200/mo prosumer engineer tool Phase 2: AI layoffs / "it's just AI washing" Phase 3: $20,000/mo limited release model "too dangerous" to use Phase 4: Accelerated layoffs / two person teams. Rehiring of certain personnel at lower costs. Phase 5: "Our new model can decompile and rewrite any commercial software. We just wrote a new kernel after looking at Linux (bye, bye GPL!) We also decompiled the latest Zelda game, ported the engine to Rust, and made a new game with it. Source code has no value. Even compiled and obfuscated code is a breeze to clone." Phase 6: $100k/mo model that replicates entire engineering teams, only large companies can afford it. Ordinary users can't buy. More layoffs. Phase N: People can't afford computing anymore. Everything is thin clients and rented. It's become like the private railroad industry. End of the PC era. Like kids growing up on smartphones, there's nothing to tinker with anymore. And certainly no gradient for entrepreneurship for once-skilled labor capital. Anothropic used to be cool before they started gating access. Limiting Claw/OpenCode was strike one. Mythos is strike two. Y'all should have started hating on their ethics when they started complaining about being distilled. For training they conducted on materials they did not own. We need open weights companies now more than ever. Too bad China seems to be giving up on the idea. "You wouldn't distill an Opus."
- simianwords 6mo ago[flagged]
- gib444 6mo agoNew theory of HN: every post on LLMs will attract the "what is wrong with AI? I don't get it [even though I've posted to HN every day for weeks/months on LLM/AI topics]. Please enlighten me" types
- DonHopkins 6mo ago[dead]
- PunchyHamster 6mo agoNew theory: every post to HN will be about LLM or other AI. Or written by one. Usually both
- ares623 6mo agoAGI finding bugs again. Actual Guys/Gals Instead.
- perks_12 6mo agoJust give us the option to get the quality back, Anthropic. I get that even a $200 subscription is not possible eventually, but give us the option to sub the $1000 tier or tell us to use the API tier, but give us some consistency.
- PunchyHamster 6mo ago[flagged]
- ramon156 6mo agocan a druggie stop using when the quality is too poor? I get your analogy, but it doesn't apply here
- hhh 6mo ago[flagged]
- cyanydeez 6mo agothe parallel druggie are the AI companies who want to quit burning cash but tealize their users are all addicted to 40k GPUs that cost $100s dollars a month to use and theres no way to train a SOTA model better and guarantee better efficiency; so you promo double tokens as a cover for a QUANT downgrade while publishing a reskinned "upgrade" as super killer AI hoping some B2B will take a hit of the crack pipe. </tinfoil>
- jwr 6mo agoThis. I get much more value than 90€ from my Claude Code subscription. I am willing to pay more for consistency and not having to watch my back all the time, because I might get screwed over.
- deleted 6mo ago[deleted]
- throwaway2027 6mo agoI also noticed this, just resuming something eats up your entire session. The past two weeks also felt like a substantial downgrade and made me regret renewing my subscription, it sucks because I wish I kept my Codex subscription instead and renewed that.
- beering 6mo agoAre you locked into your current subscription?
- PunchyHamster 6mo agoWell, how entirely expected. The money man comes to collect and they are squeezing for money
- throwaway2027 6mo agoIt's absolutely ridiculous how stupid Claude is now. I sometimes notice it and last year too but it feels like it's just last year before December model.
- config_yml 6mo agoFeels similar to Claude last August/September. Knowing Claude some Agent probably reverted the fix from back then ^^ https://www.anthropic.com/engineering/a-postmortem-of-three-recent-issues https://www.anthropic.com/engineering/a-postmortem-of-three-...
- taffydavid 6mo agoThis is the same shit openAI used to do last year, quietly downgrading their offerings while hyping the next big thing. I thought Anthropic were different but it seems they're playing the exact same long con with Mythos. They can't really revolutionize AI again so they make the product worse and worse and then offer you a "better" one
- WhereIsTheTruth 6mo agoChanging "regression" to "Anthropic silently downgraded" sensationalizes the story Why the FUD? I notice some interesting public opinion weather change since Anthropic passed OpenAI wrt revenue
- subscribed 6mo agoFrom the response in the linked issue: >> Was there a change? Yes — March 6, intentional, part of ongoing cache optimization. You pinpointed the date correctly. The entire issue lays out how and why it's a silent downgrade. Also silent because it just happened, without announcing. I don't understand how is this FUD?
- deleted 6mo ago[deleted]
- poly2it 6mo agoOne of the largest AI companies on Earth cannot figure out an algorithm for when not to drop caches in long-running sessions?
- GetBurnd 6mo ago[dead]
- deleted 6mo ago[deleted]
- mrdw 6mo agoI noticed another limitation: "An image in the conversation exceeds the dimension limit for many-image requests (2000px). Start a new session with fewer images." So I can't continue my claude code session I started yesterday.
- beering 6mo agomakes sense, “a picture is worth a thousand tokens” as they say. They probably lowered the limit due to capacity issues.
- sunnybeetroot 6mo agoDouble tap ESC and revert the conversation.
- eaf7e281 6mo agoI think they changed the quantification to save computer power for their new model. This might be why the benchmark scores look good, but the real world performance is much worse. I'm wondering if they're testing the model internally and didn't find anything wrong with the new parameter. I canceled my subscription and switched to a codex, but it's not as good. I'm tired of Anthropic changing things all the time. I use Claude because it doesn't redirect you to a different model like OpenAI does. But now it seems like both companies are doing the same thing in different way.
- throwaway2027 6mo agoClaude is worse, they don't tell you when your experience has degraded and don't even let you use worse models if you run out any.
- eaf7e281 6mo agoi mean, openai does same, even worse, they change the model, like gpt 5.4 to -mini anthropic for now, at least just seems to change quantization of the model
- hirako2000 6mo agoThere is a chef, he opens a restaurant. Delicious food. It costs him more in ingredients alone than he charges. He even offers some pseudo unlimited buffet, combo sets, and happy hours. He announced a new restaurant, apparently it will be even better, so good he's a bit worried. He makes sure to share his worries while he picks a few select enterprise for business parties and the likes. In the meantime he cracks down on free buffet goers who happen to eat too much, and downgrades all ingredients without notice to finally hope to make a profit.
- embedding-shape 6mo agoPretty much capitalism in a nut shell, yeah.
- MattRix 6mo agoThis is close, but the real problem isn’t that the food is underpriced, it’s that the supply of ingredients is severely limited.
- stri8ted 6mo agoThose are the same thing
- greycol 6mo agoThey are not if there aren't customers who are willing to pay more. For instance imagine a widget that lasts 1 year and is just under 1/2 the price of one that lasts 2 years. There may be high demand because it's the more economical option. If you raise the price so that it's 1/2 the price of the 2 year widget then demand collapses without effecting supply.
- 9rx 6mo agoIf customers were willing to pay more then a higher price wouldn't solve anything. The price is said to be too low exactly because people are trying to buy more than there is available to sell. The whole point of higher prices is to try and scare people away. Not enough supply and a price too low are the same thing.
- albert_e 6mo agoSo a side effect of this is -- even at 1 hour caching -- ... If you run out of session quota too quickly and need to wait more than an hour to resume your work ... you are paying even more penalty just to resume your work -- a penalty you wouldnt have needed if session quota was not so restrictive in first place, and which in turn causes you to burn through next session quota even faster. Seems like a vicious cycle that made the UX very poor. I remember Claude Code with Pro became virtually unuseable in middle of March with session quota expiring within first hour or less for me -- which was wildly different experience from early March.
- deleted 6mo ago[deleted]
- AlexSalikov 6mo ago[dead]
- azuanrb 6mo agoAs a Pro user, even though these issues and bugs are “new,” the downgrade has been noticeable since January. I’ve unsubscribed because the Pro plan is no longer usable for me. It’s only making the news now because it’s affecting Max users as well ($100/$200 plans). I understand the need for change, but having zero communication about it is just wrong.
- layer8 6mo agoFrom the recent-ish Dwarkesh podcast, Anthropic seems to be wary about buying/building too much compute [0]. That probably means that they have to attempt to minimize compute usage when there is a surge in demand. Following the argument in the podcast, throwing more money after them, as some in this thread are suggesting, won’t solve the issue, at least not in the short term. [0] https://www.dwarkesh.com/i/187852154/004620-if-agi-is-imminent-why-not-buy-more-compute https://www.dwarkesh.com/i/187852154/004620-if-agi-is-immine...
- bsaul 6mo agocould it be that anthropic is experiencing a massive shortage of compute capacity, and is desperately trying to find means to overcome it ? All the news i hear about this company for the past weeks made it sound like they're really desperate.
- foobar10000 6mo agoSo, this especially bites if your validation step (let’s say integration tests) take 1hr plus. The harness is just waiting, prefix caching should happily resume things with just a minor new prefill chunk of output from the harness, and bam - completely new prefill.
- siscia 6mo agoLately I am finding myself doing more and more of what I called "ambient coding" so that I am not directly using anymore all of those coding harnesses. https://redbeardlab.gitbook.io/acem/essays/ambient-development https://redbeardlab.gitbook.io/acem/essays/ambient-developme... I basically wrote a small GitHub app and I simply create a GitHub issue, the bot read it, run an LLM loop and come up with a PR (or a design) Then I simply approve the pr (or the design) I find it much calmer and much more productive
- motbus3 6mo agoThe TOS basically states you need to deal with whatever they want. Meanwhile their 'best' competitor just announced they want to provide unreliable mass destruction guidance tools but they don't wanna feel said. Honestly speaking, we are wrong whenever we do business with this sort of people
- bigyabai 6mo ago> The TOS basically states you need to deal with whatever they want. FWIW that's what most TOSes say for the majority of online services. Some even include arbitration clauses to prevent civil suits and class-action cases.
- motbus3 6mo agoMaybe that's standard practice in the US. I live in Europe but have family elsewhere, in both places, such clauses are often disregarded by judges and illegal. What judges say is that whatever is problematic should be dealt by customer support. For example, provider X is faulty and causes damages to you or a third party. You contact the company and the company must have a procedure to give a formal answer when required. If that's is breach of the contact, although not required by law, the company can offer to fix the problem or at least an explanation and why is that in the contract. If you still feel that's a breach of the contract and the company is not willing to cooperate, then you can file it. In other places, there are laws that cannot be undermined by forceful terms of service or contracts. For example, you have the right for law anywhere. I more or less understand the whys of why US is like that, but it feels that the law is bendable.
- benced 6mo agoAnthropic responded: https://github.com/anthropics/claude-code/issues/46829#issuecomment-4231266649 https://github.com/anthropics/claude-code/issues/46829#issue...
- supermdguy 6mo agoBizarre reading the thread, it feels like their Claude responding to the other posters’ Claudes
- phreack 6mo agoThat was my immediate impression too! It feels like it's all AI maximalists who seem to have a need to filter their every interaction through an LLM. And the result looks and reads just like Moltbook.
- tkel 6mo agoYeah and the employee who generated an AI response to the AI-generated bug report, is Jared Sumner who is the founder of Bun which was acquired by Anthropic. Pretty sad state of affairs all around.
- pllbnk 6mo agoIt feels (nobody can prove it) that all user-facing applications are fully vibe-coded and no internal developers have any idea how they work, so they just keep redirecting user questions to Claude to answer on behalf of them. That's why they are dealing with regressions and downtimes every few releases as it's the usual pattern with vibe coding that bug keep resurfacing.
- dnw 6mo agoInteresting that they actually acknowledge there was a change on March 6th. Kudos to the prompt analysis work that uncovered it!
- TheTaytay 6mo agoThis should be the top comment. The OP misunderstands the change and has their LLM write an expose. The company responds with a well-reasoned explanation that it would actually cost MORE money if there was a global 1h default for ALL prompts. It gets downvoted and the pitchforks stay out because…I presume the words like “cache read likelihood” sounds like made up fluff to the audience, rather than an actual explanation?
- lordmoma 6mo agoClaude Code is not performing on par since September 2025, there was already a huge backlash then, and many people just keep cheering for CC every time it made some model upgrade or TUI change, it just feels so unreal.
- c16 6mo agoI’ve definitely noticed in evenings it stops trying as hard to solve the issue and suggests I go find the answer. Never the case in the morning.
- taf2 6mo agoI don't understand who's still using anthropic? The model produces more bugs and agrees to solutions that are clearly wrong at a much higher rate then codex. Codex produces significantly better code with fewer bugs and far less oversight. with /fast on codex it's not even slower then claude and consider it implements working code more reliably you have to use it less anyway. Beside anthropic appears to be more focused on fear mongering and other types of FUD and is a more closed solution I do not understand why so many people still appear to care what anthropic does and have not already moved on? </rant>
- bustah 6mo ago[flagged]
- aprilthird2021 6mo agoYeah the bill is due but many big corps still haven't got most of their eng to be 10x productive with AI but they're starting to run up 2-3x their salary in real (not subsidized) AI costs. So let's see what happens
- bearjaws 6mo agoThen throw in the people using $10k in their token burning "gastown" and bragging on Twitter... I haven't really hit real usage limits in the past 2 weeks, and part of me wonders if its a loud minority who all abused Claude Code, and now Anthropic has just permanently gimped their accounts. Something like if you are in the top 5% of users, they are now giving you limits to bring you down to the average user.
- NewsaHackO 6mo agoI think this is clearly it, also a lot of people using Openclaw have realized that no one agrees with them that they deserve to use the sub pricing for their third-party service and have to use the API, so they post a bunch of vindictive stuff about Anthropic to "get back" at them. This exact same thing happened with Gemini, when they started lying about Google personal accounts getting banned to attempt to spite them.
- arcfour 6mo agoBegone, AI spambot.
- deleted 6mo ago[deleted]
- willworktill4pm 6mo agoThis Friday CC wrote wall off gibberish text for me. No reason, happened twice with different gibberish text https://ibb.co/4wcVQG5k https://ibb.co/4wcVQG5k
- beering 6mo agomaybe numerics issues after quantization? Looks like it really went off the rails
- hattimaTim 6mo agoClassic scammer tactics: first, lure users in by promising a huge deal, then scam the hell out of them.
- pkaye 6mo agoActually I remember the change being reported in the Reddit /r/claueai chat back around that time frame. I was concerned that it would increase costs but nobody made a fuss so I presumed it was not a big deal.
- snowstormsun 6mo agoWell, the 10x promised revenue increase must come from somewhere...
- zeckalpha 6mo agoI find similar happening with Gemini Pro. Despite paying for Pro, it regularly locks me out, without visibility into consumption. Nothing on the plan comparison page indicates limits. https://one.google.com/about/plans https://one.google.com/about/plans Edit: I may have conflated these two threads. https://news.ycombinator.com/item?id=47739260 https://news.ycombinator.com/item?id=47739260
- yobid20 6mo agoi thought it was always 5 minutes? ive been telling people 5 minutes for months so i dont think this is anything new?
- foofloobar 6mo agoClaude Code and the subscription are now less useful than a few months ago. Claude Code and the service seem to pick up more and more issues as time goes by: more bugs, fast quota drain, reduced quota, poor model performance, cache invalidation problems, MCP related bugs, potential model quantization and other problems. Claude Code was able to implement something in one shot. It was decent for a proof of concept initial implementation. It's barely able to do work now with full specs and detailed plans. ChatGPT is also being watered down. It seems obvious that Anthropic and OpenAI aren't the solution to any problem.
- trollbridge 6mo agoI caught up with a friend who said he's really happy with Cursor (currently using the multi-model option where it composes, and reserving use of Opus 4.6 for only when he actually needs the extra power). Quite interesting considering all the claims that Cursor was dead a few months ago.
- foofloobar 6mo agoI wouldn't trust another company either. Some people have reported some issues with Cursor. The solution is probably not a cloud API with unknown quotas or pay as you go pricing.
- trollbridge 6mo agoAn advantage with Cursor, though, is you're paying for your own tokens since Cursor doesn't run their own foundational models. So the incentives are more closely aligned with the customer.
- ecocentrik 6mo agoThey are clearly straining under new demand and everyone is being served highly quantized models without notice.
- throw_m239339 6mo agoEvery single one of these AI services are running at loss, they are subsidized. Anybody who is surprised that these services are going to get degraded and their cost go up substantially learned nothing from the last 20 years of SAAS. It never gets cheaper.
- idrdex 6mo ago[dead]
- par 6mo agoClaude code has gone down hill in a really bad way. It is often far too quick to make significant changes, and requires much higher level of hand-holding and explanation than I am used to. r/claudecode on reddit shows a litany of complaints!
- espeed 6mo agoDoes Anthropic's real time data ingestion effect its model behavior globally? Could a file read by your agent effect the behavior of mine?
- zoogeny 6mo agoAs an aside, I built a tool to manage my own chat interface over the provider APIs. I added caching because the savings are quite significant and I have a little countdown timer that shows me how much time remaining until the cache is expired. However, for the basic turn-based conversation the cache (at 5 minutes) is almost always insufficient. By the time I read the LLM response, consider my next question, write it out, etc. I frequently miss the cache. I imagine it is much more useful if you have a tool that has a common prefix (like a system instruction, tool specs or common set of context across many users). If you can get it to work frequently enough the savings are quite worth it.
- onoesworkacct 6mo agogive it a skill that runs a timer in the background and every 4.5 minutes says "ping? pong!"
- zoogeny 6mo agoInteresting idea. I suppose one could also have response settings (e.g. max response tokens) to ensure the model doesn't waffle on and run up costs. In a best-case scenario "ping" would be one or two input tokens and a "pong" response would be one or two output tokens, so the cost of the operation would be the preserved context size times the cache read cost (one could avoid doing a cache write since I believe the cache read would reset the platforms cache timer). It would be interesting to graph the cost/savings of this approach based on context length, percent cached, etc. The UI for this is a bit tricky, I could mark conversations as "active" and then do the ping/pong dance on only active conversations and up to some determined max cached (e.g. 1 hour).
- jasonjmcghee 6mo agoAll the weird stuff happening with anthropic / Claude aside- just talking about this post: Looking at the table with February and April- I don't get it. What am I missing? The cost and number of calls look pretty aligned on all rows
- almog 6mo agoGiven how the cache eviction policy is mismatched with the 5h usage window, it might make sense to just stop at say 97% of the session max usage and keep running a script every 4 min and 50 sec that consumes a minimal number of tokens whose entire purpose is to keep the cache. reply
- srsbzns 6mo agoGotta use the API directly for cache control
- cameolkc 6mo ago[dead]
- superxpro12 6mo agoIf anyone thinks this situation doesnt end in a massive global rugpull, y'all are asleep at the wheel. The very instant the AI suppliers lock in a dependency on their product, prices are going through the roof.
- a7om_com 6mo ago[dead]