7 ms·
> I was never under the impression that gaps in conversations would increase costs nor reduce quality. Both are surprising and disappointing. You didn't do you
by computably 5mo ago
> I was never under the impression that gaps in conversations would increase costs nor reduce quality. Both are surprising and disappointing.
You didn't do your due diligence on an expensive API. A naïve implementation of an LLM chat is going to have O(N^2) costs from prompting with the entire context every time. Caching is needed to bring that down to O(N), but the cache itself takes resources, so evictions have to happen eventually.
- solarkraft 5mo agoI somewhat disagree that this is due diligence. Claude Code abstracts the API, so it should abstract this behavior as well, or educate the user about it.
- mpyne 5mo ago> Claude Code abstracts the API, so it should abstract this behavior as well, or educate the user about it. Does mmap(2) educate the developer on how disk I/O works? At some point you have to know something about the technology you're using, or accept that you're a consumer of the ever-shifting general best practice, shifting with it as the best practice shifts.
- zem 5mo agommap(2) and all its underlying machinery are open source and well documented besides.
- mpyne 5mo agoThere are open-source and even open-weight models that operate in exactly this way (as it's based off of years of public research), and even if there weren't the way that LLMs generate responses to inputs is superbly documented. Seems like every month someone writes up a brilliant article on how to build an LLM from scratch or similar that hits the HN page, usually with fancy animated blocks and everything. It's not at all hard to find documentation on this topic. It could be made more prominent in the U/I but that's true of lots of things, and hammering on "AI 101" topics would clutter the U/I for actual decision points the user may want to take action upon that you can't assume the user already knows about in the way you (should) be able to assume about how LLMs eat up tokens in the first place.
- websap 5mo agoDoes using print() in Python means I need to understand the Kernel? This is an absurd thought.
- Nevermark 5mo agoThat might be an absurd comparison, but we can fix that. If you were being charged per character, or running down character limits, and printing on printers that were shared and had economic costs for stalled and started print runs, then: You wouldn’t “need” to understand. The prints would complete regardless. But you might want to. Personal preference. Which is true of this issue to.
- Barbing 5mo ago>If you were being charged per character, or running down character limits, and printing on printers that were shared and had economic costs for stalled and started print runs, and the system was being run by some of the planet’s brightest people whose famous creation is well known to disseminate complex information succinctly, >then: You would expect to be led to understand, like… a 1997 Prius. “This feature showed the vehicle operation regarding the interplay between gasoline engine, battery pack, and electric motors and could also show a bar-graph of fuel economy results.” https://en.wikipedia.org/wiki/Toyota_Prius_(XW10) https://en.wikipedia.org/wiki/Toyota_Prius_(XW10)
- redsocksfan45 5mo ago[dead]
- computably 5mo agoI would say this is abstracting the behavior.
- someguyiguess 5mo agoYes. It’s perfectly reasonable to expect the user to know the intricacies of the caching strategy of their llm. Totally reasonable expectation.
- coldtea 5mo agoIt's not like they have a poweful all-knowing oracle that can explain it to them at their dispos... oh, wait!
- esafak 5mo agoThey have to know that this could bite them and to ask the question first.
- nixpulvis 5mo agoI do think having some insight into the current state of the cache and a realistic estimate for prompt token use is something we should demand.
- switchbak 5mo agoIf there was an affordance on the TUI that made this visible and encouraged users to learn more - that would go a long way.
- jghn 5mo agoTo some extent I'd say it is indeed reasonable. I had observed the effect for a while: if I walked away from a session I noticed that my next prompt would chew up a bunch of context. And that led me to do some digging, at which point I discovered their prompt caching. So while I'd agree with your sarcasm that expecting users to be experts of the system is a big ask, where I disagree with you is that users should be curious and actively attempting to understand how it works around them. Given that the tooling changes often, this is an endless job.
- 5mo ago
- doesnt_know 5mo agoHow do you do "due diligence" on an API that frequently makes undocumented changes and only publishes acknowledgement of change after users complain? You're also talking about internal technical implementations of a chat bot. 99.99% of users won't even understand the words that are being used.
- tempest_ 5mo agoI use CC, and I understand what caching means. I have no idea how that works with a LLM implementation nor do I actually know what they are caching in this context.
- hakanderyal 5mo agoCC can explain it clearly, which how I learned about how the inference stack works.
- libraryofbabel 5mo agoThey are caching internal LLM state, which is in the 10s of GB for each session. It's called a KV cache (because the internal state that is cached are the K and V matrices) and it is fundamental to how LLM inference works; it's not some Anthropic-specific design decision. See my other comment for more detail and a reference.
- dlivingston 5mo agoWhat is being discussed is KV caching [0], which is used across every LLM model to reduce inference compute from O(n^2) to O(n). This is not specific to Claude nor Anthropic. [0]: https://huggingface.co/blog/not-lain/kv-caching https://huggingface.co/blog/not-lain/kv-caching
- computably 5mo ago> How do you do "due diligence" on an API that frequently makes undocumented changes and only publishes acknowledgement of change after users complain? 1. Compute scaling with the length of the sequence is applicable to transformer models in general, i.e. every frontier LLM since ChatGPT's initial release. 2. As undocumented changes happen frequently, users should be even more incentivized to at least try to have a basic understanding of the product's cost structure. > You're also talking about internal technical implementations of a chat bot. 99.99% of users won't even understand the words that are being used. I think "internal technical implementation" is a stretch. Users don't need to know what a "transformer" is to understand the trade-off. It's not trivial but it's not something incomprehensible to laypersons.
- raron 5mo agoHow big this cached data is? Wouldn't it be possible to download it after idling a few minutes "to suspend the session", and upload and restore it when the user starts their next interaction?
- cyanydeez 5mo agoI often see a local model QWEN3.5-Coder-Next grow to about 5 GB or so over the course of a session using llamacpp-server. I'd better these trillion parameter models are even worse. Even if you wanted to download it or offload it or offered that as a service, to start back up again, you'd _still_ be paying the token cost because all of that context _is_ the tokens you've just done. The cache is what makes your journey from 1k prompt to 1million token solution speedy in one 'vibe' session. Loading that again will cost the entire journey.
- throwdbaaway 5mo agoShould be about 10~20 GiB per session. Save/restore is exactly what DeepSeek does using its 3FS distributed filesystem: https://github.com/deepseek-ai/3fs#3-kvcache https://github.com/deepseek-ai/3fs#3-kvcache With this much cheaper setup backed by disks, they can offer much better caching experience: > Cache construction takes seconds. Once the cache is no longer in use, it will be automatically cleared, usually within a few hours to a few days.
- nl 5mo ago> upload and restore it when the user starts their next interaction The data is the conversation (along with the thinking tokens). There is no download - you already have it. The issue is that it gets expunged from the (very expensive, very limited) GPU cache and to reload the cache you have to reprocess the whole conversation. That is doable, but as Boris notes it costs lots of tokens.
- vanviegen 5mo agoYou're quite confidently wrong! :-) The kv-cache is the internal LLM state after having processed the tokens. It's big, and you do not have it locally.
- margalabargala 5mo agoOkay, sure. There's a dollar/intelligence tradeoff. Let me decide to make it, don't silently make Claude dumber because I forgot about a terminal tab for an hour. Just because a project isn't urgent doesn't mean it's not important. If I thought it didn't need intelligence I would use Sonnet or Haiku.
- pixl97 5mo ago"Gets mad because their is no option" "Gets mad because when their is options the defaults suck" "Gets mad because the options start massively increasing costs to areospace pricing"
- margalabargala 5mo agoDid you mean to reply to someone else? Or do you misunderstand the issue? There is no option to avoid auto-dumbing after one hour of idle. I haven't complained about the cost at all, I'm happy to pay it. So yeah, I'm mad because there's no option. The other two you mentioned don't apply.
- kovek 5mo agoWhat if the cache was backed up to cold storage? Instead of having to recompute everything.
- vanviegen 5mo agoThey probably already do that. But these caches can get pretty big (10s of GBs per session), so that adds up fast, even for cold storage.
- kovek 5mo ago10s of GBs? ( 1,000,000 context * 1,000 vector size ) ^ 2 = 1,000,000,000,000,000,000… oh wow.. I must be miscalculating What about only storing the conversation and then recomputing the embeddings in the cache? Does that cost a lot? Doing a lot of matrix multiplication does not cost dollars of compute, especially on specialized hardware, right?
- Majromax 5mo agoContext length 1e6, vector length 1e3, and 1e2 model layers for 100e9 context size. Costs will go up even more with a richer latent space and more model layers, and the western frontier outfits are reasonably likely to be maximizing both.
- kang 5mo agoIt seems you haven't done the due diligence on what part of the API is expensive - constructing a prompt shouldn't be same charge/cost as llm pass.
- coldtea 5mo agoIt seems you haven't done the due diligence on what the parent meant :) It's not about "constructing a prompt" in the sense of building the prompt string. That of course wouldn't be costly. It is about reusing llm inference state already in GPU memory (for the older part of the prompt that remains the same) instead of rerunning the prompt and rebuilding those attention tensors from scratch.
- kang 5mo agoYou not only skipped the diligence but confused everyone repeating what I said :( that is what caching is doing. the llm inference state is being reused. (attention vectors is internal artefact in this level of abstraction, effectively at this level of abstraction its a the prompt). The part of the prompt that has already been inferred no longer needs to be a part of the input, to be replaced by the inference subset. And none of this is tokens.
- coldtea 5mo ago>It seems you haven't done the due diligence on what part of the API is expensive - constructing a prompt shouldn't be same charge/cost as llm pass. I think you missed what the parent meant then, and the confusing way you replied seemed to imply that they're not doing inference caching (the opposite of what you wanted to mean). The parent didn't said that caching is needed to merely avoid reconstructing the prompt as string. He just takes that for granted that it means inference caching, to avoid starting the session totally new. That's how I read "from prompting with the entire context every time" (not the mere string). So when you answered as if they're wrong, and wrote "constructing a prompt shouldn't be same charge/cost as llm pass", you seemed to imply "constructing a prompt shouldn't be same charge/cost as llm pass [but due to bad implementation or overcharging it is]".
- miroljub 5mo agoThis sounds like a religious cult priest blaming the common people for not understanding the cult leader's wish, which he never clearly stated.
- computably 5mo agoA strange view. The trade-off has nothing to do with a specific ideology or notable selfishness. It is an intrinsic limitation of the algorithms, which anybody could reasonably learn about. Sure, the exact choice on the trade-off, changing that choice, and having a pretty product-breaking bug as a result, are much more opaque. But I was responding to somebody who was surprised there's any trade-off at all. Computers don't give you infinite resources, whether or not they're "servers," "in the cloud," or "AI."
- miroljub 5mo agoHe was surprised because it was not clearly communicated. There's a lot of theory behind a product that you could (or could not) better understand, but in the end, something like price doesn't have much to do with the theoretical and practical behavior of the actual application.
- exac 5mo agoIt is more useful to read posts and threads like this exact thread IMO. We can't know everything, and the currently addressed market for Claude Code is far from people who would even think about caching to begin with.
- bontaq 5mo agoHow's that O(N^2)? How's it O(N) with caching? Does a 3 turn conversation cost 3 times as much with no caching, or 9 times as much?
- jannyfer 5mo agoI’m not sure that it’s O(N) with caching but this illustrates the N^2 part: https://blog.exe.dev/expensively-quadratic https://blog.exe.dev/expensively-quadratic
- bontaq 5mo agoIf there was an exponential cost, I would expect to see some sort of pricing based on that. I would also expect to see it taking exponentially longer to process a prompt. I don't believe LLMs work like that. The "scary quadratic" referenced in what you linked seems to be pointing out that cache reads increase as your conversation continues? If I'm running a database keeping track of a conversation, and each time it writes the entire history of the conversation instead of appending a message, are we calling that O(N^2) now?
- atq2119 5mo agoYes, that is indeed O(N^2). Which, by the way, is not exponential. Also by the way, caching does not make LLM inference linear. It's still quadratic, but the constant in front of the quadratic term becomes a lot smaller.
- computably 5mo ago> Also by the way, caching does not make LLM inference linear. It's still quadratic, but the constant in front of the quadratic term becomes a lot smaller. Touché. Still, to a reasonable approximation, caching makes the dominant term linear, or equiv, linearly scales the expensive bits.
- _flux 5mo agoWhat we would call O(n^2) in your rewriting message history would be the case where you have an empty database and you need to populate it with a certain message history. The individual operations would take 1, 2, 3, .. n steps, so (1/2)*n^2 in total, so O(n^2). This is the operation that is basically done for each message in an LLM chat in the logical level: the complete context/history is sent in to be processed. If you wish to process only the additions, you must preserve the processed state on server-side (in KV cache). KV caches can be very large, e.g. tens of gigabytes.