3 ms·
How do you do "due diligence" on an API that frequently makes undocumented changes and only publishes acknowledgement of change after users complain? You're al
by doesnt_know 6mo ago
How do you do "due diligence" on an API that frequently makes undocumented changes and only publishes acknowledgement of change after users complain?
You're also talking about internal technical implementations of a chat bot. 99.99% of users won't even understand the words that are being used.
- tempest_ 6mo agoI use CC, and I understand what caching means. I have no idea how that works with a LLM implementation nor do I actually know what they are caching in this context.
- hakanderyal 6mo agoCC can explain it clearly, which how I learned about how the inference stack works.
- libraryofbabel 6mo agoThey are caching internal LLM state, which is in the 10s of GB for each session. It's called a KV cache (because the internal state that is cached are the K and V matrices) and it is fundamental to how LLM inference works; it's not some Anthropic-specific design decision. See my other comment for more detail and a reference.
- dlivingston 6mo agoWhat is being discussed is KV caching [0], which is used across every LLM model to reduce inference compute from O(n^2) to O(n). This is not specific to Claude nor Anthropic. [0]: https://huggingface.co/blog/not-lain/kv-caching https://huggingface.co/blog/not-lain/kv-caching
- computably 6mo ago> How do you do "due diligence" on an API that frequently makes undocumented changes and only publishes acknowledgement of change after users complain? 1. Compute scaling with the length of the sequence is applicable to transformer models in general, i.e. every frontier LLM since ChatGPT's initial release. 2. As undocumented changes happen frequently, users should be even more incentivized to at least try to have a basic understanding of the product's cost structure. > You're also talking about internal technical implementations of a chat bot. 99.99% of users won't even understand the words that are being used. I think "internal technical implementation" is a stretch. Users don't need to know what a "transformer" is to understand the trade-off. It's not trivial but it's not something incomprehensible to laypersons.
- fragmede 6mo ago> 99.99% of users won't even understand the words that are being used. That's a bad estimate. Claude Code is explicitly a developer shaped tool, we're not talking generically ChatGPT here, so my guess is probably closer to 75% of those users do understand what caching is, with maybe 30% being able to explain prompt caching actually is. Of course, those users that don't understand have access to Claude and can have it explain what caching is to them if they're interested.