3 ms·
Isn't that because of the context window size?
by marknutter 5mo ago
Isn't that because of the context window size?
- SatvikBeri 5mo agoThe context window has nothing to do with RAM usage and even if it did, a million tokens of context is maybe 5mb.
- vlovich123 5mo agoIt has nothing to do with local RAM usage. But a million tokens of LLM context is decidedly not 5mb. The rough estimate is 2 * L * H_kv * D * bytes per element Where: * L = number of layers * H_kv = # of KV heads * D = head dimension * factor of 2 = keys + values The dominant factor here is typically 2 * H_kv * D since it’s usually at least 2048 bytes. Per token. For Llama3 7B youre looking at 128gib if you’re context is really 1M (not that that particular model supports a context so big). DeepSeek4 uses something called sparse attention so the above calculus is improved - 1M of context would use 5-10GiB. But regardless of the details, you’re off by several orders of magnitude.
- tujux 5mo agoPretty sure we're talking about the output text, not the tensors.
- m00x 5mo agoThese LLM replies are really getting annoying.
- vlovich123 5mo agoMine? I literally wrote what I wrote because “context window” as a term of art refers to the LLM’s context window. I guess get better at detecting LLMs instead of accusing everything of being an LLM reply?
- bluegatty 5mo ago'A million tokens of context' is literally Terrabytes of KV cache VRAM on very expensive Nvidia silicon - on the model. On the Agent, yes, the context window does relate to RAM, because the 'entire conversational history' is generally kept in memory. So ballpark 1M 'words' across a bunch of strings. It's not that-that much. Claude Code is not inneficient because 'it's not Rust' - it's just probably not very efficiently designed. Rust does not bestow magical properties that make memory more efficient really. A bit more, but it's not going to change this situation. 'Dong it in Rust' might yield amazing returns just because the very nature of the activity is 'optimization'.
- rixed 5mo agoRust "denialism" is as annoying as rust evangelism. Of course any seemingly idiomatic rust is going to run circles around TS transpiled into JIT-compiled JS.
- bluegatty 5mo agoLamenting any 'not even criticism' of Rust as 'denialism' is just evidence of the insane cult that is Rust. Rebuilding Claude Code in Rust will make almost no difference in terms of real world performance. V8 is 'relatively fast', and there wouldn't be any noticeable improvements there, and probably not memory footprint either. The source for Claude Code was leaked and it's a vibe-coded mess, there's not much thought given to clean architecture, it's unlikely they've just cleaned up a bit and given thought to memory consumption etc, if they did, they'd get by far most of the way there and likely abnegate and real want to 'do it in rust', unless there are other architectural considerations.
- imtringued 5mo agoYou're the delusional one for bringing up the memory usage of the inference server that clearly isn't running inside the coding agent. The problem with your comments is that you're showing off a fundamental lack of understanding between managed languages and unmanaged languages. The vast majority of GCs are optimized for throughput and allocate big chunks of memory. They also tend to never release it if there was a temporary memory spike. The most advanced GCs also tend to have either read or write barriers, which slow down basic object accesses. Just in time compilation and managed languages in general need to retain a runtime representation of the source code to perform JIT compilation and then they have to store the compiled code in memory as well. JavaScript uses references against dynamic objects, which means you have to pay the indirection cost of a pointer but you also need to store type information as well to monomorphize the object literals and classes at runtime and fall back to a regular hashmap when fields are added dynamically. All of these things will add up and increase the amount of memory the application uses and how slow it runs. Sure Claude Code has severe architectural issues causing it to leak hundreds of gigabytes of RAM, but if those were not there you could easily build a C++ based alternative that runs circles around a hypothetical JavaScript based Claude Code that got its act together.
- gidellav 5mo agoHi, I'm the developer of zerostack! No, the memory footprint is not beacuse of the context window size: on my benchmarks, with a 128k context loaded, and it jumped from 8MB (without any chat/context loaded) to 11MB. The reasons why the memory footprint of zerostack are: - Rust, and not JS/Python, so no interpreters/VMs on top - Load-as-needed, so we only allocate things like LLM connectors when needed - `smallvec` used for most of the array usage of the tool (up to N items are stored in stack) - `compactstring` used for most of the string usage of the tool (up to N chars are stored in stack) - `opt-level=z` to force LLVM to optimize for binary size and not for performance (even tho we still beat both in TTFT and in tool use time opencode) - heavy usage of [LTO](https://en.wikipedia.org/wiki/Interprocedural_optimization#WPO_and_LTO https://en.wikipedia.org/wiki/Interprocedural_optimization#W...)
- SwellJoe 5mo agoThe context window is not on your system. It's on the server with the model. There may be some local prompt caching, of some sort, but you're not locally hosting the context unless you're also locally hosting the model.
- bluegatty 5mo agoChat history is kept locally, generally you have to send the 'whole history' to the model 'each turn'.
- rixed 5mo agoOnly "generally"? I'm curious what API has moved away from this protocol that seems mode adapted to conversaions with humans than agentic loops.
- bluegatty 5mo agoSo the standard API you pass it all along but I think there are some odd open ai apis that are different.
- _flux 5mo agoTo me it would certainly make sense if the protocol just said "append this text to context window id/sha256", in particular as the data is cached in tensor level in the provider side, so they need to first do that lookup anyway. So I would be surprised if they don't have that. In addition, this protocol could make it more transparent to say "oh we cannot proceed as we dropped the this cache, are you sure you want to proceed and consume a whole lot of expensive uncached tokens?". Oh, maybe that's a reason not to do it..
- SwellJoe 5mo agoThat's just the plain text (or whatever files), that's not the context the model is directly working with on the server, which is tokenized, embedded, vectorized and has attention run against those vectors. The local history is generally quite small, the context generally quite a bit larger. A text conversation of a few hundred kilobytes in plain text will be gigabytes in context.