4 ms·
Yeah chatgpt pretty much nailed it.
by nodja 1y ago
Yeah chatgpt pretty much nailed it.
- bawana 1y agoBut you still have to load the data for each request. And in an LLM doesnt this mean the WHOLE kv cache because the kv cache changes after every computation? So why isnt THIS the bottleneck? Gemini is talking about a context window of a million tokens- how big would the kv cache fir this get?