5 ms·
Is this some joke? They use Llama 2 7B? What year is it?
by az226 2y ago
Is this some joke? They use Llama 2 7B? What year is it?
- x1000 2y agoIf they had experimented using a newer model (gemma 3, deepseek-1 7b, etc.) and reported better results, would that be because their newer baseline model was better than the llama 2 model used in the previous methods' experiments? A more comprehensive study would include results for as many baseline models as possible. But there are likely other researchers in the lab all waiting to use those expensive GPUs for their experiments as well.
- josephg 2y agoSure. But papers take a really long time to write and go through peer review. I think my paper on collaborative editing took about 4 months from the point where we were done writing to the point at which it appeared on arxiv. This research was almost certainly done well before Gemma 3 and Deepseek were released.
- PaulHoule 2y agoThe best model is the one you can fit in memory. About as soon as GPT-4 came out I said that OpenAI was doomed on the trajectory it was on because they could not afford to develop a GPT-5, GPT-6, etc. Real innovation comes out of doing a lot of experiments and that means doing experiments quickly with the resources you have. So you do most of your experiments with non-frontier models, enough to make a good prediction of what would happen if you maxxed out your model size, then you go big. That's how you make everyone else have a "DeepSeek moment". A company like Apple wants to pick something on the frontier and keep advancing on a straight line. Works great if you want to make an M1, M2, M3, ... ARM chip but that's not how progress works in AI today.
- monocasa 2y agoI mean, there's other, better 7B models than Lllama 2 at this point.
- hinkley 2y agoWill we see models built on b-trees to deal with memory requirements? Have we already?
- sujayakar 2y agoDeepseek is already using SSDs for their KV cache: https://github.com/deepseek-ai/3FS https://github.com/deepseek-ai/3FS
- vlovich123 2y agoYou are deeply misunderstanding what the KV cache referred to here is. It’s not for storing data. This is the KV cache that’s part of the model to reduce quadratic compute complexity into linear for self attention. This is not stored on SSD - it’s in VRAM (or CPU if you’re not using a GPU)
- boroboro4 2y agoThey, in fact, mention inference kv cache as use case in readme. The most advanced kv caching uses hierarchy of gpu ram/regular ram/ssd. Seems like they were able to use their storage abstraction for last tier.
- magicalhippo 2y agohttps://github.com/deepseek-ai/3FS?tab=readme-ov-file#3-kvcache https://github.com/deepseek-ai/3FS?tab=readme-ov-file#3-kvca... KVCache is a technique used to optimize the LLM inference process. It avoids redundant computations by caching the key and value vectors of previous tokens in the decoder layers. The top figure demonstrates the read throughput of all KVCache clients (1×400Gbps NIC/node), highlighting both peak and average values, with peak throughput reaching up to 40 GiB/s
- vlovich123 2y agoThat's because DeepSeek uses MLA which apparently does allow offloading the KV cache. That doesn't apply to all models, particularly the open-weight models that are primarily GQA AFAIK.
- krasin 2y ago> Is this some joke? They use Llama 2 7B? What year is it? They use llama2 to demonstrate that their compression method works. There are potential cases: 1. The method works on all / most LLMs. In this case, it does not matter on which model they demonstrated the effect. 2. The method only works on llama2, but not on other models. Given that they published the code, I expect that people will quickly test the method on many other models, so we will know that soon. And yet - there would be a scientific significance even if it works only on llama2, as it would mean that there's some special and good in that architecture. But I would bet it's #1 - the method works on most of the models and they just picked whatever they had already had code bindings to, to save the effort.