4 ms·
I wonder where a critical threshold might lie? For example, right now if you only have 24Gb of VRAM, the best you can run locally at usable speeds are the 4-bit
by nullsense 3y ago
I wonder where a critical threshold might lie? For example, right now if you only have 24Gb of VRAM, the best you can run locally at usable speeds are the 4-bit quantized LLaMa 30B models, which are only semi-usable. If you have 48Gb however, you can run the 4-bit quantized LLaMa 65B, which is much better.
I assume there will be advances that basically make the context window practically infinite. That ought to make the lower parameter count models much more powerful on its own. I also assume they will become more efficient to run through sparsity/pruning, though I'm a total novice on this topic.
I wonder how many generations of hardware are we talking? One? Two? Or three? It feels like it's potentially within that range.