3 ms·
With 128 GB of unified memory, large models can fit into memory, but memory bandwidth quickly becomes the limiting factor. With long contexts in particular, the
by MyPhiloEngine 2mo ago
With 128 GB of unified memory, large models can fit into memory, but memory bandwidth quickly becomes the limiting factor. With long contexts in particular, the KV cache can consume a lot of memory and bandwidth. Quantizing the KV cache to Q8 roughly halves its memory usage, making it one of the most effective optimizations. MoE models are also well suited to this hardware because only a portion of their parameters is active for each token. As a result, a 122B-A10B model can run better than a dense 70B model.