3 ms·
How much GPU RAM would be needed to run this with just one GPU?
by jftuga 2y ago
How much GPU RAM would be needed to run this with just one GPU?
- paulluuk 2y agoI haven't tested it, but likely around 170GB, regardless of if you're using only one GPU or spreading it out over several ones.
- reissbaker 2y ago144GB VRAM to load the weights at FP16, 72GB quantized to FP8. To figure out the KV cache size you'll need for an LLM, you can use the following formula: https://x.com/AlpinDale/status/1841305040545329535 https://x.com/AlpinDale/status/1841305040545329535 Simplified for posterity: kv_bytes = kv_bits / 8 hidden_per_head = hidden_size // num_attention_heads total_heads = hidden_per_head * num_key_value_heads kv_bytes_per_token = 2 * kv_bytes * num_hidden_layers * total_heads (Edit: I accidentally swapped in some of the vision config bytes in my original calculation; these are the corrected numbers.) So, for NVLM 1.0 72B, that works out to 640kb per token assuming FP16 KV cache. If you use the entire 32k context length, that's an extra ~20GB of overhead for the KV cache. Then depending on how you're running the LLM, there might be extra overhead e.g. compiled CUDA graphs. You can cut this down lower by using grouped query attention as described here: https://medium.com/@plienhar/llm-inference-series-4-kv-caching-a-deeper-look-4ba9a77746c8 https://medium.com/@plienhar/llm-inference-series-4-kv-cachi... This allows you to divide that number by the number of grouped heads, although it trades off accuracy for VRAM usage. But TLDR, a minimum of around 164GB of VRAM at full accuracy. To me that seems fairly low, and I think vLLM would OOM without significantly more than that, but that's about as low as you could go in theory if you're running everything at FP16. Half that, of course, for FP8. You'll typically need to have a copy of the KV cache per GPU, if you're using multiple GPUs, so multiply the KV cache overhead by the number of GPUs you're using. This will depend on what the specs for the GPUs you're using are; for example, you'll need 3 H100s (really four, since vLLM wants the number of heads to be evenly divisible by the number of GPUs); if you're using L40Ses, you'll need eight of them; but most likely only a single AMD MI300x.