2 ms·
> This may be obvious to people who do this regularly This is not that obvious. Calculating VRAM usage for VLMs/LLMs is something of an arcane art. There are a
by icelancer 1y ago
> This may be obvious to people who do this regularly
This is not that obvious. Calculating VRAM usage for VLMs/LLMs is something of an arcane art. There are about 10 calculators online you can use and none of them work. Quantization, KV caching, activation, layers, etc all play a role. It's annoying.
But anyway, for this model, you need 40+ GB of VRAM. System RAM isn't going to cut it unless it's unified RAM on Apple Silicon, and even then, memory bandwidth is shot, so inference is much much slower than GPU/TPU.
- cellis 1y agoAlso I think you need a 40GB "card", not just 40GB of vram. I wrote about this upthread, you're probably going to need one card, I'd be surprised if you could chain several GPUs together.
- rapfaria 1y agoNot sure what you mean or new to llms, but two RTX 3090 will work for this, and even lower-end cards will (RTX3060) once it's GGUF'd
- karolist 1y agodo you mean https://github.com/pollockjj/ComfyUI-MultiGPU https://github.com/pollockjj/ComfyUI-MultiGPU? One GPU would do the computation, but others could pool in for VRAM expansion, right? (I've not used this node)
- Auracle 1y agoNah, that won’t gain you much (if anything?) over just doing the layer swaps on RAM. You can put the text encoder on the second card but you can also just put it in your RAM without much for negatives.
- axoltl 1y agoThis isn't a transformer, it's a diffusion model. You can't split diffusion models across compute nodes.
- icelancer 1y agoOh right, I forgot some diffusion models can't offload / split layers. I don't use vision generation models much at all - was just going off LLM work. Apologies for the potential misinformation.
- xarope 1y agowill the new AMD AI CPUs work? like an AI HX 395 or the slower 370? I'm stuck on an A2000 w/16GB of VRAM and wondering what's a worthwhile upgrade.
- Auracle 1y agoIt may fit but image generation on anything but Nvidia is so slow it won’t be worth it.