3 ms·
The moment you can't fit the model in VRAM and have to use PCIE the performance impact is MASSIVE. Depending on the type of model/usecase/how much you can fit
by filterfiber 3y ago
The moment you can't fit the model in VRAM and have to use PCIE the performance impact is MASSIVE.
Depending on the type of model/usecase/how much you can fit in VRAM it very well may still be usable, however compared to fitting everything in vram it'll be abysmal for most any model type.
From my understanding the entire model must be read per iteration (or very close to all of it. This is why offloading any part of it impacts the performance so much.
Diffusion models depending on the sampler and some other stuff can potentially get away with only a few steps (usually 2-20 best case) and are relatively small (3B or less).
However LLM's are generally much bigger (they mostly start at 7B, there's a few exceptions depending on the use case), and they need an entire iteration per token which is only a few characters (think 2 to 3). So llm's especially are far more useful if you can keep them in vram.
- dale_glass 3y agoThanks, I messed around with this stuff a little bit, but have little understanding of what goes on in the guts of it.