3 ms·
You can run 4-bit quantized version at a small (though nonzero) cost to output quality, so you would only need 16GB for that. Also it's entirely possible to ru
by wgd 2y ago
You can run 4-bit quantized version at a small (though nonzero) cost to output quality, so you would only need 16GB for that.
Also it's entirely possible to run a model that doesn't fit in available GPU memory, it will just be slower.