3 ms·Honestly you can run this on a 16GB VRAM GPU with llama.cpp. Just try it!by am17an 8mo agoHonestly you can run this on a 16GB VRAM GPU with llama.cpp. Just try it!