4 ms·
Another option is to download and compile llama.cpp and you should be able to run quantized models at an acceptable speed. https://github.com/ggerganov/llama.c
by programd 2y ago
Another option is to download and compile llama.cpp and you should be able to run quantized models at an acceptable speed.
https://github.com/ggerganov/llama.cpp https://github.com/ggerganov/llama.cpp
Also, if you can spend the $60 and buy another 32GB of RAM, this will allow you to run the 30GB models quite nicely.
- eth0up 2y agoUnfortunately motherboard is capped at 16Gb ram