3 ms·
It's very slow, and for the 7b model you're still looking at a pretty hefty RAM hit whether it's CPU or GPU. The model download is something like 40GB.
by cmsj 3y ago
It's very slow, and for the 7b model you're still looking at a pretty hefty RAM hit whether it's CPU or GPU. The model download is something like 40GB.
- MacsHeadroom 3y agoThere's already support in llama.cpp. It runs faster than ChatGPT on my old laptop CPU.