4 ms·
The shell script you linked is indeed the raw weights, but if you have the storage space you can download them and use the llama-cpp repo posted here a few days
by NameError 4y ago
The shell script you linked is indeed the raw weights, but if you have the storage space you can download them and use the llama-cpp repo posted here a few days ago (https://github.com/ggerganov/llama.cpp https://github.com/ggerganov/llama.cpp) to compress them, the 'Usage' section of the readme worked for me.
I'm using a not-that-new Macbook Pro (intel, 16GB memory) and was able to run the 7B and 13B models that way, tried 30B but it seemed to hang.
- unshavedyak 4y agoI'm in the process of using a PR'd Dockerfile for Dalai, which i think will handle everything hands free. Ie download the raw files, quantize them, pick which you want (7, 13, etc). We'll see how it goes. A bit heavy handed perhaps to use NPM/etc, but the Dockerfile really helps me ignore all the dependencies i'm adding hah.
- zamnos 4y agoI might be being dense, but why not distribute the quantized files?
- turbo_fart 4y agoWith alpaca you get just the quantized file, for llama you need to do it all yourself. Why? No idea