3 ms·
Thanks for this! I was able to integrate alpaca-30B into a slack bot & a quick tkinter GUI (coded by GPT-4 tbh) by just shelling out to `./main` in both cases,
by bestcoder69 4y ago
Thanks for this! I was able to integrate alpaca-30B into a slack bot & a quick tkinter GUI (coded by GPT-4 tbh) by just shelling out to `./main` in both cases, since model loading is so quick now. (I didn't even have to ask GPT-4 to code me up Python bindings to llama's c-style api!)
- bugglebeetle 4y agoWhat’s your setup for running these? I’m not seeing performance improvements on off the shelf hardware that would allow for this.
- MacsHeadroom 4y agoI host a llama-13B IRC chatbot on a spare old android phone.
- bugglebeetle 4y agoHave a repo anywhere?
- MacsHeadroom 4y agoIt's just the same llama.cpp repo everyone else is using. You just git clone it to your android phone in termux and then run make and you're done. https://github.com/ggerganov/llama.cpp https://github.com/ggerganov/llama.cpp Assuming you have the model file downloaded (you can use wget to download it) these are the instructions to install and run: pkg install git pkg install cmake pkg install build-essential git clone https://github.com/ggerganov/llama.cpp https://github.com/ggerganov/llama.cpp cd llama.cpp make -j ./main
- bugglebeetle 4y agoYeah, I’ve already been running llama.cpp locally, but not found it to perform at the level attested in the comment (30B model as a chat bot on commodity hardware). 13B runs okay, but inference appears generally too slow on to do anything useful on my MacBook. I wondered what you might be doing to get usable performance in that context.
- MacsHeadroom 4y agoYou can change the number of threads llama.cpp uses with the -t argument. By default it only uses 4. For example, if your CPU has 16 physical cores then you can run ./main -m model.bin -t 16 16 cores would be about 4x faster than the default 4 cores. Eventually you hit memory bottlenecks. So 32 cores is not twice as fast as 13 cores unfortunately.
- bugglebeetle 4y agoThanks! Will test that out!