18 ms·
Hey guys, I was so inspired by the llama.cpp project that I spent all day today to build a weekend side project. Basically it lets you one-click install LLaMA
by cocktailpeanut 4y ago
Hey guys, I was so inspired by the llama.cpp project that I spent all day today to build a weekend side project.
Basically it lets you one-click install LLaMA on your machine with no bullshit. All you need is just run "npx dalai llama".
I see that the #1 post today is a whole long blog post about how to walk through and compile cpp and download files and all that to finally run LLaMA on your machine, but basically I have 100% automated this with a simple NPM package/application.
On top of that, the whole thing is a single NPM package and was built with hackability in mind. With just one line of JS function call you can call LLaMA from YOUR app.
Lastly, EVEN IF you don't use JavaScript, Dalai exposes a socket.io API, so you can use whatever language you want to interact with Dalai programmatically.
I discussed a bit more about this on a Twitter thread. Check it out: https://twitter.com/cocktailpeanut/status/1635040322471489537 https://twitter.com/cocktailpeanut/status/163504032247148953...
It should "just work". Have fun!
- evo_9 4y agoWhen I run this commnad: npx dalai llama I get the following output / errors? What exactly do I need to install prior to running that command? ---------------------------- >> npx dalai llama exec: git clone https://github.com/ggerganov/llama.cpp.git https://github.com/ggerganov/llama.cpp.git /Users/rickg/llama.cpp in undefined git clone https://github.com/ggerganov/llama.cpp.git https://github.com/ggerganov/llama.cpp.git /Users/rickg/llama.cpp exit The default interactive shell is now zsh. To update your account to use zsh, please run `chsh -s /bin/zsh`. For more details, please visit https://support.apple.com/kb/HT208050 https://support.apple.com/kb/HT208050. a.cpp3.2$ git clone https://github.com/ggerganov/llama.cpp.git https://github.com/ggerganov/llama.cpp.git /Users/rickg/llam fatal: destination path '/Users/rickg/llama.cpp' already exists and is not an empty directory. bash-3.2$ exit exit exec: git pull in /Users/rickg/llama.cpp git pull exit The default interactive shell is now zsh. To update your account to use zsh, please run `chsh -s /bin/zsh`. For more details, please visit https://support.apple.com/kb/HT208050 https://support.apple.com/kb/HT208050. bash-3.2$ git pull Already up to date. bash-3.2$ exit exit exec: python3 -m venv /Users/rickg/llama.cpp/venv in undefined python3 -m venv /Users/rickg/llama.cpp/venv exit The default interactive shell is now zsh. To update your account to use zsh, please run `chsh -s /bin/zsh`. For more details, please visit https://support.apple.com/kb/HT208050 https://support.apple.com/kb/HT208050. bash-3.2$ python3 -m venv /Users/rickg/llama.cpp/venv bash-3.2$ exit exit exec: /Users/rickg/llama.cpp/venv/bin/pip install torch torchvision torchaudio sentencepiece numpy in undefined /Users/rickg/llama.cpp/venv/bin/pip install torch torchvision torchaudio sentencepiece numpy exit The default interactive shell is now zsh. To update your account to use zsh, please run `chsh -s /bin/zsh`. For more details, please visit https://support.apple.com/kb/HT208050 https://support.apple.com/kb/HT208050. io sentencepiece numpy/llama.cpp/venv/bin/pip install torch torchvision torchaud Requirement already satisfied: torch in ./llama.cpp/venv/lib/python3.10/site-packages (1.13.1) Requirement already satisfied: torchvision in ./llama.cpp/venv/lib/python3.10/site-packages (0.14.1) Requirement already satisfied: torchaudio in ./llama.cpp/venv/lib/python3.10/site-packages (0.13.1) Requirement already satisfied: sentencepiece in ./llama.cpp/venv/lib/python3.10/site-packages (0.1.97) Requirement already satisfied: numpy in ./llama.cpp/venv/lib/python3.10/site-packages (1.24.2) Requirement already satisfied: typing-extensions in ./llama.cpp/venv/lib/python3.10/site-packages (from torch) (4.5.0) Requirement already satisfied: pillow!=8.3.,>=5.3.0 in ./llama.cpp/venv/lib/python3.10/site-packages (from torchvision) (9.4.0) Requirement already satisfied: requests in ./llama.cpp/venv/lib/python3.10/site-packages (from torchvision) (2.28.2) Requirement already satisfied: charset-normalizer<4,>=2 in ./llama.cpp/venv/lib/python3.10/site-packages (from requests->torchvision) (3.1.0) Requirement already satisfied: urllib3<1.27,>=1.21.1 in ./llama.cpp/venv/lib/python3.10/site-packages (from requests->torchvision) (1.26.15) Requirement already satisfied: idna<4,>=2.5 in ./llama.cpp/venv/lib/python3.10/site-packages (from requests->torchvision) (3.4) Requirement already satisfied: certifi>=2017.4.17 in ./llama.cpp/venv/lib/python3.10/site-packages (from requests->torchvision) (2022.12.7) [notice] A new release of pip available: 22.3.1 -> 23.0.1 [notice] To update, run: python3 -m pip install --upgrade pip bash-3.2$ exit exit exec: make in /Users/rickg/llama.cpp make exit The default interactive shell is now zsh. To update your account to use zsh, please run `chsh -s /bin/zsh`. For more details, please visit https://support.apple.com/kb/HT208050 https://support.apple.com/kb/HT208050. bash-3.2$ make I llama.cpp build info: I UNAME_S: Darwin I UNAME_P: arm I UNAME_M: arm64 I CFLAGS: -I. -O3 -DNDEBUG -std=c11 -fPIC -pthread -DGGML_USE_ACCELERATE I CXXFLAGS: -I. -I./examples -O3 -DNDEBUG -std=c++11 -fPIC -pthread I LDFLAGS: -framework Accelerate I CC: Apple clang version 12.0.5 (clang-1205.0.22.9) I CXX: Apple clang version 12.0.5 (clang-1205.0.22.9) cc -I. -O3 -DNDEBUG -std=c11 -fPIC -pthread -DGGML_USE_ACCELERATE -c ggml.c -o ggml.o ggml.c:1364:25: error: implicit declaration of function 'vdotq_s32' is invalid in C99 [-Werror,-Wimplicit-function-declaration] int32x4_t p_0 = vdotq_s32(vdupq_n_s32(0), v0_0ls, v1_0ls); ^ ggml.c:1364:19: error: initializing 'int32x4_t' (vector of 4 'int32_t' values) with an expression of incompatible type 'int' int32x4_t p_0 = vdotq_s32(vdupq_n_s32(0), v0_0ls, v1_0ls); ^ ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ ggml.c:1365:19: error: initializing 'int32x4_t' (vector of 4 'int32_t' values) with an expression of incompatible type 'int' int32x4_t p_1 = vdotq_s32(vdupq_n_s32(0), v0_1ls, v1_1ls); ^ ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ ggml.c:1367:13: error: assigning to 'int32x4_t' (vector of 4 'int32_t' values) from incompatible type 'int' p_0 = vdotq_s32(p_0, v0_0hs, v1_0hs); ^ ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ ggml.c:1368:13: error: assigning to 'int32x4_t' (vector of 4 'int32_t' values) from incompatible type 'int' p_1 = vdotq_s32(p_1, v0_1hs, v1_1hs); ^ ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 5 errors generated. make: * [ggml.o] Error 1 bash-3.2$ exit exit /Users/rickg/.npm/_npx/3c737cbb02d79cc9/node_modules/dalai/index.js:153 throw new Error("running 'make' failed") ^ Error: running 'make' failed at Dalai.install (/Users/rickg/.npm/_npx/3c737cbb02d79cc9/node_modules/dalai/index.js:153:13)
- yawnxyz 4y agoWow that's so incredible. Thanks for putting this together! Do you have any machine specs associated with this? Can an old-ish Macbook Pro run this service? I'm also curious, since I'm new to all this — is it possible to run something like this on Fly.io or does it take up way too much space?
- sp332 4y ago7B is the default. If it's quantized to 4 bits, that's a 3.9 GB file.
- teruakohatu 4y agoVery nice. Anyway to add an option to install elsewhere other than ~/ ?
- GordonS 4y agoLooks great! Does it work on Windows please?
- dragonwriter 4y agoIf it makes common unix-ish assumptions like “Python 3 executables have a ‘3’ appended to their name”, which other comments here seem to suggest it does, it won’t, even if you have the required version of python installed.
- GordonS 4y agoSo, I actually got it working on Windows, pretty easily! The provided `main.exe` binary worked as-is, but `quantize.exe` did not - I built myself with CMake, and `quantize.exe` started working too.
- volaski 4y agoCurious too. Let me know if you try it out. Technically I think it should work.
- buzzier 4y agoFor Windows: 1. Binary build https://github.com/jaykrell/llama.cpp/releases/tag/1 https://github.com/jaykrell/llama.cpp/releases/tag/1 2. Quantized model (7B/13B/30B) https://mega.nz/folder/UjAUES6Z#bGhKkyiZX3eRrn9HcxVVfA https://mega.nz/folder/UjAUES6Z#bGhKkyiZX3eRrn9HcxVVfA 3. main.exe -m ggml-model-q4_0.bin -t 8 -n 128
- EZ-Cheeze 4y agoAdd something like this to your instructions: "Make sure you have Node.js installed on your computer."
- deleted 4y ago[deleted]
- turbocon 4y agoYea not a nodejs/javascript dev at all but this is failing to install on Fedora. I don't have time to dig into it at the moment but if anybody has any well known gotchas that could be the issue that would be helpful :) Edit: I do have nodejs and npx installed
- vorticalbox 4y agoMaybe make, python and pip. From what I gather this is a node wrapper it's actually python that runs the model
- AlecSchueler 4y agoOne step install after the steps that lead up to it.
- holtkam2 4y agoThis is awesome! I've wanted to try llama.cpp and you just reduced my to-do list significantly on my Sunday :) Thanks!
- sieste 4y agoDoes anyone know how to avoid downloading the model weights when doing `npx dalai llama`, and instead telling the install process where they are on my drive?
- gregsadetsky 4y agoyou could clone the repo and comment out https://github.com/cocktailpeanut/dalai/blob/main/index.js#L85 https://github.com/cocktailpeanut/dalai/blob/main/index.js#L... i.e. the specific synchronous download call..?
- skykooler 4y agoHow powerful of a computer does this need? It would be useful to see, for one thing, minimum RAM requirements for these models.
- spion 4y agollama.cpp needs 40GB for the 65B model (due to int4 quantization) RamNeeded(other_size) ~= 40GB * other_size/65B
- pmarreck 4y agoI ran "npx dalai llama" and it's just... sitting there (after I hit "y" to confirm). I checked btop++ and there's barely any downloading or CPU activity occurring, so not sure what it's doing... but does "pip3 install torch torchvision torchaudio sentencepiece numpy" take a while? If it's actually downloading the 3.9GB of model weights or whatever, it would be pretty cool if it showed a progress bar of some sort. Stretch goal, for sure, but a very nice nicety for users. anyway, I'll leave it be and check on it to see when it's complete. Super cool if this works!!
- jacooper 4y agoDoes this use the GPU? If not why? Aren't GPUs much faster than CPUs at AI?
- boredemployee 4y agoI think thats exactly the point so everyone can run it on their PCs with no GPU.
- lolinder 4y agoOr without a beefy GPU. I've got 8GB VRAM, which is great for Stable Diffusion but not useful for any of the language models released so far. I think the 4-bit 7B LLaMA would work, but the 7B is pretty fast anyway without GPU.
- boredemployee 4y agoI'm installing it here. How's the 7B model going so far?
- lolinder 4y agoHaha, I just finished ordering 32GB of additional memory for my PC so I can run the 65B model, if that tells you anything. I'm upgrading from 32GB -> 64GB. 7B is fine, 13B is better. Both are fun toys and almost make sense most of the time, but even with a lot of parameter tuning they're often incoherent. You can tell that they have encoded fewer relationships between concepts than the higher-parameter models we've gotten used to--it's much closer to GPT-2 than GPT-3. They're good enough to whet my appetite and give me a lot of ideas of what I want to do, they're just not quite good enough to make those applications reliably useful. Based on the reports I'm hearing here of just how much better the 65B model is than the 7B, I decided it was worth $80 for a few new sticks of RAM to be able to use the full model. Still way cheaper than buying a graphics card capable of handling it.
- iambateman 4y ago
- cocktailpeanut 4y agoUPDATE: Thanks for all the feedback! I went outside to take a walk after posting this and just came back, and went through them to summarize what needs to be improved. Basically looks like it comes down to the following: - *customize features:* Should not be difficult (will add flag features) - *path:* customize the home directory (instead of automatically storing to $HOME) - *python:* some people are having issues with the python binary (since the package is essentially calling these shell commands). Maybe add a flag to specify the exact name of the python binary (such as "--python python3") - *avoid downloading files:* I have this issue too when I just want to install the code instead of downloading the full model which takes a long time. Might add a flag to avoid downloading models in case you already have them (EDIT: actually upon thinking about it, it's better to just set the source model folder, something like --model) - *other flags:* The rest of the flags natively supported by the llama.cpp project, such as top_k, top_p, temp, batch_size, threads, seed, n_predict, etc. (They are already in the code but just was not exposed for CLI and not documented) - *documentation* - document the machine spec - document the storage spec: how much space is used? - node version: which version of node.js is required? - python version: which version of python doesn't work? Am I missing anything? Feel free to leave comments, will try to roll out some updates as soon as I can. To stay updated, feel free to follow me on twitter https://twitter.com/cocktailpeanut https://twitter.com/cocktailpeanut (or you could create issues on GitHub too!)
- cocktailpeanut 4y agoUPDATE 2: Thanks to all the pull requests, we've managed to solve most of these issues in the most optimal manner. Version 0.1.0 released: https://news.ycombinator.com/item?id=35143171 https://news.ycombinator.com/item?id=35143171
- icosahedron 4y agoI followed the initial instructions and the 7B model worked just fine. I tried the supplementary instructions to download some of the models (7B, 13B, and 30B), and it didn't seem to work. The prompt returned nothing after waiting for several minutes. Is there a way to run just one of the larger models?
- m3kw9 4y agoMade a comment on the other thread: why can’t we have a one click install thing and here it is. Nice!
- upghost 4y agoMy biggest concern about these LLMs was the corporate sequestration and the potential socioeconomic imbalances it would create. The work you are doing here is part of some amazing work to check that back. In summary—- Bruhhhhhh. THANK YOU!
- sebastianconcpt 4y agoThis is something to keep an eye, really. The solution for making that sequestration impossible is twofold: 1. to know how to architect and create LLMs (including training data readiness) 2. have them produced in hardware that is acquirable at reasonable cost for a normal citizen
- davidy123 4y agoYou, sir or madam, are a hero.
- anigbrowl 4y agoWell that's pretty wild. I was wondering whether I wanted to build LLaMA tomorrow but you upended my plans in the space of 2 minutes. 10/10 well done.
- Tepix 4y agoThere's an elephant in the room, or is it just me? Is your script making users violate the original license agreement(§)? For the record, i don't think Meta will go after you or anyone else. But they may decide not to make their future models available after what is happening with the Llama weights. I realize that some people are of the opinion that AI models (weights) cannot be copyrighted at all. -- § the license agreement is at https://forms.gle/jk851eBVbX1m5TAv5 https://forms.gle/jk851eBVbX1m5TAv5
- deleted 4y ago[deleted]
- Tiberium 4y agoYes, you are right, every project that distributes LLaMA right now is violating Meta's agreement.
- pksebben 4y agoI've got a weird, probably untrue conspiracy theory about this. Hugging face releases stable diffusion. It goes viral and vastly outpaces the competition in the blink of an eye. Then they get sued. Meta sees both of these things go down. Meta needs a leg up on chat GPT, but worries about legal repercussions similar to stable diffusion. Whoops, it leaked! Hey, we didn't say those dastardly devs could use it.
- deleted 4y ago[deleted]
- sebzim4500 4y ago>But they may decide not to make their future models available after what is happening with the Llama weights. I think that ship has probably sailed, in that no one is going to release weights in this way again. Either they will publish them outright (like Whisper) or they will keep them (almost) completely closed.