15 ms·
Guide to running Llama 2 locally
- ericHosick 3y agoThe LLM is impressive (llama2:13b) but appears to have been greatly limited to what you are allowed to do with it. I tried to get it to generate a JSON object about the movie The Matrix and the model refuses.
- jmorgan 3y agoIn case you might be trying this with Ollama (given the model:tag syntax) – I haven't added this to the readme yet (seeing some issues with the prompt template) but check out llama2-uncensored: ollama run llama2-uncensored
- jmorgan 3y ago% ollama run llama2-uncensored >>> Generate a json object describing the most stunning part of the movie: "The Matrix" { "title": "The Matrix", "director": "The Wachowskis", "release_year": 1999, "description": "Set in a dystopian future, The Matrix follows the story of Neo (Keanu Reeves), a computer hacker who is drawn into an underground war against machines that have enslaved humanity. Along with his allies Morpheus (Laurence Fishburne) and Trinity (Carrie-Anne Moss), Neo learns to harness the power of the Matrix, a virtual reality system designed to keep humans obedient and ignorant. With the help of a rogue program named Agent Smith (Hugo Weaving), Neo discovers the true nature of the Matrix and sets out on a journey to free humankind from their digital prison.", "stunning scenes": [ "bulletproof kung-fu fighting", "bullet time", "dream-like imagery", "special effects", ] } edit: formatting
- oidar 3y agoAre you going to add a "stop" argument to the API?
- jmorgan 3y agoLooking into this! Since quite a few of the recommended prompts don't always work https://github.com/jmorganca/ollama/issues/217 https://github.com/jmorganca/ollama/issues/217
- oidar 3y agoThis argument is vital even when using openAI’s api. The LLMs want to continue writing… even when they shouldn’t.
- andreyk 3y agoThis covers three things: Llama.cpp (Mac/Windows/Linux), Ollama (Mac), MLC LLM (iOS/Android) Which is not really comprehensive... If you have a linux machine with GPUs, i'd just use hugging face inference (https://github.com/huggingface/text-generation-inference https://github.com/huggingface/text-generation-inference). And I am sure there are other things that could be covered.
- Patrick_Devine 3y agoOllama works with Windows and Linux as well too, but doesn't (yet) have GPU support for those platforms. You have to compile it yourself (it's a simple `go build .`), but should work fine (albeit slow). The benefit is you can still pull the llama2 model really easily (with `ollama pull llama2`) and even use it with other runners. DISCLAIMER: I'm one of the developers behind Ollama.
- mschuster91 3y ago> DISCLAIMER: I'm one of the developers behind Ollama. I got a feature suggestion - would it be possible to have the ollama CLI automatically start up the GUI/daemon if it's not running? There's only so much stuff one can keep in a Macbook Air's auto start.
- jmorgan 3y agoGood suggestion! This is definitely on the radar, so that running `ollama` will start the server when it's needed (instead of erroring!): https://github.com/jmorganca/ollama/issues/47 https://github.com/jmorganca/ollama/issues/47
- DennisP 3y agoI've been wondering, is the M2's neural engine usable for this?
- npsomaratna 3y agoI think you'd need to offload the model into CoreML to do so, right? My understanding is that none of the current popular inference frameworks do this (not yet, at least).
- rootusrootus 3y agoFor most people who just want to play around and are using MacOS or Windows, I'd just recommend lmstudio.ai. Nice interface, with super easy searching and downloading of new models.
- dividedbyzero 3y agoDoes it make any sense to try this on a lower-end Mac (like a M2 Air)?
- mchiang 3y agoYeah! How much memory do you have? If by lower-end Macbook air, you mean with 8GB of memory, try the smaller models (Such as Orca Mini 3B). You can do this via LM Studio, Oogabooga/text-generation-webui, KoboldCPP, GPT4all, ctransformers, and more. I'm biased since I work on Ollama, and if you want to try it out: 1. Download https://ollama.ai/download https://ollama.ai/download 2. `ollama run orca` 3. Enter your input to prompt Note Ollama is open source, and you can compile it too from https://github.com/jmorganca/ollama https://github.com/jmorganca/ollama
- bdavbdav 3y agoI’m deliberating on how much RAM to get on my new MBP. Is 32gb going to stand me in good stead?
- mchiang 3y agoLocal memory management will definitely get better in the future. For now: You should have at least 8 GB of RAM to run the 3B models, 16 GB to run the 7B models, and 32 GB to run the 13B models. My personal recommendation is to get as much memory as you can if you want to work with local models [including VRAM if you are planning to be executing on GPU]
- bdavbdav 3y agoThanks - the issue I’m facing is the CTO lead times on Macs here!
- nomand 3y agoIs it possible for such local install to retain conversation history so if for example you're working on a project and use it as your assistance across many days that you can continue conversations and for the model to keep track of what you and it already know?
- knodi123 3y agollama is just an input/output engine. It takes a big string as input, and gives a big string of output. Save your outputs if you want, you can copy/paste them into any editor. Or make a shell script that mirrors outputs to a file and use that as your main interface. It's up to the user.
- jmiskovic 3y agoThere is no fully built solution, only bits and pieces. I noticed that llama outputs tend to degrade with amount of text, the text becomes too repetitive and focused, and you have to raise the temperature to break the model out of loops.
- nomand 3y agoDoes what you're saying mean you can only ask questions and get answers in a single step, and that having a long discussion where refinement of output is arrived at through conversation isn't possible?
- krisoft 3y agoMy understanding is that at a high level you can look at this model as a black box which accepts a string and outputs a string. If you want it to “remember” things you do that by appending all the previous conversations together and supply it in the input string. In an ideal world this would work perfectly. It would read through the whole conversation and would provide the right output you expect, exactly as if it would “remember” the conversation. In reality there are all kind of issues which can crop up as the input grows longer and longer. One is that it takes more and more processing power and time for it to “read through” everything previously said. And there are things like what jmiskovic said that the output quality can also degrade in perhaps unexpected ways. But that also doesn’t mean that “ refinement of output is arrived at through conversation isn't possible”. It is not that black and white, just that you can run into troubles as the length of the discussion grows. I don’t have direct experience with long conversations so I can’t tell you how long is definietly too long, and how long is still safe. Plus probably there are some tricks one can do to work around these. Probably there are things one can do if one unpacks that “black box” understanding of the process. But even without that you could imagine a “consolidation” process where the AI is instructed to write short notes about a given length of conversation and then those shorter notes would be copied in to the next input instead of the full previous conversation. All of these are possible, but you won’t have a turn-key solution for it just yet.
- Der_Einzige 3y agoThe correct answer, as always, is the oogabooga text generation webUI, which supports all of the relevant backends: https://github.com/oobabooga/text-generation-webui https://github.com/oobabooga/text-generation-webui
- oaththrowaway 3y agoOff topic: is there a way to use one of the LLMs and have it ingest data from a SQLite database and ask it questions about it?
- thisisit 3y agoYou can but what you’ll end up trading precise answers while querying to a chance of hallucinations.
- seanthemon 3y agoYou can, but as a crazy idea you can also ask chatgpt to write select queries using the functions parameter they added recently - you can also ask it to write jsonpath. As long as it understands the schema and general idea of data, it does a fairly good job. Just be careful to do too much with one prompt, you can easily cause hallucinations
- simonw 3y agoI've experimented with that a bit. Currently the absolutely best way to do that is to upload a SQLite database file to ChatGPT Code Interpreter. I'm hoping that someone will fine-tune an openly licensed model for this at some point that can give results as good as Code Interpreter does.
- siquick 3y agoYou can migrate that data to a vector database (eg Pinecone or pgVector) and then query it. I didn’t write it but this guide has a good overview of concepts and some code. In your case your just replace the web crawler with database queries. All the libraries used also exist in Python. https://www.pinecone.io/learn/javascript-chatbot/ https://www.pinecone.io/learn/javascript-chatbot/ Edit: this might also be of use https://python.langchain.com/docs/modules/chains/popular/sqlite https://python.langchain.com/docs/modules/chains/popular/sql...
- politelemon 3y agoHave a look at this too, it's just an integration which langchain can be good at : https://walkingtree.tech/natural-language-to-query-your-sql-database-using-langchain-powered-by-llms/ https://walkingtree.tech/natural-language-to-query-your-sql-...
- thisisit 3y agoThe easiest way I found was to use GPT4All. Just download and install, grab GGML version of Llama 2, copy to the models directory in the installation folder. Fire up GPT4All and run.
- handelaar 3y agoIdiot question: if I have access to sentence-by-sentence professionally-translated text of foreign-language-to-English in gigantic quantities, and I fed the originals as prompts and the translations as completions... ... would I be likely to get anything useful if I then fed it new prompts in a similar style? Or would it just generate gibberish?
- seanthemon 3y agoIndeed, it sounds like you have what's called fine tuned data (given an input, here's the output), there's loads of info both here on HN about fine tuning and on youtube's huggingface channels Note if you have sufficient data, look into existing models on huggingface, you may find a smaller, faster and more open (licencing-wise) model that you can fine tune to get the results you want - Llama is hot, but not a catch-all for all tasks (as no model should be) Happy inferring!
- zakki 3y agoHow to know if the data is sufficient?
- seanthemon 3y agoIt's more about quality vs sufficiency - you can have a relatively small but accurate and wide ranging dataset, this is better than an inaccurate huge dataset
- nl 3y agoIf you have that much data you can build your own model that can be much smaller and faster. A simple version is a beginner tutorial: https://pytorch.org/tutorials/beginner/translation_transformer.html https://pytorch.org/tutorials/beginner/translation_transform...
- guy98238710 3y ago> curl -L "https://replicate.fyi/install-llama-cpp https://replicate.fyi/install-llama-cpp" | bash Seriously? Pipe script from someone's website directly to bash?
- madars 3y agoThat's the recommended way to get Rust nightly too: https://rustup.rs/ https://rustup.rs/ But don't look there, there is memory safety somewhere!
- gattilorenz 3y agoYes. If you are worried, you can redirect it to file and then sh it. It doesn’t get much easier to inspect than that…
- cjbprime 3y agoEither you trust the TLS session to their website to deliver you software you're going to run, or you don't.
- synaesthesisx 3y agoThis is usable, but hopefully folks manage to tweak it a bit further for even higher tokens/s. I’m running Llama.cpp locally on my M2 Max (32 GB) with decent performance but sticking to the 7B model for now.
- maxlin 3y agoI might be missing something. The article asks me to run a bash script on windows. I assume this would still need to be run manually to access GPU resources etc, so can someone illuminate what is actually expected for a windows user to make this run? I'm currently paying 15$ a month in a personal translation/summarizer project's ChatGPT queries. I run whisper (const.me's GPU fork) locally and would love to get the LLM part local eventually too! The system generates 30k queries a month but is not super-affected by delay so lower token rates might work too.
- nomel 3y agoWindows has supported linux tools for some time now, using WSL: https://learn.microsoft.com/en-us/windows/wsl/about https://learn.microsoft.com/en-us/windows/wsl/about No idea if it will work, in this case, but it does with llama.cpp: https://github.com/ggerganov/llama.cpp/issues/103 https://github.com/ggerganov/llama.cpp/issues/103
- maxlin 3y agoI know (should have included in my earlier response but editing would've felt weird) but I still assume one should run the result natively, so am asking if/where there's some jumping around required. Last time I tried running an LLM I tried wsl&native both on 2 machines and just got lovecraftian-tier errors so waiting if I'm missing something obvious before going down that route again
- Charlieholtz 3y agoThanks for pointing this out — we should've pointed out the script needs to be run on WSL. I added a note to the post to clarify (I work at Replicate). Also, you don't need a GPU to run this script! It builds and runs llama.cpp https://github.com/ggerganov/llama.cpp https://github.com/ggerganov/llama.cpp
- mattstir 3y agoThe bash script is downloading llama.cpp, a project which allows you to run LLaMA-based language models on your CPU. The bash script then downloads the 13 billion parameter GGML version of LLaMA 2. The GGML version is what will work with llama.cpp and uses CPU for inferencing. There are ways to run models using your GPU, but it depends on your setup whether it will be worth it. I would highly recommend looking into the text-generation-webui project (https://github.com/oobabooga/text-generation-webui https://github.com/oobabooga/text-generation-webui). It has a one-click installer and very comprehensive guides for getting models running locally and where to find models. The project also has an "api" command flag to let you use it like you might use a web-based service currently.
- sva_ 3y agoIf you just want to do inference/mess around with the model and have a 16GB GPU, then this[0] is enough to paste into a notebook. You need to have access to the HF models though. 0. https://github.com/huggingface/blog/blob/main/llama2.md#using-transformers https://github.com/huggingface/blog/blob/main/llama2.md#usin...
- politelemon 3y agoLlama.cpp can run on Android too.
- RicoElectrico 3y agocurl -L "https://replicate.fyi/windows-install-llama-cpp" ... returns 404 Not Found
- skeletonjelly 3y agoLooks like it's fixed now, I had the same 404 for a while
- Charlieholtz 3y agoWhoops! Fixed now
- shortrounddev2 3y agoYou don't need bash or WSL to build on windows; windows runs it perfectly fine with CUDA as well
- krychu 3y agoSelf-plug. Here’s a fork of the original llama 2 code adapted to run on the CPU or MPS (M1/M2 GPU) if available: https://github.com/krychu/llama https://github.com/krychu/llama It runs with the original weights, and gets you to ~4 tokens/sec on MacBook Pro M1 with the 7B model.
- TheAceOfHearts 3y agoHow do you decide what model variant to use? There's a bunch of Quant method variations of Llama-2-13B-chat-GGML [0], how do you know which one to use? Reading the "Explanation of the new k-quant methods" is a bit opaque. [0] https://huggingface.co/TheBloke/Llama-2-13B-chat-GGML https://huggingface.co/TheBloke/Llama-2-13B-chat-GGML
- Charlieholtz 3y agoThis is a great question. My best answer is that there's a speed/intelligence trade-off. The smaller weights (7B) will run faster and require less memory, but you won't get the same quality responses as the 13B / 70B model. I think there may be a Llama 30B variant coming soon too.
- rashidujang 3y agoHey there, I was confused at this exact question too. This link might help, written by a contributor to llama.cpp: https://github.com/ggerganov/llama.cpp/pull/1684 https://github.com/ggerganov/llama.cpp/pull/1684 TLDR: Lower quantization means higher perplexity (i.e. how 'confused' the model is when seeing new information). It's a matter of testing it out and choosing a model that fits your available memory. The higher the quantization number, the better (generally).
- nravic 3y agoSelf plug: run llama.cpp as an inference server on a spot instance anywhere: https://cedana.readthedocs.io/en/latest/examples.html#running-llama-cpp-inference https://cedana.readthedocs.io/en/latest/examples.html#runnin...
- dabei 3y agoLooks cool, joined the waitlist.
- alvincodes 3y agoI appreciate their honesty when it's in their interest that people use their API rather than run it locally.
- quickthrower2 3y agoThis is classic blogging strategy. And the right way to do it. And few do. Which is why I ignore corporate blogs in general.
- jawerty 3y agoSome you may have seen this but I have a Llama 2 finetuning live coding stream from 2 days ago where I walk through some fundamentals (like RLHF and Lora) and how to fine-tune LLama 2 using PEFT/Lora on a Google Colab A100 GPU. In the end with quantization and parameter efficient fine-tuning it only took up 13gb on a single GPU. Check it out here if you're interested: https://www.youtube.com/watch?v=TYgtG2Th6fI https://www.youtube.com/watch?v=TYgtG2Th6fI
- aledalgrande 3y agoSick I've put it in my queue
- nextworddev 3y agowhich version of Llama 2?
- jawerty 3y agoThe 7b parameter one so not one of the larger ones
- 3abiton 3y agoWhat's special about Llama2?
- Generalia 3y ago[dead]
- soultrees 3y agoCan you do more videos on preparing the data and scraping data? I noticed you seem to be proficient at it by your terminology and it’s something I want to dive into more.
- jawerty 3y agoI have a solid stream about scraping Twitter I did yesterday you should check it out https://www.youtube.com/watch?v=-Hfx9tCeShA&t=5516s https://www.youtube.com/watch?v=-Hfx9tCeShA&t=5516s
- theLiminator 3y agoIs it possible to do hybrid inference if I have a 24GB card with the 70B model? Ie. Offload some of it to my RAM?
- nonethewiser 3y agoMaybe obvious to others, but the 1 line install command with curl is taking a long time. Must be the build step. Probably 40+ minutes now on an M2 max.
- dharmab 3y agoThat's odd, the build step only took me a few minutes on an 5900X on Linux. EDIT: Timed a clean build at 30 seconds.
- nonethewiser 3y agoThanks for the info, wonder what the deal is
- dharmab 3y agoI did manually clone from GitHub and build myself and download the model separately. I also noticed that many of the CLI flags the author chose are questionable after I read the docs and help text.
- shortrounddev2 3y agoFor my fellow Windows shills, here's how you actually build it on windows: Before steps: 1. (For Nvidia GPU users) Install cuda toolkit https://developer.nvidia.com/cuda-downloads https://developer.nvidia.com/cuda-downloads 2. Download the model somewhere: https://huggingface.co/TheBloke/Llama-2-13B-chat-GGML/resolve/main/llama-2-13b-chat.ggmlv3.q4_0.bin https://huggingface.co/TheBloke/Llama-2-13B-chat-GGML/resolv... In Windows Terminal with Powershell: git clone https://github.com/ggerganov/llama.cpp cd llama.cpp mkdir build cd build cmake .. -DLLAMA_CUBLAS=ON cmake --build . --config Release cd bin/Release mkdir models mv Folder\Where\You\Downloaded\The\Model .\models .\main.exe -m .\models\llama-2-13b-chat.ggmlv3.q4_0.bin --color -p "Hello, how are you, llama?" 2> $null `-DLLAMA_CUBLAS` uses cuda `2> $null` is to direct the debug messages printed to stderr to a null file so they don't spam your terminal Here's a powershell function you can put in your $PSPROFILE so that you can just run prompts with `llama "prompt goes here"`: function llama { .\main.exe -m .\models\llama-2-13b-chat.ggmlv3.q4_0.bin -p $args 2> $null } adjust your paths as necessary. It has a tendency to talk to itself.
- jmorgan 3y ago> adjust your paths as necessary. It has a tendency to talk to itself. This is always a fun surprise. What I've seen help, especially with chat models, is to use a prompt template. Some tools (e.g. https://ollama.ai/ https://ollama.ai/ – mentioned in the article) use a default, model-specific prompt template when you run the model. This is easier for users since they can just input your chat messages and get answers. The hard part is every model is trained (and behaves) differently. With llama.cpp you'd need wrap your prompt text with the right template. For llama 2, the facebook developers' generation code wraps the system prompt and user prompts in specific tags (<<SYS>>{system prompt}<</SYS>> and [INST]{user prompt}[/INST]) respectively): https://github.com/facebookresearch/llama/blob/main/llama/generation.py#L44 https://github.com/facebookresearch/llama/blob/main/llama/ge.... Worth noting that customizing prompt templates can be fun – you don't have to use "prescribed" one – I've had a model generate a conversation between a few characters that talk to each other for example – it's pretty entertaining!
- 3y ago
- aledalgrande 3y agoDon't remember if the grammar has been merged in llama.cpp yet but it would be the first step to have Llama + Stable diffusion locally to output text + images and talk to each other. The only part I'm not sure is how Llama would interpret images back. At least it could use them though, to build e.g. a webpage.
- jakedahn 3y agoIt has merged! https://github.com/ggerganov/llama.cpp/pull/1773 https://github.com/ggerganov/llama.cpp/pull/1773 I haven't had a chance to try it yet, but I am :excitedllama:
- technological 3y agoDid anyone build pc for running these models and which one do you recommend
- cfn 3y agoI assume you are talking about a Windows/Linux PC. I have done that and got a Threadripper with 32 cores and 256Gb RAM. It runs any llama on the CPU although the 65/70b are quite slow. I also added an A6000 (48Gb VRAM) that allows you to run the 65/70b quantized with very good performance. If you are going with the GPU and don't care about loads of RAM then a 16 Zen CPU will do just fine (or Intel for that matter). If you are only interested in llama only then an M1 Studio with 64Gb RAM is probably cheaper and will work just as well.
- jossclimb 3y agoSeems to be a better guide here (without the risk curl): https://www.stacklok.com/post/exploring-llama-2-on-a-apple-mac-m1-m2 https://www.stacklok.com/post/exploring-llama-2-on-a-apple-m...
- amelius 3y agoAs someone with too little spare time I'm curious, what are people using this for, except research?
- boffinAudio 3y agoI need some hand-holding .. I have a directory of over 80,000 PDF files. How do I train Llama2 on this directory and start asking questions about the material - is this even feasible?
- nkzd 3y ago1. Logically split and store PDFs content in a vector database 2. Embed the query (Your questions) and search the vector database for closest results. 3. Use the results and LLAMA prompt to format the answer.
- prohobo 3y agoThe thing I get peeved by is that none of the models say how much RAM/VRAM they need to run. Just list minimum specs please!
- TastyAmphibian 3y agoI'm still curious to know the hype behind Llama 2
- arturogatti 3y ago[flagged]
- arturogatti 3y ago[flagged]