8 ms·
Continue with LocalAI: An alternative to GitHub's Copilot that runs locally
- TekMol 3y agoA demo would be nice.
- Jeff_Brown 3y agoThe funny thing about the commercial model for code-helping AI is programmers are unusually capable of running their own AI, and also unusually concerned with digital privacy, so as soon as open-source alternatives are good enough, this market seems likely to evaporate. But I don't know if there's enough good public data for open source models to get there.
- bavell 3y agoGreat observation!
- npsomaratna 3y agoYou don't need so many layers of stuff (or API keys, signups, or other nonsense). Llama.cpp (to serve the model) + the Continue VS Code extension are enough. The rough list of steps to do so are: Part A: Install llama.cpp and get it to serve the model: -------------------------------------------------------- 1. Install the llama.cpp repo and run make. 2. Download the relevant model (e.g. wizardcoder-python-34b-v1.0.Q4_K_S.gguf). 3. Run the llama.cpp server (e.g., ./server -t 8 -m models/wizardcoder-python-34b-v1.0.Q4_K_S.gguf -c 16384 --mlock). 4. Run the OpenAI like API server [also included in llama.cpp] (e.g., python ./examples/server/api_like_OAI.py). Part B: Install Continue and connect it to llama.cpp's OpenAI like API: ----------------------------------------------------------------------- 5. Install the Continue extension in VS Code. 6. In the Continue extension's sidebar, click through the tutorial and then type /config to access the configuration. 7. In the Continue configuration, add "from continuedev.src.continuedev.libs.llm.ggml import GGML" at the top of the file. 8. In the Continue configuration, replace lines 57 to 62 (or around) with: models=Models( default=GGML( max_context_length=16384, server_url="http://localhost:8081" ) ), 9. Restart VS Code, and enjoy! You can access your local coding LLM through the Continue sidebar now.
- vanillax 3y agowodner if you can pair with https://github.com/getumbrel/llama-gpt https://github.com/getumbrel/llama-gpt
- noiv 3y agoThx. Where can I send flowers to?
- ignoramous 3y agoTo any person you're in a position to be kind to.
- redox99 3y agoIs there a way to make it work with ooba+exllama? (much faster than llamacpp)
- thelastparadise 3y agoYou should be able to turn on the API in booba: https://github.com/oobabooga/text-generation-webui#api https://github.com/oobabooga/text-generation-webui#api
- redox99 3y agoBut that API isn't OpenAI compatible AFAIK
- k4rli 3y agoThanks, works nicely and easy to set up. Is it possible to use GPU for this? With R9 7900x and 32GB RAM it takes 15-30sec to generate response. I have a 6900XT which might be more suited for this.
- npsomaratna 3y agoYes. In the llama.cpp server command, specify the number of layers you'd like offloaded to your GPU via the -ngl parameter, e.g.: ./server -t 8 -m models/wizardcoder-python-34b-v1.0.Q4_K_S.gguf -c 16384 --mlock -ngl 60 (You might need to play around with the number of layers.) [Edit: make sure to compile llama.cpp with GPU support first, e.g., "make clean && LLAMA_CUBLAS=1 make -j"]
- willsmith72 3y agoCan anyone share how the computer's performance is impacted by running the model locally? And what your specs are?
- npsomaratna 3y agoM1 with 32 GB RAM. I can just about fit the 4-bit quantized 33 GB Code Llama model (and it's finetunes, e.g. WizardCoder, etc.) into memory. It's somewhat slow, but good enough for my purposes. Edit: when I bought my Macbook in 2021, I was like "Ok, I'll just take the base model and add another 16 GB of RAM. That should future proof it for at least another half-decade." Famous last words.
- filmgirlcw 3y agoThis is why my rule for laptops with non-upgradable memory has been to max out the RAM at purchase -- and that has been my rule since 2012/2013 or whenever that trend really started. (written from a 64GB M1 Pro Max)
- joombaga 3y agoI wish they offered that much in the Air.
- SparkyMcUnicorn 3y ago34B Q4 will use around 20GB of memory. If it's running slow, make sure metal is actually being used[0]. You can get as much as a 50-100% boost in tokens/s, if by chance it's not enabled. I'm averaging 7 to 8 tokens/s on an M1 Max 10 core (24 GPU cores). [0] if using llama-cpp-python (or text-generation-webui, ollama, etc) try: `pip uninstall llama-cpp-python && CMAKE_ARGS="-DLLAMA_METAL=on" FORCE_CMAKE=1 pip install llama-cpp-python`
- npsomaratna 3y agoThank you. I had to reduce the context length to get this to work without crashing (from 16k to 8k)—and I'm seeing the ~100% speed up you mentioned. However, when I run the LLM, OSX becomes sluggish. I assume this is because the GPU's utilized to the point where hardware-based rendering slows down due to insufficient resources. I wonder if there's a way to avoid that slowdown?
- sailfast 3y agoOther options for this might include Code Llama (which runs on ollama locally) and looks interesting: https://about.fb.com/news/2023/08/code-llama-ai-for-coding/ https://about.fb.com/news/2023/08/code-llama-ai-for-coding/
- jmorgan 3y agoContinue has a great guide on using the new Code Llama model launched by Facebook last week: https://continue.dev/docs/walkthroughs/codellama https://continue.dev/docs/walkthroughs/codellama Continue also works with various backends and fine-tuned versions of Code Llama. E.g. for a local experience with GPU acceleration on macOS, continue can be used with Ollama (https://github.com/jmorganca/ollama https://github.com/jmorganca/ollama): ollama pull codellama from continuedev.src.continuedev.libs.llm.ollama import Ollama config = ContinueConfig( models=Models( default=Ollama(model="wizardcoder:34b-python") ) )
- hcrisp 3y agoOllama only works on Mac. Here is a portable option: https://github.com/xnul/code-llama-for-vscode https://github.com/xnul/code-llama-for-vscode
- mchiang 3y agoPeople have been compiling Ollama to run on Linux. The reason why it's not packaged yet for Linux is due to packaging it with GPU support - at the very least with nvidia support. Almost there!
- niux 3y agoHow does it actually compare to GitHub Copilot?
- rjmacarthy 3y agohttps://github.com/rjmacarthy/twinny https://github.com/rjmacarthy/twinny