5 ms·
The answer to your question is: ollama run mixtral That's it. You're running a local LLM. I have no clue how to run llama.cpp I got Stable Diffusion running
by nerdix 3y ago
The answer to your question is:
ollama run mixtral
That's it. You're running a local LLM. I have no clue how to run llama.cpp
I got Stable Diffusion running and I wish there was something like ollama for it. It was painful.
- viraptor 3y agoOn a mac, https://drawthings.ai https://drawthings.ai is the ollama of Stable Diffusion.
- ghurtado 3y agoFor me, ComfyUI made the process of installing and playing with SD about as simple as a Windows installer.
- jameshart 3y agoThe README is pretty clear, albeit it talks about a lot of optional steps you don’t need, but it’s essentially gonna be something like: git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp make wget https://huggingface.co/TheBloke/Mixtral-8x7B-v0.1-GGUF/resolve/main/mixtral-8x7b-v0.1.Q4_K_M.gguf?download=true ./main -m ./mixtral-8x7b-v0.1.Q4_K_M.gguf -n 128
- vidarh 3y agoLast time I tried llama.cpp I got errors when running make that were way too time consuming to bother tracking down. It's probably a simple build if everything is how it wants it, but it wasn't in my machine, while running ollama was.
- verdverm 3y agoThis shows the value ollama provides I only need to know the model name and then run a single command
- eclectic29 3y agoAnd what will you do after trying it? Sure, you saved a few mins in trying out a model or models. What next?
- verdverm 3y agoI focus on building the application rather than figuring out someone else preferred method for how I should work? I use Docker Compose locally, Kubernetes in the cloud I run in hot-reload locally, I build for production I often nuke my database locally, but I run it HA in production It is very rare to use the same technology locally (or the same way) as in production
- ramblerman 3y agoRelax. Not everything in this world was built exactly for you. You almost seem to have a problem with this.
- imtringued 3y agoThere is no "next", there is a whole world of people running LLMs locally on their computer and they are far more likely to switch between models on a whim every few days.
- jameshart 3y agoIt should be fairly obvious that one can find alternative models and use them in the above command too. Look, I’m not arguing that a prebuilt binary that handles model downloading has no value over a source build and manually pulling down gguf files. I just want to dispel some of the mystery. Local LLM execution doesn’t require some mysterious voodoo that can only be done by installing and running a server runtime. It’s just something you can do by running code that loads a model file into memory and feeds tokens to it. More programmers should be looking at llama.cpp language bindings than at Ollama’s implementation of the openAI api.
- verdverm 3y agoI'd rather focus on building on top of of LLMs than going lower level Ollama makes that super easy. I tried llama.cpp first and hit build issues. Ollama worked out of the box
- kergonath 3y agoCompared to “ollama pull mixtral”? And then actually using the thing is easier as well.
- ies7 3y agoFor us this may like a walk in the park. For non technical people there is a possibility their os don't have git, wget and c++ compiler (especially in windows) This is just like dropbox case years ago.
- cjbprime 3y agoThis will likely build a version without GPU acceleration, I think?
- jameshart 3y agoBuilds with Metal support on my Mac M2
- UncleEntity 3y agoI was trying to get AMD GPU support going in llama.cpp a couple weeks ago and just gave up after a while. 'rocminfo' shows that I have a GPU and, presumably, rocm installed but there were build problems I didn't feel like sorting out just to play with a LLM for a bit. Kudos if Ollama has this sorted out.
- imtringued 3y agoThe average user isn't going to compile llama.cpp. They will either download a fully integrated application that contains llama.cpp and is able to read gguf files directly, like kobold.cpp or they are going to use any arbitrary front end like Silly Tavern which needs to connect to an inference server via an API and ollama is one of the easier inference servers to install and use.
- icelain 3y agoCheck out EasyDiffusion.