9 ms·
I'm not sure why Ollama garners so much attention. It has limited value - used for only experimenting with models + cannot support more than 1 model at a time.
by eclectic29 3y ago
I'm not sure why Ollama garners so much attention. It has limited value - used for only experimenting with models + cannot support more than 1 model at a time. It's not meant for production deployments. Granted that it makes the experimentation process super easy but for something that relies on llama.cpp completely and whose main value proposition is easy model management I'm not sure it deserves the brouhaha people are giving it.
Edit: what do you do after the initial experimentation? you need to deploy these models eventually to production. I'm not even talking about giving credit to llama.cpp, just mentioning that this product is gaining disproportionate attention and kudos compared to the value it delivers. Not denying that it's a great product.
- reustle 3y agoEven for just running a model locally, Ollama provided a much simpler "one click install" earlier than most tools. That in itself is worth the support.
- Aka456 3y agoKoboldcpp is also very, very good, plug and play, very complet web UI, nice little api with sse text streaming, vulkan accelerated, have an AMD fork...
- crooked-v 3y ago> Granted that it makes the experimentation process super easy That's the answer to your question. It may have less space than a Zune, but the average person doesn't care about technically superior alternatives that are much harder to use.
- thejohnconway 3y ago*Nomad And lame.
- davidhariri 3y agoAs it turns out, making it faster and better to manage things tends to get people’s attention. I think it’s well deserved.
- vikramkr 3y agoIt's nice for personal use which is what I think it was built for, has some nice frontend options too. The tooling around it is nice, and there are projects building in rag etc. I don't think people are intending to deploy days services through these tools
- andrewstuart 3y agoSounds like you are dismissing Ollama as a "toy". Refer: https://paulgraham.com/startupideas.html https://paulgraham.com/startupideas.html
- nerdix 3y agoThe answer to your question is: ollama run mixtral That's it. You're running a local LLM. I have no clue how to run llama.cpp I got Stable Diffusion running and I wish there was something like ollama for it. It was painful.
- viraptor 3y agoOn a mac, https://drawthings.ai https://drawthings.ai is the ollama of Stable Diffusion.
- ghurtado 3y agoFor me, ComfyUI made the process of installing and playing with SD about as simple as a Windows installer.
- jameshart 3y agoThe README is pretty clear, albeit it talks about a lot of optional steps you don’t need, but it’s essentially gonna be something like: git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp make wget https://huggingface.co/TheBloke/Mixtral-8x7B-v0.1-GGUF/resolve/main/mixtral-8x7b-v0.1.Q4_K_M.gguf?download=true ./main -m ./mixtral-8x7b-v0.1.Q4_K_M.gguf -n 128
- vidarh 3y agoLast time I tried llama.cpp I got errors when running make that were way too time consuming to bother tracking down. It's probably a simple build if everything is how it wants it, but it wasn't in my machine, while running ollama was.
- verdverm 3y agoThis shows the value ollama provides I only need to know the model name and then run a single command
- eclectic29 3y agoAnd what will you do after trying it? Sure, you saved a few mins in trying out a model or models. What next?
- cbhl 3y agoIn my opinion, pre-built binaries and an easy-to-use front-end are things that should exist and are valid as a separate project unto themselves (see, e.g., HandBrake vs ffmpeg). Using the name of the authors or the project you're building on can also read like an endorsement, which is not _necessarily_ desirable for the original authors (it can lead to ollama bugs being reported against llama.cpp instead of to the ollama devs and other forms of support request toil). Consider the third clause of BSD 3-Clause for an example used in other projects (although llama.cpp is licensed under MIT).
- airocker 3y agoMore than one is easy: put it behind a load balancer. Put one ollama in one container or one port.
- Zambyte 3y agoThat is still one model per instance of Ollama, right?
- airocker 3y agoyes, not sure you can do better than that. You cannot still have one instance of LLM in (GPU) memory answer two queries at one time.
- eclectic29 3y agoOf course, you can support concurrent requests. But Ollama doesn't support it and it's not meant for this purpose and that's perfectly ok. That's not the point though. For fast/perf scenarios, you're better off with vllm.
- airocker 3y agoThanks! This is great to know.
- eclectic29 3y agoFWIW Ollama has no concurrency support even though llama.cpp's server component (the thing that Ollama actually uses) supports it. Besides, you can't have more than 1 model running. Unloading and loading models is not free. Again, there's a lot more and really much of the real optimization work is not in Ollama; it's in llama.cpp which is completely ignored in this equation.
- airocker 3y agoThanks! Great to know. I did not know llama.cpp could do this. It should be pretty straight forward to support, not sure why they would not do it.
- Karrot_Kream 3y agoI mean, it takes something difficult like an LLM and makes it easy to run. It's bound to get attention. If you've tried to get other models like BERT based models to run you'll realize just how big the usability gains are running ollama than anything else in the space. If the question you're asking is why so many folks are focused on experimentation instead of productionizing these models, then I see where you're coming from. There's the question of how much LLMs are actually being used in prod scenarios right now as opposed to just excited people chucking things at them; that maybe LLMs are more just fun playthings than tools for production. But in my experience as HN has gotten bigger, the number of posters talking about productionizing anything has really gone down. I suspect the userbase has become more broadly "interested in software" rather than "ships production facing code" and the enthusiasm in these comments reflects those interests. FWIW we use some LLMs in production and we do not use ollama at all. Our prod story is very different than what folks are talking about here and I'd love to have a thread that focuses more on language model prod deployments.
- jart 3y agoWell you would be one of the few hundred people on the planet doing that. With local LLMs we're just trying to create a way for everyone else to use AI that doesn't require sharing all their data with them. First thing everyone asks for of course is how to turn the open source local llms into their own online service.
- Karrot_Kream 3y agoOllama's purpose and usefulness is clear. I don't think anyone is disputing that nor the large usability gains ollama has driven. At least I'm not. As far as being one of the few hundred on the planet, well yeah that's why I'm on HN. There's tons of publications and subreddits and fora for generic tech conversation. I come here because I want to talk about the unknowns.
- kergonath 3y ago> I come here because I want to talk about the unknowns. Your knowns an are unknowns to some people and vice versa. This is a great strength of HN; on a whole lot of subject you’ll find people ranging from enthusiastic to expert. There are probably subreddits or discord servers tailored to narrow niches and that’s cool, but HN is not that. They are complementary, if anything. In contrast, HN is much more interesting and with a much better S/N ratio than generic tech subreddits, it’s not even comparable.
- evilduck 3y agoI am 100% uninterested in your production deployment of rent seeking behavior for tools and models I can run myself. Ollama empowers me to do more of that easier. That’s why it’s popular.
- jameshart 3y agoOP’s point is more that Ollama isn’t what’s doing the empowering. Llama.cpp is.
- brucethemoose2 3y agoInterestingly, Ollama is not popular at all in the "localllama" community (which also extends to related discords and repos). And I think thats because of capabilities... Ollama is somewhat restrictive compared to other frontends. I have a littany of reasons I personally wouldn't run it over exui or koboldcpp, both for performance and output quality. This is a necessity of being stable and one-click though.
- elwebmaster 3y agoYou are not making any sense. I am running ollama and Open WebUI (which takes care of auth) in production.
- ecnahc515 3y agoOllama is the Docker of LLMs. Ollama made it _very_ easy to run LLMs locally. This is surprisingly not as easy as it seems, and incredibly useful.
- kergonath 3y ago> It's not meant for production deployments. I am probably not the demographics you expect. I don’t do “production” in that sense, but I have ollama running quite often when I am working, as I use it for RAG and as a fancy knowledge extraction engine. It is incredibly useful: - I can test a lot of models by just pulling them (very useful as progress is very fast), - using their command line is trivial, - the fact that it keeps running in the background means that it starts once every few days and stays out of the way, - it integrates nicely with langchain (and a host of other libraries), which means that it is easy to set up some sophisticated process and abstract away the LLM itself. > what do you do after the initial experimentation? I just keep using it. And for now, I keep tweaking my scripts but I expect them to stabilise at some point, because I use these models to do some real work, and this work is not monkeying about with LLMs. > I'm not even talking about giving credit to llama.cpp, just mentioning that this product is gaining disproportionate attention and kudos compared to the value it delivers. For me, there is nothing that comes close in terms of integration and convenience. The value it delivers is great, because it enables me to do some useful work without wasting time worrying about lower-level architecture details. Again, I am probably not in the demographics you have in mind (I am not a CS person and my programming is usually limited to HPC), but ollama is very useful to me. Its reputation is completely deserved, as far as I am concerned.
- rkwz 3y ago> I use it for RAG and as a fancy knowledge extraction engine Curious, can you share more details about your usecase?
- idncsk 3y agoTry ollama webui(now open-webui). Sry on my phone now => no links
- kergonath 3y agoThe use case is exploratory literature review in a specific scientific field. I have a setup that takes pdfs and does some OCR and layout detection with Amazon, and then bunch them with some internal reports. Then, I have a pipeline to write summaries of each document and another one to slice them into chunks, get embeddings and set up a vector store for a RAG chat bot. At the moment it’s using Mixtral and the command line. But I like being able to swap LLMs to experiments with different models and quantisation without hassle, and I more or less plan to set this up on a remote server to free some resources on my workstation so the web UI could come in handy. Running this locally is a must for confidentiality reasons. I’d like to get rid of Textract as well, but unfortunately I haven’t found a solution that’s even close. Tesseract in particular was very disappointing.
- Abishek_Muthian 3y agoFor me all the projects which enable running & fine-tuning LLMs locally like llama.cpp, ollama, open-webui, unsolth etc. play a very important part in democratizing AI. > what do you do after the initial experimentation? you need to deploy these models eventually to production I built GaitAnalyzer[1], to analyze my gait laptop; I had deployed it briefly in production when I had enough credits to foot the AWS GPU bills. Ollama made it very simple to deploy the application, Anyone who has used docker before can now run GaitAnalyzer in their computer. [1] https://github.com/abishekmuthian/gaitanalyzer https://github.com/abishekmuthian/gaitanalyzer