40 ms·
Show HN: LlamaGPT – Self-hosted, offline, private AI chatbot, powered by Llama 2
- albert_e 3y agoOh I thought this was a quick guide to host it on any server (AWS / other clouds) of our choosing.
- mayankchhabra 3y agoYes! It can run on any home server or cloud server.
- ryanSrich 3y agoInteresting. I might try to get this to work on my NAS.
- reneberlin 3y agoGood luck! The token/sec will be under your expectations or it will overheat. You really shouldn't play games with your data-storage. You could try it with an old laptop to see how bad it performs. Ruining your NAS for this is a bit over the top to show, that "it worked somehow". But i don't know, maybe your NAS has a powerful processor and is tuned to the max and you have redundancy and don't care to loose a NAS? Or this was just a joke and i fell for it! ;)
- mayankchhabra 3y agoNot sure how powerful their NAS is, but on Umbrel Home (which has an N5105 CPU), it's pretty useable with ~3 tokens generated per second.
- samspenc 3y agoI had the same question initially, was a bit confused by the Umbrel reference at the top, but there's a section right below it titled "Install LlamaGPT anywhere else" which I think should work on any machine. As an aside, UmbrelOS actually seems like a cool concept by itself btw, good to see these "self hosted cloud" projects coming together in a unified UI, I may investigate this more at some point.
- QuinnyPig 3y agoI've been looking for something like this for a while. Nice!
- belval 3y agoNice project! I could not find the information in the README.md, can I run this with a GPU? If so what do I need to change? Seems like it's hardcoded to 0 in the run script: https://github.com/getumbrel/llama-gpt/blob/master/api/run.sh#L12 https://github.com/getumbrel/llama-gpt/blob/master/api/run.s...
- crudgen 3y agoHad the same thought, since it is kinda slow (only have 4 pyhsical/8 logical cores though). But I think vRAM might be a problem (8gb can work, if one has a rather recent gpu (here m1/2 might be interesting)).
- mayankchhabra 3y agoAh yes, running on GPU isn't supported at the moment. But CUDA (for Nvidia GPUs) and Metal support is on the roadmap!
- samspenc 3y agoAh fascinating, just curious, what's the technical blocker? I thought most of the Llama models were optimized to run on GPUs?
- mayankchhabra 3y agoIt's fairly straightforward to add GPU support when running on the host, but LlamaGPT runs inside a Docker container, and that's where it gets a bit challenging.
- chasd00 3y agois it a free model or is the politically-correct-only response constraints in place?
- benreesman 3y agoI’m a little out of date (busy few weeks), didn’t the Vicuna folks un-housebreak the LLaMA 2 language model (which is world class) with a slightly less father-knows-best Instruct tune?
- Havoc 3y agoLlama is definitely "censored" though I've not found this to be an issue in practice. Guess it depends on what you want to do with it
- mayankchhabra 3y agoIt's powered by Nous Hermes Llama2 7b. From their docs: "This model stands out for its long responses, lower hallucination rate, and absence of OpenAI censorship mechanisms. [...] The model was trained almost entirely on synthetic GPT-4 outputs. Curating high quality GPT-4 datasets enables incredibly high quality in knowledge, task completion, and style."
- lee101 3y ago[dead]
- Atlas-Marbles 3y agoVery cool, this looks like a combination of chatbot-ui and llama-cpp-python? A similar project I've been using is https://github.com/serge-chat/serge https://github.com/serge-chat/serge. Nous-Hermes-Llama2-13b is my daily driver and scores high on coding evaluations (https://huggingface.co/spaces/mike-ravkine/can-ai-code-results https://huggingface.co/spaces/mike-ravkine/can-ai-code-resul...).
- netdur 3y agono not llama-cpp-python, it uses llama.cpp's built in server.
- lazzlazzlazz 3y ago(1) What are the best more creative/less lobotomized versions of Llama 2? (2) What's the best way to get one of those running in a similarly easy way?
- mritchie712 3y agotry llama2-uncensored https://github.com/jmorganca/ollama https://github.com/jmorganca/ollama
- lkbm 3y agohttps://github.com/jmorganca/ollama https://github.com/jmorganca/ollama was extremely simple to get running on my M1 and has a couple uncensored models you can just download and use.
- brucemacd 3y agohttps://github.com/jmorganca/ollama/tree/main/examples/privategpt https://github.com/jmorganca/ollama/tree/main/examples/priva... there's an example using PrivateGPT too
- dealuromanet 3y agoIs it private and offline via ollama? Are all ollama models private and offline?
- brucemacd 3y agoYes, they are private and offline in the sense that they are running entirely locally and do not send any information off your local system.
- freedomben 3y agoThe uncensored model isn't very uncensored. It refused a number of test prompts for me, telling me that things were unsafe and telling me to consult a professional
- caesil 3y agoSo many projects still using GPT in their name. Is the thinking here that OpenAI is not going to defend that trademark? Or just kicking the can down the road on rebranding until the C&D letter arrives?
- schappim 3y agoThey don’t have the trademark yet. OpenAI has applied to the United States Patent and Trademark Office (USPTO) to seek domestic trademark registration for the term “GPT” in the field of AI.[64] OpenAI sought to expedite handling of its application, but the USPTO declined that request in April 2023.
- khaledh 3y agoThis reminds me of the first generation of computers in the 40s and early 50s following the ENIAC: EDSAC, EDVAC, BINAC, UNIVAC, SEAC, CSIRAC, etc. It took several years for the industry to drop this naming scheme.
- super256 3y agoWell, GPT is simply an initialism for "Generative Pre-trained Transformer". In Germany, a trademark can be lost if it becomes a "Gattungsbegriff" (generic term). This happens when a trademark becomes so well-known and widely used that it becomes the common term for a product or service, rather than being associated with a specific company or brand. For example, if a company invented a new type of vacuum cleaner and trademarked the name, but then people started using that name to refer to all vacuum cleaners, not just those made by the company, the trademark could be at risk of becoming a generic term; which would lead to a deletion of the trademark. I think this is basically what happens to GPT here. Btw, there are some interesting exampls from the past were trademarks were lost due to the brand name becoming too popular: Vaseline and Fön (hairdryer; everyone in Germany uses the term "Fön"). I also found some trademarks which are at risk of being lost: "Lego", "Tupperware", "Post" (Deutsche Post/DHL), and "Jeep". I don't know how all this stuff works in America though. But it would honestly suck if you'd approve such a generic term as a trademark :/
- raffraffraff 3y ago
- ccozan 3y agoOk, since is running all private, how can I add my own private data? For example I have a 20+ years of an email archive that I'd like to be ingested.
- rdedev 3y agoA simple way would be to do some form of retrieval on those emails and add those back to the original prompt
- cromka 3y agoI imagine this means you’d need to come up with own model, even if based on existing one.
- ravishi 3y agoAnd is that hard? Sorry if this is a newbie question, I'm really out of the loop on this tech. What would be required? Computing power and tagging? Or can you like improve the model without much human intervention? Can it be done incrementally with usage and user feedback? Would a single user even be able to generate enough feedback for this?
- phillipcarter 3y agoYes, this would be quite hard. Fine-tuning an LLM is no simple task. The tools and guidance around it are very new, and arguably not meant for non-ML Engineers.
- aryamaan 3y agoWhat are some ways people to get familiar with machine learning engineering who are also working adults
- influxmoment 3y agoThat would require custom training. This project only does inference
- SubiculumCode 3y agoI didn't see any info on how this is different than installing/running llamacpp or koboldcpp. New offerings are awesome of course, but what is it adding?
- mayankchhabra 3y agoThe main difference is setting everything up yourself manually, downloading the modal, optimizing the parameters for best performance, running an API server and a UI front-end - which is out of reach for most non-technical people. With LlamaGPT, it's just one command: `docker compose up -d` or one click install for umbrelOS home server users.
- SubiculumCode 3y agothanks. yeah, that IS useful. Anyone see if it contains utilities to import models from huggingface/github?
- DrPhish 3y agoMaybe I've been at this for too long and can't see the pitfalls of a normal user, but how is that easier than using an oobabooga one-click installer (an option that's been around "forever")? I guess ooba one-click doesn't come with a model included, but is that really enough of a hurdle to stop someone from getting it going? Maybe I'm not seeing the value proposition of this. Glad to be enlightened!
- ShamelessC 3y agoThe difference is that this project has both "GPT" and "llama" in its name, and used the proper HN-bait - "self hosted, offline, private". HN users (mostly) don't actually read or check anything and upvote mostly based on titles and subsequent early comments.
- Multicomp 3y agoAgreed. Gpt4all[1] offers a similar 'simple setup' but with application exe downloads, but is arguably more like open core because the gpt4all makers (nomic?) want to sell you the vector database addon stuff on top. [1]https://github.com/nomic-ai/gpt4all https://github.com/nomic-ai/gpt4all I like this one because it feels more private / is not being pushed by a company that can do a rug pull. This can still do a rug pull, but it would be harder to do.
- synaesthesisx 3y agoHow this compare to just running llama.cpp locally?
- mayankchhabra 3y agoIt's an entire app (with a chatbot UI) that takes away the technical legwork to run the model locally. It's a simple one line `docker compose up -d` on any machine, or one click install on umbrelOS home servers.
- avivo 3y agoWhat is the advantage of this versus running something like https://github.com/simonw/llm https://github.com/simonw/llm , which also gives you options to e.g. use https://github.com/simonw/llm-mlc https://github.com/simonw/llm-mlc for accelerated inference?
- stormfather 3y agoWhich layers are best to use as vector embeddings? Is it the initial embedding layer afer tokenization? First hidden layer? Second?