15 ms·
An LLM playground you can run on your laptop
- underlines 4y agoAwesome! Does it support safetensors, new ggml format, triton or cuda model files? I added the playground to the GUI list here: to https://github.com/underlines/awesome-marketing-datascience/blob/master/awesome-ai.md https://github.com/underlines/awesome-marketing-datascience/...
- dventimihasura 3y agoWhy does it use port 5432 by default? That's the default port for PostgreSQL. Does this use PostgreSQL?
- kristjansson 4y agoAn LLM playground whose UI you can run on your laptop.
- Cyphase 4y agoIt supports local models.
- Zetobal 4y agoI wonder how people that don't read properly before doing something will fare in a world of text interfaces/ai.
- Lockal 4y agoThey will answer "Please don't comment on whether someone read an article, please review https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html (actually, a bad excuse, but it is what it is).
- manojlds 4y agoBetter? Because the AI will help them?
- groby_b 4y agoDoes it matter who generated the text people don't bother to read? And how will lack of attention to detail affect prompt generation?
- psychphysic 4y agoAre there any good local models? Gpt-2 is pants once you've gotten used to 4.
- danielbln 4y agoLlama, Alpaca, Dolly, Vicuna
- psychphysic 4y agoI just tried llama (30B) and alpaca. I'd say they are closer to GOT2.5
- newswasboring 4y agoGPT2 is ancient news. There are now local running models which, allegedly, can reach the same performance as GPT-4. Look up llama.cpp[1] and all the various community generated models. [1] https://github.com/ggerganov/llama.cpp https://github.com/ggerganov/llama.cpp
- napier 4y agoVery much allegedly. They’re good, but not great.
- holografix 4y agoDoes llama.cop not run on GPU?
- MacsHeadroom 4y agollama.cpp is CPU only but llama runs on GPU using the HuggingFace Transformers library. You can run the llama model which far outpaces GPT-3.5 on a couple of $200 Tesla P40 GPUs at faster speeds than GPT-3.5 Turbo, completely locally.
- LoganDark 4y agoIt doesn't, it's CPU-only. -Emily
- nl 4y agoNo, this is running LLMs on your laptop, although it can also connect to remote ones. Eg: "Automatically detects local models in your HuggingFace cache, and lets you install new ones." Llamma.cpp and HuggingFace models are all local.
- kristjansson 4y agoMy bad, I was wrong to miss the huggingface support.
- penny10k 4y agoIts amazing can run Alpaca llama 33B parameters, totally can handle japanese and korean where the earlier ones like 7B parameters could only do english (any other languages was horrible). All able to run on my M1 macbook.
- glandium 4y agoHow good is it with Japanese compared to GPT?
- a5huynh 4y agoAs an alternative for purely local LLMs, I've been having fun with this setup: https://github.com/oobabooga/text-generation-webui https://github.com/oobabooga/text-generation-webui
- specproc 4y agoThe oobabooga setup feels a lot more mature and has a larger community. Skimming OP's repo, there seems like a lot of fiddling and faffing with JSON to get things running. Nice stuff all the same.
- simion314 4y agoDo you know what is the way to find models compatible with oobabooga/text-generation-webui ? I downloaded one with the included script and that worked, but if I try different ones it seems there are so many formats so no idea say how do I search huggingface or google to find the correct format. I would like to try this new quantized LLAMA versions with the GUI, I can run them in the CLI on the CPU but llama.cpp uses ggml formats .
- enlyth 4y agohttps://rentry.org/nur779 https://rentry.org/nur779 (scroll down to the Huggingface ones)
- xenodium 4y agoNeat, eventually would like to run a purely local LLM Emacs shell https://github.com/xenodium/chatgpt-shell https://github.com/xenodium/chatgpt-shell. For now ChatGPT only, but working on making more generic/reusable.
- la64710 4y agoAwesome !! Great … I haven’t gone through it yet but from just what I saw on the GitHub page it would be awesome to have a CLI with standard arguments built into it for everything that can be done through the flask interface. Thanks !!
- Cyphase 4y agoSomeone with more inclination than I at the moment might be able to say something interesting about this being from Nat Friedman (former CEO of GitHub).
- amrb 4y agoGood to see a CEO still coding
- quickthrower2 4y agoHe wants my $5 though!
- fahrradflucht 4y agoHe didn’t build this. This is a repl.co bounty project he paid for.
- dark-star 4y agoDang, I don't have a laptop, can I also run it on my desktop...? ;-)
- Tepix 4y agoYou can even upgrade it to 128GB cheaply nowadays and run the much bigger models (LLaMA 65B 16bit).
- yreg 4y agoBy laptop they mean you don't need a beefy gaming PC with 200 RGB lights.
- amrb 4y agoCan it run Crysis?
- manojlds 4y agoAs long as it you take it and keep it on your laps while running this, should be fine.
- Takennickname 4y agoEvery desktop is a laptop if you're strong enough.
- ted_bunny 4y agoAnd every laptop is a desktop if you're one of those deskful types.
- ryan-allen 4y agoThanks for this! The compare feature is very cool, especially being able to play around with GPT4 settings (I'm still on the waiting list for GPT4 API so having access to this now is fantastic).
- d4rkp4ttern 4y agoThis is very neat, thanks for sharing. I was wondering about a related thing — is there a way to query a llama.cpp (or other such local model) via an API from Python? In other words, I see a lot of cool applications being built with langchain + ClosedAPI, so I’m wondering if an API call to a local model could be a drop-in replacement for the ClosedAPI call?
- nl 4y agoThere's a llama.cpp fork (I think?) with a build in HTTP server for an API. I'm on mobile and can't find it right now though.
- mark_l_watson 4y agoI have a short example in my recently published book [1] on downloading a HuggingFace model and using it locally with LangChain. [1] https://leanpub.com/langchain https://leanpub.com/langchain EDIT: GitHub repo https://github.com/mark-watson/langchain-book-examples https://github.com/mark-watson/langchain-book-examples
- MacsHeadroom 4y agoYes, there are python bindings for llama.cpp and the text-generation-webui already uses them for local inference. https://github.com/oobabooga/text-generation-webui/wiki/llama.cpp-models https://github.com/oobabooga/text-generation-webui/wiki/llam... "pip install llamacpp" or https://github.com/thomasantony/llamacpp-python https://github.com/thomasantony/llamacpp-python
- boppo1 4y agoI think langchain lets you do this.
- DougBTX 4y agoIn principle you can use subprocess.run with llama.cpp (especially now that the mmap patch has landed, so model load time is negligible on subsequent runs) and then use stdin and stdout to interact with it. Multiple local sessions should work too. I’ve not tried it out yet... though I was looking for an afternoon project to work on!