8 ms·
I Self-Hosted Llama 3.2 with Coolify on My Home Server
- seungwoolee518 2y agoGreat post! However, Do I need to Install CUDA toolkit on host? I haven't install CUDA toolkit when I use on Containerized platform (like docker)
- thangngoc89 2y agoYou don't need to install CUDA toolkit on host system. Nvidia driver + Nvidia container toolkit would do the job. You could check official instructions at [0] [0] https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/install-guide.html https://docs.nvidia.com/datacenter/cloud-native/container-to...
- seungwoolee518 2y agoThanks, I was bit confused on install a CUDA Toolkit on the Host. (Because I don't install any software except Driver && Toolkit)
- varun_ch 2y agoI’m curious about how good the performance with local LLMs is on ‘outdated’ hardware like the author’s 2060. I have a desktop with a 2070 super that it could be fun to turn into an “AI server” if I had the time…
- taosx 2y agoLast time I tried a local llm was about a year ago with a 2070S and 3950x and the performance was quite slow for anything beyond phi 3.5 and the small models quality feels worse than what some providers offer for cheap or free so it doesn't seem worth it with my current hardware. Edit: I've loaded llama 3.1 8b instruct GGUF and I got 12.61 tok/sec and 80tok/sec for 3.2 3b.
- magicalhippo 2y agoI've been playing with some LLMs like Llama 3 and Gemma on my 2080Ti. If it fits in GPU memory the inference speed is quite decent. However I've found quality of smaller models to be quite lacking. The Llama 3.2 3B for example is much worse than Gemma2 9B, which is the one I found performs best while fitting comfortably. Actual sentences are fine, but it doesn't follow prompts as well and it doesn't "understand" the context very well. Quantization brings down memory cost, but there seems to be a sharp decline below 5 bits for those I tried. So a larger but heavily quantized model usually performs worse, at least with the models I've tried so far. So with only 6GB of GPU memory I think you either have to accept the hit on inference speed by only partially offloading, or accept fairly low model quality. Doesn't mean the smaller models can't be useful, but don't expect ChatGPT 4o at home. That said if you got a beefy CPU then it can be reasonable to have it do a few of the layers. Personally I found Gemma2 9B quantized to 6 bit IIRC to be quite useful. YMMV.
- magicalhippo 2y agoYes, gemma-2-9b-it-Q6_K_L is the one that works well for me. I tried gemma-2-27b-it-Q4_K_L but it's not as good, despite being larger. Using llama.cpp and models from here[1]. [1]: https://huggingface.co/bartowski https://huggingface.co/bartowski
- whitefables 2y agoHere's how it looks like in real time: https://youtu.be/3vhJ6fNW8AI https://youtu.be/3vhJ6fNW8AI
- thisguyagain 2y agoWhat’d you use to record that? Looks really great.
- whitefables 2y agoScreen studio
- khafra 2y agoIf you want to set up an AI server for your own use, it's exceedingly easy to install LM Studio and hit the "serve an API" button. Testing performance this way, I got about 0.5-1.5 tokens per second with an 8GB 4bit quantized model on an old DL360 rack-mount server with 192GB RAM and 2 E5-2670 CPUs. I got about 20-50 tokens per second on my laptop with a mobile RTX 4080.
- taosx 2y agoLM studio is so nice, I'm up and running in 5 minutes. ty
- dtquad 2y agoI am using an old laptop with a GTX 1060 6 GB VRAM to run a home server with Ubuntu and Ollama. Because of quantization Ollama can run 7B/8B models on an 8 year old laptop GPU with 6 GB VRAM.
- nubinetwork 2y agoI'm happy with a Radeon VII, unless the model is bigger than 16gb...
- alias_neo 2y agoYou can get a relative idea here: https://developer.nvidia.com/cuda-gpus https://developer.nvidia.com/cuda-gpus I use a Tesla P4 for ML stuff at home, it's equivalent to a 1080 Ti, and has a score of 7.1. A 2070 (they don't list the "super") is a 7.5. For reference, 4060 Ti, 4070 Ti, 4080 and 4090 are 8.9, which is the highest score for a gaming graphics card.
- keriati1 2y agoWhat model size is used here? How much memory does the GPU have?
- thawab 2y agohe is using the 3b one, since it's the default when downloading it from ollama: https://ollama.com/library/llama3.2 https://ollama.com/library/llama3.2
- taosx 2y agoFor the people who self-host LLMs at home: what use cases do you have? Personally, I have some notes and bookmarks that I'd like to scrape, then have an LLM summarize, generate hierarchical tags, and store in a database. For the notes part at least, I wouldn't want to give them to another provider; even for the bookmarks, I wouldn't be comfortable passing my reading profile to anyone.
- segalord 2y agoI use it exclusively for users on my personal website to chat with my data. I've given the setup tools to have read access my files and data
- netdevnet 2y agoIs this not something that you can with non-hosted LLMs like ChatGPT? If you expose your data, it should be able to access it iirc
- worldsayshi 2y agoYou can absolutely do that but then you pay by the token instead of a big upfront hardware cost. It feels different I suppose. Sunk cost and all that.
- xyc 2y agollama3.2 1b & 3b is really useful for quick tasks like creating some quick scripts from some text, then pasting them to execute as it's super fast & replaces a lot of temporary automation needs. If you don't feel like invest time into automation, sometimes you can just feed into an LLM. This is one of the reason why recently I added floating chat to https://recurse.chat/ https://recurse.chat/ to quickly access local LLM. Here's a demo: https://x.com/recursechat/status/1846309980091330815 https://x.com/recursechat/status/1846309980091330815
- taosx 2y agoLooks very nice, saved it for later. Last week, I worked on implementing always-on speech-to-text functionality for automating tasks. I've made significant progress, achieving decent accuracy, but I imposed some self-imposed constraints to implement certain parts from scratch to deliver a single binary deployable solution, which means I still have work to do (audio processing is new territory for me). However, I'm optimistic about its potential. That being said, I think the more straightforward approach would be to utilize an existing library like https://github.com/collabora/WhisperLive/ https://github.com/collabora/WhisperLive/ within a Docker container. This way, you can call it via WebSocket and integrate it with my LLM, which could also serve as a nice feature in your product.
- _blk 2y agoWhy disable LVM for a smoother reboot experience? For encryption I get it since you need a key to mount, but all my setups have LVM or ZFS and I'd say my reboots are smooth enough.
- satvikpendem 2y agoI love Coolify, used to use v3, anyone know how their v4 is going? I thought it was still a beta release from what I saw on GitHub.
- whitefables 2y agoI'm using v4 beta in the blog post. Didn't try v3 so there's no point of comparison but I'm loving it so far! It was so easy to get other non-AI stuffs running!
- j12a 2y agoCoolify is quite nice, have been running some things with the v4 beta. It reminds a bit of making web sites with a page builder. Easy to install and click around to get something running without thinking too much about it fairly quickly. Problems are quite similar also, training wheels getting stuck in the woods more easily, hehe.
- raybb 2y agoV4 beta is working well for me. Also the new core dev Coolify hired mentioned in a Tweet this week that they're fixing up lots of bugs to get ready for V4 stable.
- netdevnet 2y agoAm I right thinking that a self-hosted llama wouldn't have the kind restrictions ChatGPT has since it has no initial system prompt?
- Kudos 2y agoMany protections are baked into the models themselves.
- dtquad 2y agoAll the self-hosted LLM and text-to-image models come with some restrictions trained into them [1]. However there are plenty of people who have made uncensored "forks" of these models where the restrictions have been "trained away" (mostly by fine-tuning). You can find plenty of uncensored LLM models here: https://ollama.com/library https://ollama.com/library [1]: I personally suspect that many LLMs are still trained on WebText, derivatives of WebText, or using synthetic data generated by LLMs trained on WebText. This might be why they feel so "censored": >WebText was generated by scraping only pages linked to by Reddit posts that had received at least three upvotes prior to December 2017. The corpus was subsequently cleaned The implications of so much AI trained on content upvoted by 2015-2017 redditors is not talked about enough.
- nubinetwork 2y ago> All the self-hosted [...] text-to-image models come with some restrictions trained into them https://github.com/huggingface/diffusers/issues/3422 https://github.com/huggingface/diffusers/issues/3422
- thrdbndndn 2y agoMy to-go test for uncensoring is to ask the LLM to write erotic novel. But I haven't yet find any "uncensored" ones (on ollama) that works. Did I miss something? (On the contrary: when ChatGPT first came out, it was trivial to jailbreak it to make it write erotica.)
- gjs278 2y ago
- ragebol 2y agoProbably saves a bit on the gas bill for heating too
- szundi 2y agoIf only we had heat-pump computers
- ragebol 2y agoI'd gladly run whatever model you want at home, rent it out so you can pay for both heating, the GPU and the power consumed :-)
- CraigJPerry 2y agoI don’t know, it’s kind of amazing how good the lighter weight self hosted models are now. Given a 16gb system with cpu inference only, I’m hosting gemma2 9b at q8 for llm tasks and SDXL turbo for image work and besides the memory usage creeping up for a second or so while i invoke a prompt, they’re basically undetectable in the background.
- rglullis 2y agoSnark aside, even in Germany (where electricity is very expensive) it is more economical to self host than to pay for a subscription to any of the commercial providers.
- bambax 2y ago> I decided to explore self-hosting some of my non-critical applications Self-hosting static or almost-static websites is now really easy with a Cloudflare front. I just closed my account on SmugMug and published my images locally using my NAS; this costs no extra money (is basically free) since the photos were already on the NAS, and the NAS is already powered on 24-7. The NAS I use is an Asustor so it's not really Linux and you can't install what you want on it, but it has Apache, Python and PHP with Sqlite extension, which is more than enough for basic websites. Cloudflare free is like magic. Response times are near instantaneous and setup is minimal. You don't even have to configure an SSL certificate locally, it's all handled for you and works for wildcard subdomains. And of course if one puts a real server behind it, like in the post, anything's possible.
- Reubend 2y agoIs the NAS exposed to the whole internet? Or did you find a clever way to get CloudFlare in front of it despite it just being local?
- cheema33 2y agoYou can use CloudFlare Tunnel (https://www.cloudflare.com/products/tunnel/ https://www.cloudflare.com/products/tunnel/) to connect a system to your cloudflare gateway, without exposing it to the Internet.
- rmbyrro 2y agoOr Tailscale, which is pretty cool piece of tech.
- telgareith 2y agoTailscale is wireguard with advertising, a convenient UI, and a STUN/TURN server.
- calgoo 2y ago
- eloycoto 2y agoI have something like this, and I'm super happy with anythingLLM, which allows me to add a custom board with my workspaces, RAG, etc.. I love it!
- ossusermivami 2y agoai generated blog post (or reworded, whatever) are kinda getting very irritating, like playing chess against the computer, it feel soulless
- hmcamp 2y agoHow can you tell this post was ai generated? I’m curious.
- cranberryturkey 2y agoHow is coolify different than ollama? is it better? worse? I like ollama because I can pull models and it exposes a rest api to me. which is great for development
- deleted 2y ago[deleted]
- grahamj 2y agoMight want to skim the article
- cranberryturkey 2y agoi did. just realized its a totally different tool for deploying apps.
- grahamj 2y agofwiw that was my reaction to the title too :D
- vincentclee 2y agoInstead of `watch -n 0.5 nvidia-smi` to track GPU usage. One can use `nvtop` https://github.com/Syllo/nvtop https://github.com/Syllo/nvtop
- sorenjan 2y agoCan you use a selfhosted LLM that fits in 12 GB VRAM as a reasonable substitute for copilot in VSCode? And if so, can you give it documentation and other code repositories to make it better at a particular language and platform?
- 0xedd 2y agoTechnically, yes, but will yield poor results. We did it internally at big corp n+1 and it, frankly, blows. Other than menial tasks, it's good for nothing but a scout badge.
- tbrownaw 2y agoIs that really that much worse than full copilot, though? When we tried it this past spring, it was really cool but not quite useful enough to actually stick with.
- deleted 2y ago[deleted]