4 ms·
This is the diametric opposite of the rent-vs-buy scenario that this entails. Local: You need to invest $thousands into GPU and/or very-high-end CPU+Memory har
by jiggawatts 2mo ago
This is the diametric opposite of the rent-vs-buy scenario that this entails.
Local: You need to invest $thousands into GPU and/or very-high-end CPU+Memory hardware.
Vendor: You can use any existing device, even a phone or tablet. A very low-end laptop is fine.
> takes literal minutes to get started
Local: Typical scenario is hours just to download the software, the model weights, and then faffing around with CUDA and matching your GPU drivers.
Vendor: Free-tier available instantly on a web URL. Even local agents have free tiers from multiple vendors. Install is a single command and/or download and "next,next,next,finish" wizard that takes ~1 minute.
> you can just `rm -fr` it and forget the whole thing existed.
I'm still cleaning up multi-GB model weights floating around in hidden subdirectories under my user profile from months ago when I was experimenting with local models!
Meanwhile I simply... stopped using Gemini. That was the entire process: I no longer actively use it. They stopped billing me for my token usage, because it is now zero. That's... it.
You have it totally backwards.
- palata 2mo ago> I'm still cleaning up multi-GB model weights floating around in hidden subdirectories under my user profile from months ago when I was experimenting with local models! Are you trying to say that local models are hard to use because... you're having issues handling files properly? I am not sure I get the argument. I get the rest of the comment: local models require an investment upfront, and it is less convenient. It doesn't say that it is not cheaper, though.
- gpugreg 2mo ago> you're having issues handling files properly? I guess they were using ollama, which does not tell you where it puts the models it downloads.
- accrual 2mo agoFilelight / ncdu are my friends for finding random 30GB directories containing cached models.
- gpugreg 2mo agoPersonally, I prefer QDirStat. I just tried to use FileLight to compare, but the package seems to be broken on Lubuntu.
- jiggawatts 2mo agoYou two have just listed several tools needs to clean up after another tool: not a strong argument for “simpler”.
- deleted 2mo ago[deleted]
- ahartmetz 2mo agoIt took me about three hours total to set up a local model. I already have a GPU and I have fiber for the download. llama.cpp is not difficult to compile and has many backends. It can run parts of the model on different backends, like in the common case that the GPU doesn't have enough VRAM for everything. There are many step-by-step guides available.
- AbsurdCensor 2mo agoTakes even less depending on your system. LM Studio or Lemonade and you are set up in minutes and now they can even tell you what models will fit with the memory you have.
- dgellow 2mo agoAnd it would be in seconds if models weren’t that large and slow-ish to download! LM studio is such a noob friendly experience, pretty neat first experience!
- AbsurdCensor 2mo agoAt least for the most part, if you are downloading from huggingface, you should be able to saturate your connection. I know I usually can pretty easily even with a 5gig connection at home.
- dgellow 2mo agoHmm, let’s not talk about my German poor internet connection please :)
- jiggawatts 2mo agoThree hours is a lot longer than one minute.
- ahartmetz 2mo agoNo shit, but the huge ordeal you described is an exaggeration.
- leansensei 2mo agoHuh what? Qwen3.5-35B-A3B runs just fine with maximum context, on an RTX SUPER 12 GB, with offloading of some expert layers to DDR4-3200. Same story on an RTX 4060 Ti 16 GB. MTP is a serious boost to tg. Downloading the model is a simple hf command that HuggingFace's web UI even gives you. llama.cpp is trivial to use, and so is llama-swap, if you want to use other models too. If you don't know what arguments to run it with, you download ggrun and use that. Local LLMs are incredibly capable and don't need expensive hardware. A $500 GPU will do. Or even cheaper. This is all trivial.
- tsss 2mo agoRTX Super 12GB costs $700. An openrouter account costs nothing.
- freehorse 2mo agoA lot of people already have 12GB+ GPUs lying around for playing games, doing video editing, etc. I would not get a GPU or mac just to run LLMs personally, but if one wants to get such a device for other tasks too, it may make sense to eg choose a slightly higher (v)RAM variant if they want to run some bigger models. Then what you pay for the local llms is just the difference.
- Anonyneko 2mo agoFull model or a 4-bit quant? I have a 5090 and I'm not sure whether I should use a quant that fits within the VRAM or a much bigger version where I'd have to offload a lot to 64GB RAM and a beefy CPU (but still a CPU)
- NekkoDroid 2mo agoI personally run the Q6 quant on my RX 9070 XT (16GB VRAM). On r/LocalLlama there was a post recently as well, which talked about the degradation of different quants (for the 27B version)[0] [0]: https://www.reddit.com/r/LocalLLaMA/comments/1vef79c/quantization_hurts_knowledge_nonlinearly_qwen36/ https://www.reddit.com/r/LocalLLaMA/comments/1vef79c/quantiz...
- peri-cl 2mo ago> "I'm still cleaning up multi-GB model weights floating around in hidden subdirectories under my user profile from months ago when I was experimenting with local models!" I used to deal with these kinds of frustrations too. fd --unrestricted --size +1G fd --help -u, --unrestricted... Perform an unrestricted search, including ignored and hidden files. This is an alias for '--no-ignore --hidden'. -S, --size size Limit results based on the size of files using the format <+-><NUM><UNIT>
- Forgeties79 2mo ago> Local: Typical scenario is hours just to download the software, the model weights, and then faffing around with CUDA and matching your GPU drivers. Download LM studio, search models, click download, wait minutes, prompt and have fun
- jiggawatts 2mo ago“If you have the prerequisite hardware, then… know which model you want out of thousands of a variants… and your drivers are up to date, then it is fast!”
- deleted 2mo ago[deleted]
- Forgeties79 2mo ago> If you have the prerequisite hardware Literally tens of millions of people of silicon MacBooks have sold send 2020 so it’s probably safe to say hundreds of millions of people have the necessary hardware. Not even getting into smartphones. >which model you want Have you personally searched for models in LM studio? It’s actually pretty straight forward and it tells you with a very clear icon if it will all fit in your GPU or if it will offload onto ram. >drivers are up to date Are you just making things up now? I run LM studio on an M1 MBpro (albeit very small model for small tasks with tool calls) and on a Linux (fedora) PC with an AMD GPU. In both cases i downloaded LM studio, quickly found models with their search, and started messing around. I am not a coder or engineer mind you, so clearly it isn’t that difficult.