5 ms·
The problem for me with making such an investment is that next month a better model will be released. It will either require more or less RAM than the current b
by gregwebs 2y ago
The problem for me with making such an investment is that next month a better model will be released. It will either require more or less RAM than the current best model- making it either not runnable or expensive to run on an overbuilt machine.
Using cloud infrastructure should help with this issue. It may cost much more per run but money can be saved if usage is intermittent.
How are HN users handling this?
- nickthegreek 2y agoI plan to wait for the NVIDIA Digits release and see what the token/sec is there. Ideally it will work well for at least 2-3 years then I can resell and upgrade if needed.
- diggan 2y ago> How are HN users handling this? Combine the best of both worlds. I have a local assistant (communicate via Telegram) that handles tool-calling and basic calendar/todo management (running on a RTX 3090ti), but for more complicated stuff, it can call out to more advanced models (currently using OpenAI APIs for this) granted the request itself doesn't involve personal data, then it flat out refuses, for better or worse.
- idrathernot 2y agoThere is also an overlooked “tail risk” with cloud services that can end up costing you more than a a few entire on-premise rigs if you don’t correctly configure services or forget to shut down a high end vm instance. Yeah you can implement additional scripts and services as a fail-safe, but this adds another layer of complexity that isn’t always trivial (especially for a hobbyist). I’m not saying that dumping $10k into rapidly depreciating local hardware is the more economical choice, just that people often discount the likelihood and cost of making mistakes in the cloud during their evaluations and the time investment required to ensure you have the correct safeguards in-place.
- anon373839 2y agoYes. And somehow, those cloud providers just can’t seem to work out how to build a spend limit feature for customers who’d like to prevent that. It must be a really difficult engineering problem…
- 3s 2y agoExactly! While I have llama running locally on RTX and it’s fun to tinker with, I can’t use it for my workflows and don’t want to invest 20k+ to run a decent model locally > How are HN users handling this? I’m working on a startup for end-to-end confidential AI using secure enclaves in the cloud (think of it like extending a local+private setup to the cloud with verifiable security guarantees). Live demo with DeepSeek 70B: chat.tinfoil.sh
- walterbell 2y ago> expensive to run on an overbuilt machine There's a healthy secondary market for GPUs.
- xienze 2y agoThe price goes up dramatically once you go past 12GB though, that’s the problem.
- JKCalhoun 2y agoNot on these server GPUs. I'm seeing 24GB M40 cards for $200, 24GB K80 cards for $40 on eBay.
- xienze 2y agoWell OK, I should have been more specific that, even for server GPUs on eBay: * Cheap * Fast * Decent amount of RAM Pick two. These old GPUs are as cheap as they are because they don’t perform well.
- JKCalhoun 2y agoThat's fair. From what I read though I think there is some interplay between Fast and Decent amount of RAM. Or at least there is a large falloff in performance when RAM is too small. So Cheap and Decent amount of RAM work for me.
- JKCalhoun 2y agoI think the solution is already in the article and comments here: go cheap. Even next year the author will still have, at the very least, their P40 setup running late 2024 models. I'm about to plunge in as others have to get my own homelab running the current crop of models. I think there's no time like the present.
- tempoponet 2y agoMost of these new models release several variants, typically in the 8b, 30b, and 70b range for personal use. YMMV with each, but you usually use the models that fit your hardware, and the models keep getting better even in the same parameter range. To your point about cloud models, these are really quite cheap these days, especially for inference. If you're just doing conversation or tool use, you're unlikely to spend more than the cost of a local server, and the price per token is a race to the bottom. If you're doing training or processing a ton of documents for RAG setups, you can run these in batches locally overnight and let them take as long as they need, only paying for power. Then you can use cloud services on the resulting model or RAG for quick and cheap inference.
- michaelt 2y agoAmong people who are running large models at home, I think the solution is basically to be rich. Plenty of people in tech earn enough to support a family and drive a fancy car, but choose not to. A used RTX 3090 isn't cheap, but you can afford a lot of $1000 GPUs if you don't buy that $40k car. Other options include only running the smaller LLMs; buying dated cards and praying you can get the drivers to work; or just using hosted LLMs like normal people.