7 ms·
I used to be excited about running models locally (LLM, stable diffusion etc) on my Mac, PC, etc. But now I have resigned to the fact that most useful AI comput
by hagope 2y ago
I used to be excited about running models locally (LLM, stable diffusion etc) on my Mac, PC, etc. But now I have resigned to the fact that most useful AI compute will mostly be in the cloud. Sure, I can run some slow Llama3 models on my home network, but why bother when it is so cheap or free to run it on a cloud service? I know Apple is pushing local AI models; however, I have serious reservations about the impact on battery performance.
- fouc 2y agohttps://x.com/karpathy/status/1814038096218083497 https://x.com/karpathy/status/1814038096218083497 LLMs will start shrinking massively in size soon, without any loss in performance.
- Cantinflas 2y agoWhy bother running models locally? Privacy, for once, or censorship resistance.
- bongodongobob 2y agoI have a 2 year old Thinkpad and I wouldn't necessarily call llama3 slow on it. It's not as fast as ChatGPT but certainly serviceable. This should only help. Not sure why your throwing your hands up because this is a step towards solving your problem.
- nhod 2y agoIs this a hunch, or do you know of some data to back up your reservations? Copilot+ PC’s, which all run models locally, have the best battery life of any portable PC devices, ever. These devices have in turn taken a page out of Apple Silicon’s playbook. Apple has the benefit of deep hardware and software integration that no one else has, and is obsessive about battery life. It is reasonable to think that battery life will not be impacted much.
- fragmede 2y agoThat doesn't seem totally reasonable. The battery life of an iphone is pretty great if you're not actually using it, but if you're using the device hard, it gets hot to the touch, along with the battery getting drained. playing resource intensive video games, maxing out the *PU won't stop and let the device sleep at all, and has a noticable hit on battery life. Where inference takes a lot of compute to perform, it's hard to imagine inference being totally free, battery-wise. It probably won't be as hard on the device as playing specific video games non-stop, but I get into phone conversations with ChatGPT as it is, so I can imagine that being a concern if you're already low on battery.
- wokwokwok 2y ago> Sure, I can run some slow Llama3 models on my home network, but why bother when it is so cheap or free to run it on a cloud service? Obvious answer: because it's not free, and it's not cheap. If you're playing with a UI library, lets say, QT... would you: a) install the community version and play with ($0) b) buy a professional license to play with (3460 €/Year) Which one do you pick? Well, the same goes. It turns out, renting a server large enough to run big (useful, > 8B) models is actually quite expensive. The per-api-call costs of real models (like GPT4) adds up very quickly once you're doing non-trivial work. If you're just messing around with the tech, why would you pay $$$$ just to piss around with it and see what you can do? Why would you not use a free version running on your old PC / mac / whatever you have lying around? > I used to be excited about running models locally That's an easy position to be one once you've already done it and figured out, yes, I really want the pro plan to build my $StartUP App. If you prefer to pay for an online service and you can afford it, absolutely go for it; but isn't this an enabler for a lot of people to play and explore the tech for $0? Isn't having more people who understand this stuff and can make meaningful (non-hype) decisions about when and where to use it good? Isn't it nice that if meta released some 400B llama 4 model, most people can play with it, not just the ones with the $7000 mac studio? ...and keep building the open source ecosystem? Isn't that great? I think it's great. Even if you don't want to play, I do.
- FeepingCreature 2y agoI just prepay $20/mo to openrouter.ai and can instantly play with every model, no further signup required.
- itake 2y agoI’m a bit confused. Your reasoning doesn’t align with the data you shared. The startup costs for just messing around at home are huge: purchasing a server and gpus, paying for electricity, time spent configuring the api. If you want to just mess around, $100 to call the world’s best api is much cheaper than spending $2-7k Mac Studio. Even at production level traffic, the ROI on uptime, devops, utilities, etc would take years to recapture the upfront and on-going costs of self-hosting. Self hosting will have higher latency and lower throughput.
- dotancohen 2y agoI have found many similarities between home AI and home astronomy. The equipment needed to get really good performance is far beyond that available to the home user, however intellectually satisfying results can be had at home as a hobby. But certainly not professional results.
- grugagag 2y agoWhen learning and experimenting it could make a difference.
- friendly_chap 2y agoWe are running smaller models with software we wrote (self plug alert: https://github.com/singulatron/singulatron https://github.com/singulatron/singulatron) with great success. There are obvious mistakes these models make (such as the one in our repo image - haha) sometimes but they can also be surprisingly versatile in areas you don't expect them to be, like coding. Our demo site uses two NVIDIA GeForce RTX 3090 and our whole team is hammering it all day. The only problem is occasionally high GPU temperature. I don't think the picture is as bleak as you paint. I actually expect Moore's Law and better AI architectures to bring on a self-hosted AI revolution in the next few years.
- dsign 2y agoFor my advanced spell-checking use-case[^1], local LLMs are, sadly, not state-of-the-art. But their $0 price-point is excellent to analyze lots of sentences and catch the most obvious issues. With some clever hacking, the most difficult cases can be handled by GPT4o and Claude. I'm glad there is a wide variety of options. [^1] Hey! If you know of spell-checking-tuned LLM models, I'm all ears (eyes).
- bruce343434 2y agoI think the floating point encoding of LLMs is inherently lossy, add to that the way tokenization works. The LLMs I've worked with "ignore" bad spelling and correctly interpret misspelled words. I'm guessing that for spelling LLMs, you'd want tokenization at the character level, rather than a byte pair encoding. You could probably train any recent LLM to be better than a human at spelling correction though, where "better" might be a vague combination of faster, cheaper, and acceptable loss of accuracy. Or maybe slightly more accurate. (A lot of people hate on LLMs for not being perfect, I don't get it. LLMs are just a tool with their own set of trade offs, no need to get rabid either for or against them. Often, things just need to be "good enough". Maybe people on this forum have higher standards than average, and can not deal with the frustration of that cognitive dissonance)
- Hihowarewetoday 2y agoI'm not sure why you have resigned? If you don't care about running it locally, just spend it online. Everything is good. But you can run it locally already. Is it cheap? No. Are we still in the beginning? yes. We are still in a phase were this is a pure luxury and just getting into it by buying a 4090, is still relativly cheap in my opinion. Why running it locally you ask? I personally think running anythingllm and similiar frameworks on your own local data is interesting. But im pretty sure in a few years you will be able to buy cheaper ml chips for running models locally fast and cheap. Btw. aat least i don't know a online service which is uncensored, has a lot of loras as choice and is cost effective. For just playing around with LLMs for sure there are plenty of services.
- PostOnce 2y agoMaybe you want to conduct experiments that the cloud API doesn't allow for. Perhaps you'd like to plug it into a toolchain that runs faster than API calls can be passed over the network? -- eventually your edge hardware is going to be able to infer a lot faster than the 50ms+ per call to the cloud. Maybe you would like to prevent the monopolists from gaining sole control of what may be the most impactful technology of the century. Or perhaps you don't want to share your data with Microsoft & Other Evils (formerly known as dont be evil). You might just like to work offline. Whole towns go offline, sometimes for days, just because of bad weather. Nevermind war and infrastructure crises. Or possibly you don't like that The Cloud model has a fervent, unshakeable belief in the propaganda of its masters. Maybe that propaganda will change one day, and not in your favor. Maybe you'd like to avoid that. There are many more reasons in the possibility space than my limited imagination allows for.
- tarruda 2y agoIt is not like strong models are at a point where you can 100% trust their output. It is always necessary to review LLM generated text before using it. I'd rather have a weaker model which I can always rely on being available than a strong model which is hosted by a third party service that can be shut down at any time.
- Aurornis 2y ago> I'd rather have a weaker model which I can always rely on being available than a strong model which is hosted by a third party service that can be shut down at any time. Every LLM project I’ve worked with has an abstraction layer for calling hosted LLMs. It’s trivial to implement another adapter to call a different LLM. It’s often does as a fallback, failover strategy. There are also services that will merge different providers into a unified API call if you don’t want to handle the complexity on the client. It’s really not a problem.
- PostOnce 2y agoSuppose you live outside of America and the supermajority of LLM companies are American. You want to ask a question about whisky distillation or abortion or anything else that's legal in your jurisdiction but not in the US, but the LLM won't answer. You've got a plethora of cloud providers, all of them aligned to a foreign country's laws and customs. If you can choose between Anthropic, OpenAI, Google, and some others... well, that's really not a choice at all. They're all in California. What good does that do an Austrian or an Australian?
- jrm4 2y agoWhat do you mean by useful here? I'm saying because I've had the exact OPPOSITE thought. The intersection of Moore's Law and the likelihood that these things won't end up as some big unified singularity brain and instead little customized use cases make me think that running at home/office will perhaps be just as appealing.
- aftbit 2y agoWhat if you want to create transcripts for 100s of hours of private recorded audio? I for one do not want to share that with the cloud providers and have it get used as training data or be subject to warrentless search under the third party doctrine. Or what if you want to run a spicy Stable Diffusion fine-tune that you'd rather not have associated with your name in case the anti-porn fascists take over? I feel like there are dozens of situations where the cost is really not the main reason to prefer a local solution.
- dws 2y ago> Sure, I can run some slow Llama3 models on my home network, but why bother when it is so cheap or free to run it on a cloud service? Running locally, you can change the system prompt. I have Gemma set up on a spare NUC, and changed the system prompt from "helpful" to "snarky" and "kind, honest" to "brutally honest". Having an LLM that will roll its eyes at you and say "whatever" is refreshing.
- diego_sandoval 2y ago> why bother when it is so cheap or free to run it on a cloud service? For the same reasons that we bother to use Open Source software instead of proprietary software.
- cess11 2y agoI don't want people I don't know snooping around in my experiments.
- deleted 2y ago[deleted]