4 ms·
A new base model mac mini is $900. That is 45 month of Gemini. Gemini 4.7 Flash will give better OCR results that Qwen or GLM w/ 10GB.
by tiahura 26d ago
A new base model mac mini is $900. That is 45 month of Gemini. Gemini 4.7 Flash will give better OCR results that Qwen or GLM w/ 10GB.
- nrmitchi 26d ago... > They are things that I would not be comfortable sending a cloud provider It's also an old machine that the commenter already has; it's intellectually dishonest to compare it to the price of a brand new, 4-iteration-newer machine.
- Jeremy1026 26d agoThat doesn't help with the not wanting to send confidential information to a cloud though. No amount of cost savings can negate that.
- deleted 26d ago[deleted]
- wolvoleo 26d ago100% this. No way I'm giving Google my stuff. Neither openai or xai. Anthropic maybe but not likely. Mistral is the most likely one because they're under the EU laws but I think they focus more on commercial these days.
- homarp 26d agoexcept a) model I pick will not 'suddenly' go away b) I am sure my data stays where I want it c) my inference mac can run other things if I need to I pay for that.
- wilkystyle 26d agoThis is such a tired argument and it seems to be parroted every single time someone talks about local models on hacker news. Yes, of course the most economical path is to hand over all your data and become fully dependent on a cloud provider who is already operating as scale, hoping that they won't change/remove models, hamstring capabilities, or raise prices. If this were a thread about hosting your own email or blog or cloud photos, you'd have plenty of people out here telling you how easy it is to do it yourself instead of relying on Gmail for email or WordPress/Medium/Substack for blogging, or iCloud for cloud photos. And yet, without fail, every single thread about self hosting local models seems to have some copy/paste form of this cost-savings argument. Where is the appreciation for this cool thing GP built? Where is the appreciation for the desire to figure out how to host your own version of the incredible capabilities that were not available merely a few years ago? And why, on this site of all places, would someone advocate trading all of the knowledge and independence gained from learning how to host something like this ourselves in favor of throwing it all over the wall to Google? Come on.
- jay_kyburz 26d agoAnd everybody knows advertising is just around the corner. It will be horrible to be dependent on an AI who is also be trying to sell you various goods and services. We're going to need AI whose loyalty is to us and only us.
- mv4 26d agoIt's quite shocking to me how many experienced, tech-savvy people, who used to care about cookies and ad tracking - are now willingly sending their business strategies, highly confidential contracts, and intimate personal issues to a cloud provider because "it is only $0.0x per million tokens!".
- jmalicki 26d ago> It's quite shocking to me how many experienced, tech-savvy people, who used to care about cookies and ad tracking - are now willingly sending their business strategies, highly confidential contracts, and intimate personal issues to a cloud provider because "it is only $0.0x per million tokens!". Because there are more privacy guarantees there, depending on the provider. "But what if they violate their contract!" is some pretty tin-foil hat stuff. How is this any different than a business running their website out of the cloud, assuming you are using a provider with appropriate contractual terms? You can care about tracking and ads but still be comfortable storing your backups in the cloud, and many have been for quite awhile, even sometimes without encryption - that is totally different than e.g. Meta actively trying to track you and understand your relationship graph and your purchases etc.
- 0cf8612b2e1e 26d agoBut what if they violate their contract!" is some pretty tin-foil hat stuff. The foundation of these businesses is stealing IP in bulk.
- jmalicki 26d agoThe foundation of the businesses training AI models. So don't use them for inference.
- m4rtink 26d agoDo you believe Gemini will costs the same in 45 months or even exist, given Google track record ?
- Aurornis 26d agoThe options available across the board are getting cheaper and better all the time. There is no reason to believe that equivalent level model output will be more expensive in 12 months, let alone almost 4 years from now. Of all the good reasons to use local AI (privacy, etc), worrying about not having access to cheap models in 4 years is not one of them.
- overfeed 26d ago> There is no reason to believe that equivalent level model output will be more expensive in 12 months It's almost never a drop-in replacement, and having to check and adjust integrations and workflows with new models gets old fast. My task was perfectly solved by the old model, I don't need a newer, "better" one - especially at higher prices ("more cost-effective" my foot). Local models lets one choose a model and freeze the downstream integrations forever, without being forced on the 6/8-month upgrade treadmill by aggressively short, scarcity-driven hosted model-deprecation schedules.
- mixmastamyk 26d agoThe big providers are losing on average tens of billions a year on these services, so yes prices must go up. Even Moore’s won’t help in the medium-term due to shortages and difficulty/reluctance to vastly increase capacity.
- deleted 26d ago[deleted]
- Aurornis 26d agoYou can buy tokens on OpenRouter right now from companies who serve tokens as a business and who are not selling at a loss. The frontier labs have very high prices for inference. The prices are actually going down, not up.
- jtbaker 26d agohaving a 64GB mac mini m4 pro the last few years with some increasingly capable usefulness has kept me interested in this stuff in a way that using a paid platform wouldn't have. Similar to running K8s in a homelab, something about interacting with the hardware makes it more engaging/interesting, for me at least. In general, I'm a big believer in doing more with fewer resources, within reason, and think having local setups really helps me be mindful with what's happening under the hood with these systems and managing context efficiently to get high quality results.
- socalgal2 25d agoI guess it depends on what you're trying to do. I've run a few LLMs on my 64gig Mac but anything image or video related is ridiculously slow compared to even old NVidia on my PC.
- jtbaker 25d agoHave you been using an MTP setup? I've been having pretty good luck with the Qwen models with the built in MTP heads via https://mtplx.com https://mtplx.com at around Q4. My main driver rig is an M5 Max MBP work got for me a few weeks ago. Hitting 1000-1200 TPS prefill on Qwen 3.8-Flash-Next there.
- booty 26d agoTwo thoughts. A $20/month Gemini subscription is truly all you need, then yeah, sure.... obviously a homelab setup is a ridiculous alternative on a pure cost basis. For most people doing "real" work with LLMs 40+ hours per week, a more apt comparison would be one or multiple $200/month subscriptions. At which point the break-even point of a homelab is much sooner. However, most people running homelabs are doing it for other reasons. Independence, learning, and/or privacy issues.
- gessha 25d agoI love this quote from a homelab reddit: - Is that even worth the electricity price compared to api? - We don't ask that here
- booty 25d agoHahaha. It's an excellent question, though! I have not spun up my own homelab, but, I have researched it and the electricity costs can be mitigated to a large extent. One... the GPUs can be massively clocked down during idle state, to the point where the fans can be shut off as well. The machines themselves can be shut down and wait for a magic wake-on-LAN packet if needed. Two... even under load, the GPU cores can be significantly underclocked with very little performance loss. The GPU and VRAM/HBM clocks are independent, and the GPU is largely bottlenecked on VRAM/HBM so you can just drop the GPU speed. A commonly reported figure I saw was, basically, 350W nominal cards being underclocked to consume "only" 200W under load with ~5-10% perf loss. This all assumes you're using discrete GPUs and not AIO systems like a Mac Studio which is going to be pretty efficient just by design; they idle at 35W or so and in practice max out at a few hundred watts. I believe DGX Spark and Strix Halo are similar. I must stress that this is second-hand anecdata here, admittedly, but my understanding is that it can be pretty manageable.
- bahmboo 26d agoYes but I also get a full fledged computer in the deal. I can sell it later. I can use it for all sorts of things like games and browsing and video editing. Paying for Gemini for other tasks is also in the mix but at the end of 4 years I get...nothing.