3 ms·
I dunno why Apple fans do this thing where they pretend that their specific workflow is the standard for how things work, and because they can do it so well on
by ActorNightly 1mo ago
I dunno why Apple fans do this thing where they pretend that their specific workflow is the standard for how things work, and because they can do it so well on their Macs, that means Macs are the best.
To be specific, 5 models at 57 gb means you are using crap quantized models, which suck for any real agentic work. I mean, sure they give you some inference, but compared to the full parameter models like Qwen3.8 and Gemma4 that can run full agentic loops, you may as well just use cloud inference for the price.
You of course could "run" those larger models, but we both know that the tok/sec is dogshit on Macs for those.
And 57 gb is split across 3 cards quite easily, which will all be cheaper than your comparable Mac and way faster.
You really need to educated yourself on how running local models works and what the models like Gemma 4 are capable of, so you don't continue to waste money on Macs.
- EagnaIonat 1mo ago> I dunno why Apple fans You mean someone who uses a Mac? I find the term "Apple fan" is used as a way to attack the person. > that can run full agentic loops, you may as well just use cloud inference for the price. You can run full agentic loops fine with quantised models. In fact that is likely what you are doing with a 32GB PC Graphics card. Or what model are you using? You use the right model for the right job. For example granite4.2 is optimised for agentic work and only needs 5GB of memory. Gemma4 MLX runs fine with the larger model needing 19GB. Prior to that I had OpenClaw (in Parallels VM) create an application with local models that worked the exact same way as created by Claude. It was more an experiment in how OpenClaw works, hence the VM. > And 57 gb is split across 3 cards quite easily, I assume you are talking about a good graphics card. A good 32GB will run you $2K a card, so that $6K to beat out a laptop of similar price, and where the difference doesn't matter. [edit] Anyway my main point is Macs work fine for local models. I have a 6 year old machine that proves that.
- ActorNightly 1mo agoRunning agentic loops doesnt mean just being able to execute them. When your llm is so slow that you are faster writing the code yourself with free gemini that comes with google account, local llm is no longer worth it. Tok/sec is everything. And 3090 is $1500 used, and 24gb gb of ram. 3x is $4500 for 72gb of ram.
- EagnaIonat 1mo ago> When your llm is so slow that you are faster writing the code yourself I have never had that issue, and LLMs are used for more than just code generation. > And 3090 is $1500 used Second hand Macs are also cheaper. Not sure what your point is, but you just seem to arguing because of your hatred of Macs.