4 ms·
Apart from learning, I cant see the point of spending that kind of money and energy to get such awfull token generation speed. Assuming memory bandwith is the b
by mickael-kerjean 1mo ago
Apart from learning, I cant see the point of spending that kind of money and energy to get such awfull token generation speed. Assuming memory bandwith is the bottleneck, is it just a matter of time until we start to see hbm4 based chip able to run qwen3.8 for normal people running at more than 500 token / seconds? Is memory speed the only technological bottlenecks that prevent us from having fast local model?
- wolvoleo 1mo agoFor me, a HUGE benefit to running local models is that my data stays mine, on my computers only. Even if Google, OpenAI, Anthropic whatever promise not to use it, how sure can I be of that? Data is gold anyway. And we're in the middle of a massive gold rush. And they've already been caught scraping sites they had no business to, and pirating books. Clearly their promises and the law mean nothing to them. It's just something you pay off in a settlement if you get caught, a cost of doing business. With local models besides the speed you also lose a lot of inference quality but it helps to mitigate that. For example making sure your RAG inputs are properly prepared and categorised so the model can find them easily without having to wade through a bunch of misdirected crap. For example what I do with my bookmarks and chats (the latter are recorded per day), is before I enter them into a RAG corpus I run a small LLM over it to summarise what's being discussed or what the webpage is about. That really helped retrieval quality, and doing this is a batch task that can run asynchronously so speed is not very relevant. This way I get a lot closer to SOTA-model retrieval quality (like with Office Copilot 365 looking for conversations in Teams). PS: I wouldn't be surprised if Microsoft runs something similar on their end :) But yes the Mac Studio is outrageously priced especially for running such a small model as qwen 27b. For that price you can use something much much cheaper. It only shines for models that are much bigger, because there simply is not much hardware that can address fast 512GB banks.
- rbanffy 1mo ago> my data stays mine, on my computers only. Which is a regulation constraint on many professions, BTW. Some people simply can't give some data to say, ChatGPT, without comprehensive guarantees written in Sam Altman's blood.
- voakbasda 1mo agoGuarantees that are completely worthless, because these companies will simply pay the fine for breaking their word. There need to be second order consequences to those who delegate their responsibility to companies that behave like this. They know the contract is worthless, but then proceed to use it as defense for their own gross negligence.
- rbanffy 1mo ago> Guarantees that are completely worthless They ensure you are not liable when they misuse the information you entrusted them because you took “adequate precautions”.
- wolvoleo 1mo agoYes that's cool for a company but not when I'm the end user, because it's my data that's being misused. It's my problem and having someone to blame doesn't actually solve the problem.
- rbanffy 1mo agoMy original point was that, for many professionals, there is a strong incentive to go local with their AI models the same way you wouldn’t be able to use cloud storage without the assurance the data would be physically stored within a specific region.
- someguydave 1mo agoIs there an “easy button” software for macos for making the RAG corpus you mention? Ideally installed from homebrew?
- dannyw 1mo agoI just used Qwen3.8 & Deepseek v4 flash to make it for me.
- wolvoleo 1mo agoSimilar here, except I used Opus 4.8. I don't mind using cloud models to make code as long as I don't put actual private data into the model. And I can't run qwen 27b reliably at the moment, I need a better gpu. It was pretty easy mode like this tbh, though you do need to know what you're doing. I use it with llama-server and openwebui. I guess you can get it a lot more click to go with something like LM Studio though but I want to use it on a server and call upon its services from multiple sources.
- wolvoleo 1mo agoI'm not sure. I don't use Mac anymore. It used to be my daily driver but they pushed me away a few years ago with the constant iOSification. I've heard good thing about LM Studio, that's about it. https://lmstudio.ai/ https://lmstudio.ai/ I just run a server with Linux (previously multiple servers but I found a way to add multiple GPUs to a single one).
- wolvoleo 1mo agoSorry I made a typo. When I said: > For that price you can use something much much cheaper. I meant "For that model size you can use something much cheaper".
- bicepjai 1mo ago> For me, a HUGE benefit to running local models is that my data stays mine, on my computers only. This is a huge motivation for me to build one myself
- dannyw 1mo agoThe M3 Ultra was a lot cheaper for most of its lifetime, and really the main reason for buying it is if you want a lot of unified RAM, to run bigger models than 27B models. The new M5 Ultra should deliver ~50% faster token generation (1.2TB/s mem bandwidth); and extrapolating from my M5 Max (since the M5 Ultra is literally just 2x Maxes), probably ~3x faster PP. But I don't think it's fair to look at this only from monetary ROI vs API. With local models, you get privacy and ownership. I do not trust _any_ API provider with my most personal information; such as for example, all my messages, emails, daily journals spanning a decade+, all my photos and videos, etc. So it unlocks new use cases that I simply don't feel comfortable with via API. And a personal assistant with ALL my context and data, locally, has been incredibly useful for me :) Zero outages either, zero "overloaded", etc. Nearly-zero refusals too (I don't run abliterated models; thinking prefill has worked for anything I've wanted to do)
- Larrikin 1mo agoA slow AI tasks that can process my self hosted personal journal, my medical history in fasten, my diet and exercise in Mealie and Sparky Fitness to offer insights once a day or even once a week is better than nothing. Because I would never upload that data to any of the AI companies.
- oceanplexian 1mo agoI have 2x 3090s and I get 260 tokens/sec peak (110 avg) with Dflash2 and about 1500t/s prefill with a Q4 quant of Qwen 27b. I didn't buy an overpriced Apple product and it performs much better. It's extremely reliable for Agentic coding and I can run 2-3 simultaneous agents with a full ~260k context window. I realize the price of NVIDIA has gone up but there are plenty of GPU options from others like AMD and Intel with reasonable performance.