4 ms·
What would that breakthrough be?
by i_love_retros 5mo ago
What would that breakthrough be?
- Waterluvian 5mo agoMagic math and computer science that allows us to get the same quality response for a fraction of the GPU.
- deleted 5mo ago[deleted]
- deleted 5mo ago[deleted]
- toufka 5mo agoI mean, the most cutting edge of iPhones, iPads and MacBook Pros _today_ are quite capable of running in realtime today’s high-end local LLMs. If you project out that hardware just a couple of years, and the trained models out a couple of years, you end up in a place where it makes so much more sense to run them locally, for all sorts of latency, privacy, efficacy, and domain-specific reasons. Not all that different from the old terminal & mainframe->pc shifts. Finally - hardware has seemingly gotten out ahead of software that most folks use - watching YouTube, listening to music, playing a game or two. There was a time when playing an mp3 or watching a 4k video really taxed all but the nicest systems. Hardware fixed that problem, like it very well could this one.
- sofixa 5mo ago> I mean, the most cutting edge of iPhones, iPads and MacBook Pros _today_ are quite capable of running in realtime today’s high-end local LLMs Definitely not the high end local LLMs. The small ones, yes, absolutely. > If you project out that hardware just a couple of years One of the biggest bottlenecks for LLMs is memory capacity and bandwidth. With the current glut for memory, it's unlikely we'll see lots of advancements in terms of average memory available or its bandwidth on regular (not super high end devices) in the coming years. Alternatively, it's possible we get dedicated SMLs for e.g. phone specific use cases, that are optimised and run well.
- intothemild 5mo agoThat's already happening. Qwen3.6 and Gemma4. Basically small and medium models that are crazy well trained for their sizes. Then we have a lot of specular decoding stuff like MTP and others coming to speed up responses, and finally better quantisation to use less memory. Local LLM is the future, and the larger labs know that the open models will eat their lunch once people realise that the gap is only a few months. If we were good with LLMs a couple months ago, we're good with the open models now.
- krupan 5mo agoAnd how were those models developed and trained?
- lelanthran 5mo ago> And how were those models developed and trained? That's irrelevant to my decision to use local or not.
- deleted 5mo ago[deleted]
- krupan 5mo agoThat's not what this thread is about? We're saying some new breakthrough is needed, someone said it already has happened, and I'm asking if it really has. Has it? I don't think so, those models are not in some way fundamentally different than other LLMs
- lelanthran 5mo ago> We're saying some new breakthrough is needed, someone said it already has happened, and I'm asking if it really has. I didn't read "and how were those models trained" as "Are we there yet?"
- intothemild 5mo ago
- YZF 5mo agoThe current LLMs are also "magic" so anything is possible. AFAIK there is no proof that the current architecture is optimal. And we have our brains as a pretty powerful local thinking machine as a counter-example to the idea that thinking has to happen in data centers.
- _heimdall 5mo agoI want to ask what makes them magic, but even those building LLMs don't really know what happens when they run inference... I have to assume current architectures aren't optimal though, the idea that we stumbled into the one and only optimal solution seems almost impossible.
- _heimdall 5mo agoI'd assume its a totally different architecture that isn't based on storing a compressed dataset of all digital human text.