4 ms·
Local LLMs are the future, and one of the reasons I think the data centre furore is just going to end in a market crash. LLMs that can reliably be used for con
by PaulRobinson 8d ago
Local LLMs are the future, and one of the reasons I think the data centre furore is just going to end in a market crash.
LLMs that can reliably be used for control problems are the future, and I think classic/deep RL has generally been overlooked for years for a whole host of problems by wider industry because it felt inaccessible. The first thing I thought of when I saw Jev (and then Laya), was "this might move the needle in a really, really interesting way".
Local LLMs that can reliably be used for control problems smash through a lot of barriers I'm interested in, and this intrigues me a lot. Guess I'm about to become a big Laya fan if it can run on this kind of hardware to this performance.
- frag 8d agothat's not a local LLM. If it's local, it doesn't matter in this case. Laya is a System 1 "AI", namely works like a classifier, given a state and questions, it shoots probabilities for each. I publish an episode tomorrow about Laya and Typesafe AI on https://www.youtube.com/@DataScienceatHome https://www.youtube.com/@DataScienceatHome Stay tuned ;)
- putna 8d agocool, will check it
- EagnaIonat 7d agoI’ve found they start to fail the more classifications you have, long before your typical ML classifier. To me it’s like a solution looking for a problem that is already solved.
- bigyabai 8d agoIt won't. Laya is a finetuned version of Google's BeRT model, which is almost 10 years old right now. If BeRT had any potential to disrupt the datacenter buildout, it already would have.
- viraptor 8d agoModernbert is from 2024. It's also trained from scratch, not a fine tune.
- ipsi 8d agoThe future for whom? The general public? Not a chance, no way, not unless it's able to run on a phone (anywhere from 20-40% of internet users, world-wide, are phone-only). For companies? I think that's a lot more plausible, as that's mostly just a question of money - is it cheaper to run and administrate our own models, or outsource that? For technically inclined users? I think that's unlikely unless they're able to operate on relatively cheap hardware while still being just as good as the hosted models. And by that I don't mean "a mac studio," that's far more money than I think is reasonable. A single RTX 5080, maybe, once memory prices start to drop.
- Izmaki 8d agoCompare the games your average high-end smartphone can run to the AAA titles of the 2010's. It's not a matter of "unless it is able to" but "when it is able to".
- bigyabai 8d agoThat's going from 150w 720p gaming to ~15w 720p gaming in ~10 years. Let's say an inference cluster draws 1500w to deliver a small-ish 500b model at reasonable speeds/quantization. Extrapolating from your gaming example, it will take smartphones only... *checks clipboard* ...100 years to achieve datacenter-level performance at the pace of 2010's improvements.
- mynegation 8d agoUh, my clipboard says 20 years
- bigyabai 7d agoMea culpa, but it's still a decent while.
- itemize123 7d agono way. one is functionally (slightly overhyped) magic. another is better tool.