6 ms·
I strongly believe this premise in the article is correct - we will see a lot of tiny, hyper specialized models for individual tasks, and perhaps that will conv
by N_Lens 3mo ago
I strongly believe this premise in the article is correct - we will see a lot of tiny, hyper specialized models for individual tasks, and perhaps that will converge with an orchestration layer for a generalized intelligence that controls these specialized tiny models, that will be quite capable.
I don't foresee AGI arising out training bigger LLMs (Though investors won't realise that for a while yet).
It's actually how organic brains work - specialized tasks are offloaded to local cortical columns. The overall coordination between these sub-brains creates emergent skills/abilities.
- looofooo0 3mo agoWhat about recent models providing correct proofs to open math problems?
- 16bitvoid 3mo agoWhat about it?
- TJSomething 3mo agoI haven't tried it, but I saw Leanstral, an LLM specialized in writing Lean proofs, posted on HN recently and it claims to outperform some larger general purpose models. It didn't beat Claude Opus, but it seems to do decently at one tenth the cost. It's plausible that further research could yield other models that are smaller and more effective at limited tasks, reversing the trend of ever growing models.
- stingraycharles 3mo ago> It's actually how organic brains work - specialized tasks are offloaded to local cortical columns. How are small isolated language models more similar to that than MoE in LLMs?
- simianwords 3mo agoRight MoE is a tradeoff between efficiency and intelligence.
- roadside_picnic 3mo agoMoEs don't route like most people imagine. They aren't learning topic based experts despite the name The original Mixtral paper [0] (in the "Routing analysis" section) found: "surprisingly, we do not observe obvious patterns in the assignment of experts based on the topic" A quick skim of more recent analysis on MoE shows that this hasn't changed. MoE models do appear to work, but don't appear to do what the name implies, if anything they're routing based on the structure of the text and not the semantic content (and we're still not entirely sure what they're doing). 0. https://arxiv.org/pdf/2401.04088 https://arxiv.org/pdf/2401.04088
- andy99 3mo agoGeneral purpose models are always more robust and generally better than smaller narrower models. My bet is that compute will catch up and any “small” model will still be generally capable, just smaller than sota, rather than intentionally narrow. The exception would be for very well defined tasks where the data distribution never varies, but these are rare and don’t really need “AI” anyway when they do exist.
- plastic-enjoyer 3mo ago> General purpose models are always more robust and generally better than smaller narrower models. What do you mean with more robust?
- deleted 3mo ago[deleted]
- ACCount37 3mo agoLess weird unexpected failures, more innate ability to handle edge cases gracefully. Quite important when you're running high on automation and low on oversight.
- plastic-enjoyer 3mo agoThis may be speculative, but couldn't robustness emerge by having a number of specialized models, that are interconnected and feed into each other? Are there any arguments from ML that would speak against this?
- ACCount37 3mo agoThat does work. Even if you drop the "specialized" part. Ensembles of the same architecture at the same scale trained on the same data do outperform a singular model of the same line - especially on corner cases. Successes of an ensemble correlate stronger than failures do. The usual argument against is that if you have "a number of specialized models" that perform well in ensemble, you can take that ensemble, and distill it into a single larger model (dense or integrated sparse, like MoE), and get the same improvement in performance with an efficiency win. This works because having those "specialized models" duplicates a lot of the highly conserved "low level" wiring that's required for a model to function at all. As such, you end up running a small scale version of the same "backbone" computational processes many times. "Merging" those models into a larger, denser model allows for a singular strong "backbone" to be used for everything.
- simianwords 3mo agoNo this will never work. Domain specific models will never be a thing because intelligence carries over and compounds. Why didn’t OpenAI release a math specific model? Why not a literature specific one? Why do they instead have generic models of different sizes? And how did all labs converge on this? Why does Fable just not train on non cybersec and non biology data but instead have clearly costly and annoying classifiers?
- thereitgoes456 3mo agoDeepMind did release a math specific model. And OpenAI has released a coding specific model. The answer to your question is “because the market isn’t big enough”, not because it doesn’t work. Why would knowing about 2019 internet memes help you in any way at coding?
- simianwords 3mo ago> And OpenAI has released a coding specific model They did and retracted it because they found that GPT 5.5 beat codex pareto optimally. This keeps happening. > because the market isn’t big enough Huuh? market isn't big enough for AGI? The parent suggested that AGI would emerge from this process.
- InsomniacL 3mo ago> Why would knowing about 2019 internet memes help you in any way at coding? https://github.com/Brainrotlang/brainrot https://github.com/Brainrotlang/brainrot "Brainrot is a meme-inspired programming language that translates common programming keywords into internet slang and meme references."
- andy99 3mo ago> Why would knowing about 2019 internet memes help you in any way at coding? 99.99% of the knowledge an LLM has is useless for a given scenario, the hard part is knowing what the .01% that’s needed is. Knowing as much as it can means the model can handle edge cases, turns of phrase, etc. Put another way, it avoids overfitting. That’s basically the insight that’s given way to the current AI boom.
- visarga 3mo agoI think the harness and local context should supply that missing piece between general model and bespoke application. Each application has its own context and action quirks that don't generalize well. Maybe it's just 5% but that is genuinely specific. So its rightful place is in context engineering. I have a long-ass post about how this could be implemented. https://old.reddit.com/r/VisargaPersonal/comments/1um9uyv/state_space_policies_making_expert_judgment/ https://old.reddit.com/r/VisargaPersonal/comments/1um9uyv/st...
- chris_money202 3mo agoI think future is probably more similar to speculative execution (inference/decoding). A small LLM is used to speculate and a large LLM is used to confirm if needed. If the small LLM is accurate enough on N tokens it’s cheap for the large LLM to say looks good and keep moving along.
- hiyfsch 3mo ago[dead]
- piotrekno1 3mo ago[flagged]