4 ms·
I don't understand LLMs enough to know if this is a silly question or not. Is it possible to build domain specific smaller models and merge/combine them at que
by devsda 2y ago
I don't understand LLMs enough to know if this is a silly question or not.
Is it possible to build domain specific smaller models and merge/combine them at query/run time to give better response or performance instead of one large all knowing model that learns everything ?
- RossBencina 2y agoI think that's the intuition behind MoE (Mixture of Experts). Train separate subnets for different tasks, train a router that selects which subnets to activate at inference time. Mixtral is a current open model which I believe implements this.
- ljlolel 2y agoNo. MoE tends to change expert every other word. There’s a bit of pattern (like a lot of punctuation to one expert) but it’s not clear what. Nobody understands how or why the router chooses the expert. It’s so early.
- j16sdiz 2y ago> Nobody understands how or why the router chooses the expert. It’s so early. Nobody understand how LLM works either. Is LLM as "early" as MoE ?
- xvector 2y agoLLMs are really well understood, what do you mean? You can see the precise activations and token probabilities for every next token. You can abliterate the network however you'd like to suppress or excite concepts of your choosing.
- ben_w 2y agoThere's various layers of understanding. If you will excuse analogy and anthropomorphism, the human analogy of what we do and don't understand about LLMs is, I think, that we understand quantum mechanics, cell chemistry, and overall connectivity (perceptrons, activation functions, and architecture) and group psychology (general dynamics of the output), but not specifically how some belief is stored (in both humans and LLMs).
- menaerus 2y agoMathematically speaking LLMs have very precise formulation and can be seen as F(context, X0, X1, ..., XP) = next_token. What science behind the LLMs is still lacking is how all these parameters are correlated one to each other and why one set of values is giving a better prediction than the other set of values. Right now, we arrive to these values through experimental approach, that is, through trainings.
- currymj 2y agoi think the younger generation who came up post deep learning, has a very very low bar for “understanding” because they never knew a world where SotA models worked in a way that made sense.
- htrp 2y ago> MoE tends to change expert every other word Any citation on this one?
- crystal_revenge 2y agoIt's covered in the original Mistral "Mixtral of Experts" paper [0]. 0. https://arxiv.org/abs/2401.04088 https://arxiv.org/abs/2401.04088
- Ey7NFZ3P0nzAe 2y agoI believe it's actually a per token routing, not a "every few words"
- regularfry 2y agoIt's mechanically capable of per-token routing, but the routing tends to be stable across more than one token. It's weird.
- qeternity 2y agoIt's got nothing to do with words, and many MoEs route to multiple experts per token (the well known Mixtral variants for example activates 2 experts per token).
- regularfry 2y agoWeirdly it does have to do with words, but not intentionally. Mechanically the routing is per-token, but the routing is frequently stable across a word as an emergent property. At least, that's how I read the mixtral paper.
- ljlolel 2y agoYep. Also note that I’m ELI5 so saying word is fine.
- deleted 2y ago[deleted]
- qeternity 2y agoThis is not how MoEs work at all. They are all trained together, often you have multiple experts activated for a single token. They are not domain specific in any way that is understandable by humans.
- zwaps 2y agoThis is called speculative decoding
- qeternity 2y agoNo, speculative decoding is when you use a smaller draft model to propose tokens and then use the larger target model to verify the proposals. It has got nothing to do with domain specialization.
- benob 2y agoYou might want to look into "task arithmetic" which aims at combining task-specific models post-training. For example: https://proceedings.neurips.cc/paper_files/paper/2023/file/d28077e5ff52034cd35b4aa15320caea-Paper-Conference.pdf https://proceedings.neurips.cc/paper_files/paper/2023/file/d...
- elcomet 2y agoIt's possible, the question is how to choose which submodel will be used for a given query. You can use a specific LLM, or a general larger LLM to do this routing. Also, some work suggest using smaller llms to generate multiple responses and use a stronger and larger model to rank the responses (which is much more efficient than generating them)
- dhash 2y agoTaking a further step back from LLM’s, this is called portfolio / ensemble techniques in the literature. A common practice in more formal domains is to have a portfolio of solvers and race them, allowing for the first (provably correct) solver to “win” In less formal domains, adding/removing nodes/trees in an online manner is part of the deployment process for random forests.