3 ms·
Does MoE help with multimodality? Can it in general enable reasoning in imagery (technical drawings, diagrams, schematics) rather than text-based?
by ddevnyc 2mo ago
Does MoE help with multimodality? Can it in general enable reasoning in imagery (technical drawings, diagrams, schematics) rather than text-based?
- dannyw 2mo agoMoE has nothing to do with multimodality. MoE is a concept proposed in 1991, before the deep learning era (which is before what I call the transformers era). You can think of it like sharing. Contrary to popular belief; 'experts' in MoE LLMs do not specialize. There's no expert trained to be good at maths, or python, or writing, or whatever. It's an inference optimization. As for reasoning in non-text modalities, you might find this paper interesting :) https://huggingface.co/papers/2502.05171 https://huggingface.co/papers/2502.05171
- dummydummy1234 2mo agoWait I thought the router ends up specializing the experts? Like there is no explicit goal aside from each 'expert' getting roughly equal weight? And it happens that when you train the router you do end up passing certain classes of problem to each expert - just as a training result nothing as clean as a python expert. But math vs creative writing will tend to rely on different experts over the majority of the inference? I do not know what I am talking about, this is my limited understanding...
- juliendorra 2mo agoYou are right that at training the main goal is balancing between the sections to avoid certain paths becoming the only path. In the end the inference will be routed token by token to a mixture of say 3 or 4 sections of the model. The combination can change at each turn. It’s really a statistical optimization. As for many things in neural networks, the original intuition coming from anthropomorphism once implemented becomes something very non human!
- ahepp 2mo agoMy understanding of it is also pretty surface level, but I was not under the impression that it develops "expertise" in a particular subject matter, at least not in a way that's easy to harness. From what I've read, it develops expertise at the token level. Because the natural continuation of this is to say like, "Ok I want to load the bird detection expert and the navigation expert but leave the medieval European history expert behind", and my understanding is that this is not really how it works. At least at the moment.
- ahepp 2mo ago> You can think of it like sharing. was this meant to read "sharding"?
- xg15 2mo ago> There's no expert trained to be good at maths, or python, or writing, or whatever. It's an inference optimization. Huh, always thought one would sort of require the other. For a good MoE model, wouldn't I want to minimize the "churn" between experts, i.e. the amount of time one expert model has to be swapped in for another expert model? That would be naturally the way if experts correspond to semantic categories. E.g. suppose I have a model that can answer questions in 100s of languages. Then while any of those languages might be requested by some caller, it's highly unlikely a caller will request all languages in the same session - realistically, there might be one or two languages in a session and those will then span the entire session. There will also be languages that are requested very often and others that are extremely rare. (Let's say my model also supports Klingon and Sindarin. Those are important for marketing reasons and because I genuinely like to make the occasional nerd happy - but practically, I get maybe a handful of requests for those every few months. So it would make sense to centralize the knowledge for those languages in some specific part of the model, so I can keep that part out of VRAM - and probably RAM as well - during the 99% of time where it's not needed) So wouldn't it make sense to make the expert models language specific here? Then you could take advantage of the fact that a language rarely changes inside a session and keep that expert in VRAM for the entire session. You could also avoid dragging parameters along with you for languages that are practically never used.