4 ms·
Isn't evaluating against different effective "experts" within the model effectively what MoE [1] does? > Mixture of experts (MoE) is a machine learning techniq
by byteknight 2y ago
Isn't evaluating against different effective "experts" within the model effectively what MoE [1] does?
> Mixture of experts (MoE) is a machine learning technique where multiple expert networks (learners) are used to divide a problem space into homogeneous regions.[1] It differs from ensemble techniques in that for MoE, typically only one or a few expert models are run for each input, whereas in ensemble techniques, all models are run on every input.
[1] https://en.wikipedia.org/wiki/Mixture_of_experts https://en.wikipedia.org/wiki/Mixture_of_experts
- HarHarVeryFunny 2y agoNo - MoE is just a way to add more parameters to a model without increasing the cost (number of FLOPs) of running it. The way MoE does this is by having multiple alternate parallel paths through some parts of the model, together with a routing component that decides which path (one only) to send each token through. These paths are the "experts", but the name doesn't really correspond to any intuitive notion of expert. So, rather than having 1 path with N parameters, you have M paths (experts) each with N parameters, but each token only goes through one of them, so number of FLOPs is unchanged. With tree search, whether for a game like Chess or potentially LLMs, you are growing a "tree" of all possible alternate branching continuations of the game (sentence), and keeping the number of these branches under control by evaluating each branch (= sequence of moves) to see if it is worth continuing to grow, and if not discarding it ("pruning" it off the tree). With Chess, pruning is easy since you just need to look at the board position at the tip of the branch and decide if it's a good enough position to continue playing from (extending the branch). With an LLM each branch would represent an alternate continuation of the input prompt, and to decide whether to prune it or not you'd have to pass the input + branch to another LLM and have it decide if it looked promising or not (easier said than done!). So, MoE is just a way to cap the cost of running a model, while tree search is a way to explore alternate continuations and decide which ones to discard, and which ones to explore (evaluate) further.
- PartiallyTyped 2y agoHow does MoE choose an expert? From the outside and if we squint a bit; this looks a lot like an inverted attention mechanism where the token attends to the experts.
- telotortium 2y agoUsually there’s a small neural network that makes the choice for each token in an LLM.
- HarHarVeryFunny 2y agoI don't know the details, but there are a variety of routing mechanisms that have been tried. One goal is to load balance tokens among the experts so that each expert's parameters are equally utilized, which it seems must sometimes conflict with wanting to route to an expert based on the token itself.
- magicalhippo 2y agoFrom what I can gather it depends, but could be a simple Softmax-based layer[1] or just argmax[2]. There was also a recent post[3] about a model where they used a cross-attention layer to let the expert selection be more context aware. [1]: https://arxiv.org/abs/1701.06538 https://arxiv.org/abs/1701.06538 [2]: https://arxiv.org/abs/2208.02813 https://arxiv.org/abs/2208.02813 [3]: https://news.ycombinator.com/item?id=40675577 https://news.ycombinator.com/item?id=40675577