5 ms·
One key idea here is to use a very large number of parameters (model weights), but only use some subset of the parameters on each example. The parameters are di
by gradys 6y ago
One key idea here is to use a very large number of parameters (model weights), but only use some subset of the parameters on each example. The parameters are divided up into blocks called "experts", and then some subset of experts are used on any given input. Which subset is used is chosen by the model itself in a data-dependent manner. This can be thought of as letting the model specialize different experts to handle different situations.
The advantage, as they show, is that the model can train to a given level of performance much faster with a fixed amount of computing power compared to an architecture that uses all parameters on every step. This might be because it allows you to have a very large number of parameters that can store a lot more specialized information without incurring as much of a computational cost. Of course the downside is that you end up with a very large model that literally won't fit in a lot of environments.
- stingraycharles 6y agoAnother naive question: why is this better than creating a separate, smaller model for each expert?
- p1esk 6y agoYes, you can think of this as a collection of small models, with another model choosing which smaller model to use for each input.
- gwenzek 6y agoNot really, the dispatching between experts happen for every word and every layer, so you can't easily isolate distinct models.
- probably_wrong 6y agoThe common argument I've heard: because then you would have to decide how many experts models are required, train and evaluate them separately, and overall make your architecture dependent on this choice. If your expert is wrong and miscalculates how many models are required then your entire architecture is also likely to be wrong (humans, am I right?). Researchers at Google's scale prefer a single model where you throw all your data in a single bin and get perfect performance out, no tweaking and no pesky humans required.
- stingraycharles 6y agoBut this is something you could just use a hp search for, right, to determine the amount of models? Or are hp searches generally not used anymore at that scale?
- thomasahle 6y agoHow would you know which inputs to use for training each expert?