3 ms·
The common argument I've heard: because then you would have to decide how many experts models are required, train and evaluate them separately, and overall make
by probably_wrong 6y ago
The common argument I've heard: because then you would have to decide how many experts models are required, train and evaluate them separately, and overall make your architecture dependent on this choice. If your expert is wrong and miscalculates how many models are required then your entire architecture is also likely to be wrong (humans, am I right?).
Researchers at Google's scale prefer a single model where you throw all your data in a single bin and get perfect performance out, no tweaking and no pesky humans required.
- stingraycharles 6y agoBut this is something you could just use a hp search for, right, to determine the amount of models? Or are hp searches generally not used anymore at that scale?