2 ms·
For clarity, i believe these are all mixture of expert models, where each input only sparsely activates some subset subset of the full model. This is why they w
by atty 5y ago
For clarity, i believe these are all mixture of expert models, where each input only sparsely activates some subset subset of the full model. This is why they were able to make such a big jump over the “dense” GPT3. Not really an apples-to-apples comparison.