3 ms·
My burning question: Why not also make a slightly larger model (100B) that could perform even better? Is there some bottleneck there that prevents RL from scal
by paradite 2y ago
My burning question: Why not also make a slightly larger model (100B) that could perform even better?
Is there some bottleneck there that prevents RL from scaling up performance to larger non-MoE model?